SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

時間: 2018年做成網站 https://chengyu.18dao.net 2024年用AI將文本生成圖片 2025年上傳到Hugging Face的Datasets 数据集中的文件总数: 20609 目录 "Text-to-Image/" 下的文件数量: 10296,子目錄數:5148,每個子目錄兩個文件,一個原始的文生圖png圖片,一個圖片解釋txt文件 目录 "image-chengyu/" 下的文件数量: 5155,加字的圖片jpg文件 目录 "text-chengyu/" 下的文件数量: 5156,文字解釋txt文件

Publicly accessible cc-by-nc-4.0 1K<n<10K

Dataset · Image to video

RealCam-Vid

MuteApo

25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid. - 25/03/26: Release our dataset RealCam-Vid v1 for metric-scale camera-controlled video generation, containing ~100K video clips with dedicated short/long captions and metric-scale camera annotations. - 25/02/18: Initial commit of the project, we plan to release the full dataset and data processing code in several weeks. DiT-based models (e.g., CogVideoX) trained on our dataset will be available at RealCam-I2V. Current datasets for camera-controllable video generation face critical limitations that hinder the development of robust and versatile models. Our curated dataset and data-processing pipeline uniquely…

Publicly accessible mit

Dataset · Question answering

piqa

Yonatan Bisk

To apply eyeshadow without a brush, should I use a cotton swab or a toothpick? Questions requiring this kind of physical commonsense pose a challenge to state-of-the-art natural language understanding systems. The PIQA dataset introduces the task of physical commonsense reasoning and a corresponding benchmark dataset Physical Interaction: Question Answering or PIQA. Physical commonsense knowledge is a major challenge on the road to true AI-completeness, including robots that interact with the world and understand natural language. PIQA focuses on everyday situations with a preference for atypical solutions. The dataset is inspired by instructables.com, which provides users with instructions…

Publicly accessible unknown 10K<n<100K

https://meta-math.github.io/ see our paper at https://arxiv.org/abs/2309.12284 All MetaMathQA data are augmented from the training sets of GSM8K and MATH. You can check the originalquestion in meta-math/MetaMathQA, each item is from the GSM8K or MATH train set. MetaMath-Mistral-7B is fully fine-tuned on the MetaMathQA datasets and based on the powerful Mistral-7B model. It is glad to see using MetaMathQA datasets and changing the base model from llama-2-7B to Mistral-7b can boost the GSM8K performance from 66.5 to 77.7. To fine-tune Mistral-7B, I would suggest using a smaller learning rate (usually 1/5 to 1/10 of the lr for LlaMa-2-7B) and staying other training args unchanged. More…

Publicly accessible mit

This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting, noise injection, pitch variation). The dataset is organized in folders per command label. Metadata files are included to facilitate training and evaluation. - One folder per speech command (e.g., yes/, no/, go/, stop/, etc.) - traininglist.txt - validationlist.txt…

Publicly accessible cc-by-4.0

Dataset

recipes

Feyn

This repository contains all the recipes that you can use with Chonkie to manage various documents, languages, and more. To use the recipes, you need to install chonkie with the hub feature, with the following command: This would enable Hubie which is used internally to get the recipes from this repository. So, you can do things like use the fromrecipe method to create a RecursiveChunker from a recipe. The above code would read the example.md file and chunk it into smaller chunks using the markdown recipe. The recipes are stored in the recipes folder. Each recipe is a JSON file that contains the rules for the recipe. The full JSON schema for the recipes is available here. You can use this…

Publicly accessible apache-2.0

Dataset · Text generation

SWE-rebench-V2

Nebius

SWE-rebench-V2 is a curated dataset of software-engineering tasks derived from real GitHub issues and pull requests. The dataset contains 32,079 samples covering Python, Go, TypeScript, JavaScript, Rust, Java, PHP, Kotlin, Julia, Elixir, Scala, Swift, Dart, C, C++, C#, R, Clojure, OCaml, and Lua. For log parser functions, base Dockerfiles, and the prompts used, please see https://github.com/SWE-rebench/SWE-rebench-V2 The detailed technical report is available at “SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale”. The dataset is licensed under the Creative Commons Attribution 4.0 license. However, please respect the license of each specific repository on which a particular…

Publicly accessible cc-by-4.0

Retargeted AMASS for Robotics Project Overview This project aims to retarget motion data from the AMASS dataset to various robot models and open-source the retargeted data to facilitate research and applications in robotics and human-robot interaction. AMASS (Archive of Motion Capture as Surface Shapes) is a high-quality human motion capture dataset, and the SMPL-X model is a powerful tool for generating realistic human motion data. By adapting the motion data from AMASS

Publicly accessible cc-by-4.0 10K<n<100K

All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes. I will update the data regularly, and you are welcome to follow this project and use the data. Each time the data is updated, I will record the update time in spec.json. Use DuckDB or Python API or Terminal or AI Agents, including Claude Desktop, Manus, Codex, Hermes Agent, OpenClaw Agent, and others, to access data. When using defeatbeta-api, version 0.0.61 or later is required. Earlier versions use the legacy flat data/ paths and are incompatible with the current market-specific directory layout. All…

Publicly accessible odc-by 100M<n<1B

2018年搭建成网站 https://zidian.18dao.net 2025年上傳到Hugging Face做成數據集。 目录 "image/" 下的文件数量: 4307,文生圖原始png圖片 目录 "image-zidian/" 下的文件数量: 4307,加字後的jpg圖片 目录 "text-zidian/" 下的文件数量: 4307,圖片解釋文字 目录 "pinyin/" 下的文件数量: 1702,拼音mp3文件

Publicly accessible cc-by-nc-4.0 1K<n<10K

This dataset is comprised of the LAMBADA test split as pre-processed by OpenAI (see relevant discussions here and here). It also contains machine translated versions of the split in German, Spanish, French, and Italian. LAMBADA is used to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative texts sharing the characteristic that human subjects are able to guess their last word if they are exposed to the whole text, but not if they only see the last sentence preceding the target word. To succeed on LAMBADA, computational models cannot simply rely on local context, but must be able to keep track of…

Publicly accessible mit 1K<n<10K

1. BigCodeBench-Complete: Code Completion based on the structured docstrings. 1. BigCodeBench-Instruct: Code Generation based on the NL-oriented instructions. The overall statistics of the dataset are as follows: The function-calling (tool use) statistics of the dataset are as follows: BigCodeBench is an easy-to-use benchmark which evaluates LLMs with practical and challenging programming tasks. The dataset was created as part of the BigCode Project, an open scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs). BigCodeBench serves as a fundamental benchmark for LLMs instead of LLM Agents, i.e., code-generating AI systems that enable…

Publicly accessible apache-2.0

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.