RACE is a large-scale reading comprehension dataset with more than 28,000 passages and nearly 100,000 questions. The dataset is collected from English examinations in China, which are designed for middle school and high school students. The dataset can be served as the training and test sets for machine comprehension. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - exampleid: a string feature. - article: a string feature. - answer: a string feature. - question: a string feature. - options: a list of string features. - exampleid: a string feature. - article: a string…
SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
A set of badges you can use anywhere. Just update the anchor URL to point to the correct action for your Space. Light or dark background with 4 sizes available: small, medium, large, and extra large. - With markdown, just copy the badge from: https://huggingface.co/datasets/huggingface/badges/blob/main/README.md?code=true - With HTML, inspect this page with your web browser and copy the outer html.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due to GitHub's limitations with large file storage, the full dataset is hosted on Hugging Face. The Hugging Face repository contains the generated videos (.mp4), cover images (.jpg), and a highly structured.jsonl file that holds all prompt metadata. No login required, lighting-fast response. Launched by GokuOpenLab, seedance-2-prompts-datasets is a prompt data infrastructure project created for developers and researchers. In the current AI ecosystem, prompts are the new…
ngii-map-full-light Light point/line extract from NGII 1/1000 topographic data for Korea. Not for shipping into GitHub — use this Hugging Face dataset instead. CRS Korea2000CentralBelt2010 projected meters [x, y] Layers (per region under byregion/ /) Layer Description C023 poles (전주/통신주) C022 lights (가로등·보안등) A002 roads (도로 중심선) B001tiny building footprints <25 m² as centroids B002 lines (구분/재질 라인) Also
This dataset was created using LeRobot.
ViFailback Dataset: Real-World Robotic Manipulation Failure Dataset with Visual Symbol Guidance A real-world dataset for diagnosing, correcting, and learning from robotic manipulation failures via visual symbols. ViFailback is a large-scale, real-world robotic manipulation failure dataset introduced in the CVPR 2026 paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols". It introduces visual
We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. This is the variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here. The original JSONL files (which take ~4x fewer disk space than this HF version) and the larger non-cleaned version can be found at https://hplt-project.org/datasets/v2.0. We conducted the FineWeb-style…
P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If {Premise} is true, is it also true that {Hypothesis}?, whereas a target template can be defined with the label choices Choices[label]. Here Choices is prompt-specific metadata that consists of the options yes, maybe, no corresponding to label being entailment (0)…
C-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details. Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model evaluation. More details on loading and using the data are at our github page. Please cite our paper if you use our dataset.
https://huggingface.co/spaces/HorizonRobotics/EmbodiedGen-Gallery-Explorer
TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions. English. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - question: a string feature. - questionid: a string feature. - questionsource: a string feature. - entitypages: a dictionary feature containing…
We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on the Internet. We identified high quality math-related code from large code training corpus, Stack-V2, further enhancing data diversity. We synthesized QA-style text, math-related code, and interleaved text-code blocks from web data or code data. MegaMath is the largest open math pre-training dataset to date, surpassing DeepSeekMath (120B)…
Completely uncurated collection of IRC logs from the Ubuntu IRC channels
EgoDemo A 50-hour sample from EgoSuite-Open100K, covering every annotated subset plus two EgoSuite-Open100K ↗ EgoSuite-Open100K Overview Collection: EgoSuite-Open100K SKU Sub-SKU Format Planned Duration EgoStandard EgoStand Hand Pose 80,000 h
BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - sentencegood: a string feature. - sentencebad: a string feature. - field: a string feature. - linguisticsterm: a string…
ReActor Assets The Fast and Simple Face Swap Extension Models
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.
