SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Multiple choice

race

Eduard Hovy

RACE is a large-scale reading comprehension dataset with more than 28,000 passages and nearly 100,000 questions. The dataset is collected from English examinations in China, which are designed for middle school and high school students. The dataset can be served as the training and test sets for machine comprehension. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - exampleid: a string feature. - article: a string feature. - answer: a string feature. - question: a string feature. - options: a list of string features. - exampleid: a string feature. - article: a string…

Publicly accessible other 10K<n<100K

A set of badges you can use anywhere. Just update the anchor URL to point to the correct action for your Space. Light or dark background with 4 sizes available: small, medium, large, and extra large. - With markdown, just copy the badge from: https://huggingface.co/datasets/huggingface/badges/blob/main/README.md?code=true - With HTML, inspect this page with your web browser and copy the outer html.

Publicly accessible mit

Dataset · Text to video

seedance-2-prompts-datasets

MaBiao

This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due to GitHub's limitations with large file storage, the full dataset is hosted on Hugging Face. The Hugging Face repository contains the generated videos (.mp4), cover images (.jpg), and a highly structured.jsonl file that holds all prompt metadata. No login required, lighting-fast response. Launched by GokuOpenLab, seedance-2-prompts-datasets is a prompt data infrastructure project created for developers and researchers. In the current AI ecosystem, prompts are the new…

Publicly accessible cc-by-4.0 1K<n<10K

Dataset · Other

ngii-map-full-light

Solkyu Park

ngii-map-full-light Light point/line extract from NGII 1/1000 topographic data for Korea. Not for shipping into GitHub — use this Hugging Face dataset instead. CRS Korea2000CentralBelt2010 projected meters [x, y] Layers (per region under byregion/ /) Layer Description C023 poles (전주/통신주) C022 lights (가로등·보안등) A002 roads (도로 중심선) B001tiny building footprints <25 m² as centroids B002 lines (구분/재질 라인) Also

Publicly accessible other 1M<n<10M

Dataset · Visual question answering

ViFailback-Dataset

SII-RHOS

ViFailback Dataset: Real-World Robotic Manipulation Failure Dataset with Visual Symbol Guidance A real-world dataset for diagnosing, correcting, and learning from robotic manipulation failures via visual symbols. ViFailback is a large-scale, real-world robotic manipulation failure dataset introduced in the CVPR 2026 paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols". It introduces visual

Publicly accessible mit

Dataset · Fill mask

HPLT2.0_cleaned

HPLT

We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. This is the variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here. The original JSONL files (which take ~4x fewer disk space than this HF version) and the larger non-cleaned version can be found at https://hplt-project.org/datasets/v2.0. We conducted the FineWeb-style…

Publicly accessible cc0-1.0 n>1T

Dataset · Other

P3

BigScience Workshop

P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If {Premise} is true, is it also true that {Hypothesis}?, whereas a target template can be defined with the label choices Choices[label]. Here Choices is prompt-specific metadata that consists of the options yes, maybe, no corresponding to label being entailment (0)…

Publicly accessible apache-2.0 100M<n<1B

Dataset · Text classification

ceval-exam

Ceval

C-Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi-choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details. Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few-shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model evaluation. More details on loading and using the data are at our github page. Please cite our paper if you use our dataset.

Publicly accessible cc-by-nc-sa-4.0 10K<n<100K

Dataset · Question answering

trivia_qa

Mandar Joshi

TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions. English. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - question: a string feature. - questionid: a string feature. - questionsource: a string feature. - entitypages: a dictionary feature containing…

Publicly accessible unknown 10K<n<100K

Dataset · Text generation

MegaMath

Institute of Foundation Models

We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on the Internet. We identified high quality math-related code from large code training corpus, Stack-V2, further enhancing data diversity. We synthesized QA-style text, math-related code, and interleaved text-code blocks from web data or code data. MegaMath is the largest open math pre-training dataset to date, surpassing DeepSeekMath (120B)…

Publicly accessible odc-by 1B<n<10B

Dataset · Text generation

UbuntuIRC

Nils

Completely uncurated collection of IRC logs from the Ubuntu IRC channels

Publicly accessible cc0-1.0

Dataset · Video classification

EgoDemo

Lightwheel

EgoDemo A 50-hour sample from EgoSuite-Open100K, covering every annotated subset plus two EgoSuite-Open100K ↗ EgoSuite-Open100K Overview Collection: EgoSuite-Open100K SKU Sub-SKU Format Planned Duration EgoStandard EgoStand Hand Pose 80,000 h

Access requested at publisher other 1K<n<10K

Dataset · Text classification

blimp

NYU Machine Learning for Language

BLiMP is a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing 1000 minimal pairs isolating specific contrasts in syntax, morphology, or semantics. The data is automatically generated according to expert-crafted grammars. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - sentencegood: a string feature. - sentencebad: a string feature. - field: a string feature. - linguisticsterm: a string…

Publicly accessible cc-by-4.0 10K<n<100K

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.