SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Text classification

snli

Stanford NLP

The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE). Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), is the task of determining the inference relation between two (short, ordered) texts: entailment, contradiction, or neutral (MacCartney and Manning 2008). See the corpus webpage for a list of published results. The language in the dataset is English as spoken by users of the website Flickr and as spoken by…

Publicly accessible cc-by-sa-4.0 100K<n<1M

Dataset · Question answering

hotpot_qa

Hotpotqa

HotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) the questions require finding and reasoning over multiple supporting documents to answer; (2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas; (3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason with strong supervision and explain the predictions; (4) we offer a new type of factoid comparison questions to test QA systems’ ability to extract relevant facts and perform necessary comparison. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields…

Publicly accessible cc-by-sa-4.0 100K<n<1M

Dataset · Other

AgiBotWorld-Beta

AgiBot World

1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours. - 100+ real-world scenarios across 5 target domains. - 200+ types of tasks: - 87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc. Your browser does not support the video tag. Your browser does not support the video tag. Your browser does not support the video tag. - [2025/3/1] AgiBot World Beta released. - [ ] AgiBot World Colosseum:Comprehensive platform (expected release date: 2025) - [ ] 2025 AgiBot World Challenge (expected release date: 2025) To download the full dataset, you can use the following code. If you encounter any issues, please refer to the official Hugging Face documentation. If…

Access requested at publisher n>1T

We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973 Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each part. Pixabay has also been compressed into multiple parts. After extracting them, all videos should be placed into a single folder. For SAM data, please download from the official link. After downloading 1000 compressed files, extract all the…

Publicly accessible mit

Dataset · Speech recognition

fleurs

Google

Universal Representations of Speech](https://arxiv.org/abs/2205.12446) Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven geographical areas: The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The dataset can be downloaded and prepared…

Publicly accessible cc-by-4.0 10K<n<100K

Dataset · Video classification

xperience-10m

Ropedia

Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified. Interactive Intelligence from Human Xperience Xperience-10M Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial

Access requested at publisher other 1M<n<10M

Dataset · Text to image

monet

Jasper AI

MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 103.8 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with multiple vision-language models, and is further augmented with synthetically generated samples. Each image is released with pre-computed embeddings, structured annotations and pre-encoded VAE latents to accelerate downstream use. A 4B-parameter…

Publicly accessible apache-2.0 100M<n<1B

SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset only contains the problemstatement (i.e. issue text) and the basecommit which can represents the state of the codebase before the issue has been resolved. If you want to run inference using the "Oracle" or BM25 retrieval settings mentioned in the paper, consider the following datasets.…

Publicly accessible

Dataset · Text generation

AIME_2024

Minghui Jia

This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2024. AIME is a prestigious high school mathematics competition known for its challenging mathematical problems. Each record contains the following fields: - ID: Problem identifier (e.g., "2024-I-1" represents Problem 1 from 2024 Contest I) - Problem: Problem statement - Solution: Detailed solution process - Answer: Final numerical answer This dataset is primarily used for: 1. Evaluating Large Language Models' (LLMs) mathematical reasoning capabilities 2. Testing models' problem-solving abilities on complex mathematical problems 3. Researching AI performance on structured mathematical tasks - Covers…

Publicly accessible mit n<1K

Dataset · Text generation

Ultra-FineWeb

OpenBMB

English | Ultra-FineWeb is a large-scale, high-quality, and efficiently-filtered dataset. We use the proposed efficient verification-based high-quality filtering pipeline to the FineWeb and Chinese FineWeb datasets (source data from Chinese FineWeb-edu-v2, which includes IndustryCorpus2, MiChao, WuDao, SkyPile, WanJuan, ChineseWebText, TeleChat, and CCI3), resulting in the creation of higher-quality Ultra-FineWeb-en with approximately 1T tokens, and Ultra-FineWeb-zh datasets with approximately 120B tokens, collectively referred to as Ultra-FineWeb. Ultra-FineWeb serves as a core pre-training web dataset for the MiniCPM4 Series and MiniCPM5 Series models. - Ultra-FineWeb-L1: L1 filtered data…

Publicly accessible apache-2.0 n>1T

Dataset

aime25

Math AI

If you use the AIME25 dataset in your research, please consider citing it as follows

Publicly accessible apache-2.0

AI-MO Olympiad Reference Dataset This dataset contains a structured collection of Olympiad problems and their solutions, organized by competition. Contains high quality data, prioritizing "official" solutions to problems. Structure / # Problems and solutions from the International Mathematical Olympiad ├── raw/ # Raw problem/solution statements (.pdf) │ ├── file1.pdf │ ├── file2.pdf ├── downloadscript/ # the scripts used to

Publicly accessible

Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and alt texts for images in processed samples. This public extract should contain 500-650GT depending on the tokenizer you use, and can be enhanced with the curated corpora of your choosing. This public extract is about ~500GB to download…

Publicly accessible odc-by 100B<n<1T

Dataset · Question answering

Nemotron-Terminal-Corpus

NVIDIA

Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. The high-quality trajectories in Terminal-Corpus enable models of various sizes to achieve performance that rivals or exceeds much larger frontier models on the Terminal-Bench 2.0 benchmark. Training on Terminal-Corpus yields substantial gains across the Qwen3 model family: The Nemotron-Terminal-32B (27.4%) outperforms the 480B-parameter Qwen3-Coder (23.9%) and…

Publicly accessible cc-by-4.0 100K<n<1M

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.