SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Video classification

RekaDaily-10k-raw

Reka AI

Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix. This is the full RekaDaily-10k release: 10,865 hours / 412,050 videos / 11,127 shards / 80 TB. See the Additional recordings may be appended over time; the metadata/ tables and this line are updated whenever that happens. The Dataset…

Publicly accessible apache-2.0 100K<n<1M

Dataset · Question answering

MMLU-Pro

TIGER-Lab

MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. - \[2026.03.11\] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among others. Stay tuned for updated rankings and analysis. - \[2026.01.18\] Fixed leading space issue in answer options (affected chemistry, physics, and other STEM subsets). This formatting inconsistency could have been exploited as a shortcut. Thanks to @giffmana and @fujikanaeda for identifying…

Publicly accessible mit 10K<n<100K

Dataset · Robotics

AgiBotWorld2026

AgiBot World

Real-World Embodied Intelligence Dataset As robotics research advances into real-world scenarios, the demand for authentic, high-quality data has become increasingly urgent. Following AGIBOT WORLD's "ImageNet moment," we now release the AGIBOT WORLD 2026 dataset. Built upon massive real-world scenes, it systematically spans pivotal research directions in embodied intelligence, designed to power the next generation of embodied agents. The AGIBOT WORLD 2026 dataset is collected from 100% real-world environments, covering commercial spaces, home, and other general-purpose scenarios. Collected on the AGIBOT G2 robot platform through a free-form collection mode, the dataset provides developers…

Publicly accessible cc-by-nc-sa-4.0 1K<n<10K

A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents" This dataset contains around 200M samples in PIN format, with around 312 TB storage. News [ 2025.09.22 ]!NEW! We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it. [ 2024.12.06 ]!NEW! We have updated the quality signals, enabling a swift assessment of whether a sample meets the required specifications based on our quality indicators. Further detailed descriptions will be provided in the forthcoming formal publication. (Aside from the Chinese-Markdown subset, there are unresolved issues that are currently being addressed.) This dataset…

Publicly accessible apache-2.0 100M<n<1B

FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision Each of the publicly available sub-datasets present in FineVision are governed by specific licensing conditions. Therefore, when making use of them you must take into consideration each of the licenses governing each dataset. To the extent we have any rights in the prompts, these are licensed under CC-BY-4.0. If you find this dataset useful, please cite

Publicly accessible 10M<n<100M

Dataset · Robotics

stereo-550

FPV Labs

FPV Labs Open-Source Stereo Hardware A first-person calibrated stereo RGB video dataset capturing everyday human manipulation across objects, materials, tools, and multi-step activities. Every session is recorded as a synchronized left/right camera pair with per-session stereo calibration, giving the visual geometry of hands, object interaction, state

Access requested at publisher other 1K<n<10K

Production-ready pipeline (Python package videovec2wav2tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio Speech recognition — faster-whisper, CPU & CUDA, automatic language detection, word-level timestamps. Segmentation — cut audio by transcript timestamps into dataset/audio/000001.wav …. Dataset generation — metadata.csv, dataset.jsonl, ttsmetadata.csv. Feature extraction (optional) — streaming features/train.bin + train.dat with float32 samples, mel spectrograms, duration and sample rate. Statistics…

Publicly accessible

Dataset · Text generation

MATH-500

Hugging Face H4

This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits

Publicly accessible

Dataset

iDATA

AiEDA

Dataset A dataset of AI + EDA iDATA is a dataset of AI + EDA, which can be used to train AI models for design PPA prediction, PPA-aware physical design, and related tasks. Dataset structure Describe the dataset structure. aes/ ├── iEDArouteprocessdata/ # Process data exported by iEDA-iRT 2D routing ├── synnetlist/ # The synthesized netlist files、sdc files ├── place/ # The place stage def、sdc、vectors └── route/ # The route

Publicly accessible gpl

Dataset · Text generation

SWE-smith

SWE-bench

[12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the SWE-bench/SWE-smith-[lang] datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo. The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit. It is the largest dataset to date for training software engineering agents. All SWE-smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the…

Publicly accessible mit 10K<n<100K

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.