SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix. This is the full RekaDaily-10k release: 10,865 hours / 412,050 videos / 11,127 shards / 80 TB. See the Additional recordings may be appended over time; the metadata/ tables and this line are updated whenever that happens. The Dataset…
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. - \[2026.03.11\] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among others. Stay tuned for updated rankings and analysis. - \[2026.01.18\] Fixed leading space issue in answer options (affected chemistry, physics, and other STEM subsets). This formatting inconsistency could have been exploited as a shortcut. Thanks to @giffmana and @fujikanaeda for identifying…
Real-World Embodied Intelligence Dataset As robotics research advances into real-world scenarios, the demand for authentic, high-quality data has become increasingly urgent. Following AGIBOT WORLD's "ImageNet moment," we now release the AGIBOT WORLD 2026 dataset. Built upon massive real-world scenes, it systematically spans pivotal research directions in embodied intelligence, designed to power the next generation of embodied agents. The AGIBOT WORLD 2026 dataset is collected from 100% real-world environments, covering commercial spaces, home, and other general-purpose scenarios. Collected on the AGIBOT G2 robot platform through a free-form collection mode, the dataset provides developers…
A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents" This dataset contains around 200M samples in PIN format, with around 312 TB storage. News [ 2025.09.22 ]!NEW! We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it. [ 2024.12.06 ]!NEW! We have updated the quality signals, enabling a swift assessment of whether a sample meets the required specifications based on our quality indicators. Further detailed descriptions will be provided in the forthcoming formal publication. (Aside from the Chinese-Markdown subset, there are unresolved issues that are currently being addressed.) This dataset…
FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision Each of the publicly available sub-datasets present in FineVision are governed by specific licensing conditions. Therefore, when making use of them you must take into consideration each of the licenses governing each dataset. To the extent we have any rights in the prompts, these are licensed under CC-BY-4.0. If you find this dataset useful, please cite
FPV Labs Open-Source Stereo Hardware A first-person calibrated stereo RGB video dataset capturing everyday human manipulation across objects, materials, tools, and multi-step activities. Every session is recorded as a synchronized left/right camera pair with per-session stereo calibration, giving the visual geometry of hands, object interaction, state
Production-ready pipeline (Python package videovec2wav2tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio Speech recognition — faster-whisper, CPU & CUDA, automatic language detection, word-level timestamps. Segmentation — cut audio by transcript timestamps into dataset/audio/000001.wav …. Dataset generation — metadata.csv, dataset.jsonl, ttsmetadata.csv. Feature extraction (optional) — streaming features/train.bin + train.dat with float32 samples, mel spectrograms, duration and sample rate. Statistics…
MATH dataset from https://github.com/hendrycks/math
This repo contains all the docs published on https://huggingface.co/docs. The docs are generated with https://github.com/huggingface/doc-builder.
This dataset contains a subset of 500 problems from the MATH benchmark that OpenAI created in their Let's Verify Step by Step paper. See their GitHub repo for the source file: https://github.com/openai/prm800k/tree/main?tab=readme-ov-file#math-splits
Dataset A dataset of AI + EDA iDATA is a dataset of AI + EDA, which can be used to train AI models for design PPA prediction, PPA-aware physical design, and related tasks. Dataset structure Describe the dataset structure. aes/ ├── iEDArouteprocessdata/ # Process data exported by iEDA-iRT 2D routing ├── synnetlist/ # The synthesized netlist files、sdc files ├── place/ # The place stage def、sdc、vectors └── route/ # The route
This dataset was created using LeRobot.
[12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the SWE-bench/SWE-smith-[lang] datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo. The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit. It is the largest dataset to date for training software engineering agents. All SWE-smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.




