The Dataset Viewer is configured to read data/train.jsonl (see the configs: section in the YAML header above). - The actual assets (JPG / PLY / SPZ) are stored under unsplash/ /. - The image field in data/train.jsonl stores the full HF resolve URL of the JPG for Dataset Viewer previews. imageid stays as the stable identifier. - Range locks (list + oldest): range coordination is stored on HF under ranges/locks, ranges/done, and ranges/progress. Each row in data/train.jsonl is a JSON object with stable (string) types for fields that commonly drift (to keep the Dataset Viewer working reliably). - unsplash/ /.jpg - unsplash/ /.ply - unsplash/ /.spz You can reconstruct URLs from ids: - gsplat…
SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
Project page Paper Code 3dcodebench.com arXiv:2606.01057 gaoypeng/3dcodebench News [06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code. Note. This is an open-source reproduction of 3DCodeBench. Under final check. The 3DCodeData/ code is still undergoing final quality review and may contain occasional issues (non-executable scripts, mismatched captions/renders, or imperfect geometry). If you run
SRTM 30m OZT2 Elevation Tiles This dataset contains SRTM 30-meter resolution elevation data encoded in the OZT2 tile format. Format OZT2 is a high-performance elevation tile format: Compression: ~93% smaller than Terrarium PNG Prediction: Gradient-based prediction (left neighbor + vertical gradient) Quantization: Adaptive bit-depth (8/10/12/16-bit per channel) Codec: Zstd q3 (30× faster encode than Brotli, same decode speed) Each tile is 256×256 pixels in Web
CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish, Portuguese, and Spanish, all at a sampling rate of 24kHz. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. - text-to-speech, text-to-audio: The dataset can also be used to train a model for Text-To-Speech (TTS). The…
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. tinystoriesalldata.tar.gz - contains a superset of the stories together with metadata and the prompt that was used to create each story. TinyStoriesV2-GPT4-train.txt - Is a new version of the dataset that is based on generations by GPT-4 only (the original dataset also has generations by GPT-3.5 which are of lesser quality). It contains all the…
InternData-A1 InternData-A1 is a hybrid synthetic-real manipulation dataset containing over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation. Your browser does not support the video tag. Your browser does not support the video tag.
A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research Blog Website Demo GitHub Paper Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in diverse contexts. This dataset enables the development of AI technologies that understand human interactions and communication, unlocking breakthroughs in: Explore the dataset with our interactive browser: We provide comprehensive download methods supporting all research scales…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.
