SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

The Dataset Viewer is configured to read data/train.jsonl (see the configs: section in the YAML header above). - The actual assets (JPG / PLY / SPZ) are stored under unsplash/ /. - The image field in data/train.jsonl stores the full HF resolve URL of the JPG for Dataset Viewer previews. imageid stays as the stable identifier. - Range locks (list + oldest): range coordination is stored on HF under ranges/locks, ranges/done, and ranges/progress. Each row in data/train.jsonl is a JSON object with stable (string) types for fields that commonly drift (to keep the Dataset Viewer working reliably). - unsplash/ /.jpg - unsplash/ /.ply - unsplash/ /.spz You can reconstruct URLs from ids: - gsplat…

Publicly accessible other

Dataset · Text to 3d

3DCode

Yipeng Gao

Project page Paper Code 3dcodebench.com arXiv:2606.01057 gaoypeng/3dcodebench News [06/01/2026] Paper released on arXiv: 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code. Note. This is an open-source reproduction of 3DCodeBench. Under final check. The 3DCodeData/ code is still undergoing final quality review and may contain occasional issues (non-executable scripts, mismatched captions/renders, or imperfect geometry). If you run

Publicly accessible mit 10K<n<100K

Dataset · Image generation

srtm30m-ozt2-v2

Aliasfox

SRTM 30m OZT2 Elevation Tiles This dataset contains SRTM 30-meter resolution elevation data encoded in the OZT2 tile format. Format OZT2 is a high-performance elevation tile format: Compression: ~93% smaller than Terrarium PNG Prediction: Gradient-based prediction (left neighbor + vertical gradient) Quantization: Adaptive bit-depth (8/10/12/16-bit per channel) Codec: Zstd q3 (30× faster encode than Brotli, same decode speed) Each tile is 256×256 pixels in Web

Publicly accessible n<1K

Dataset · Text to speech

cml-tts

Yoach Lacombe

CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish, Portuguese, and Spanish, all at a sampling rate of 24kHz. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. - text-to-speech, text-to-audio: The dataset can also be used to train a model for Text-To-Speech (TTS). The…

Publicly accessible cc-by-4.0 1M<n<10M

Dataset · Text generation

TinyStories

Ronen Eldan

Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. tinystoriesalldata.tar.gz - contains a superset of the stories together with metadata and the prompt that was used to create each story. TinyStoriesV2-GPT4-train.txt - Is a new version of the dataset that is based on generations by GPT-4 only (the original dataset also has generations by GPT-3.5 which are of lesser quality). It contains all the…

Publicly accessible cdla-sharing-1.0

Dataset · Other

InternData-A1

Intern Robotics

InternData-A1 InternData-A1 is a hybrid synthetic-real manipulation dataset containing over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation. Your browser does not support the video tag. Your browser does not support the video tag.

Access requested at publisher n>1T

A large-scale multimodal dataset of 4,000+ hours of human interactions for AI research Blog Website Demo GitHub Paper Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. The Seamless Interaction Dataset is a large-scale collection of over 4,000 hours of face-to-face interaction footage from more than 4,000 participants in diverse contexts. This dataset enables the development of AI technologies that understand human interactions and communication, unlocking breakthroughs in: Explore the dataset with our interactive browser: We provide comprehensive download methods supporting all research scales…

Publicly accessible cc-by-nc-4.0

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.