SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…
prettyname: JesseLikesWeather API
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test This repository contains the benchmarks, generated data, and evaluation logs for CoSPlay, a training-free framework that jointly improves code generation and unit tests through cooperative self-play at inference time. Paper: CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test CoSPlay
See https://github.com/ust-archive/ust-archive for more information.
Testing new way to compress images All file contains 20,000 WebP images in 1536 resolutions instead
The DiSCo substrate (Rafael-Patino, Girard, Truffet, Pizzolato, Caruyer, Thiran, The diffusion-simulated connectivity (DiSCo) dataset, Data in Brief 38 (2021) 107429, doi:10.1016/j.dib.2021.107429; data doi:10.17632/fgf86jdfg6.3, CC BY 4.0) walked once and stored as a replay pack, so that any acquisition a human scanner can play is a replay of the same walk, voxel by voxel on the dataset's own 40³ grid of 25 µm voxels, with a per-voxel Monte-Carlo certificate. DiSCo published one acquisition of this substrate; this dataset is the substrate itself in replayable form: the same 12,196 strands, two tubes per strand (the listed inner diameter and the outer tube at 1/0.7 of it), the dataset's…
Companion artifacts for the unified jobs search Space. Contains: titles + slim metadata + full metadata (JSONL) + bge-small + te3-large @ 1024 catalog vectors + pre-encoded te3 query cache (~196k popular queries) for 347,900 job postings across 4 corpora (Open-Apply, LinkedIn, JobStreet, USAJobs).
Every game played on Faïence, a free browser implementation of the rules of Azul (Michael Kiesling) against a neural net trained by self-play, unless the player switched sharing off. This dataset is the training pile the playing page tells its players about, and it is public precisely so that a player can read everything the project collects. Records are anonymous by construction: moves, deals, which net played, and the score. No names, no accounts, no IPs, no user agents. games/YYYY-MM-DD/ -.jsonl, one file per ingest batch, one JSON object per line. Nothing is ever rewritten; new batches only add files. Each line is a canonical record rebuilt by the collector which replayed the game in…
This repository hosts tmax-compatible SIF images and a unified download manifest. Training data and task archives are in hamishivi/agent-task-recursive-task-synthesis. The manifest includes earlier images hosted under hamishivi and new images hosted under TMaxxx; the downloader selects the correct repository and immutable commit for each image. The pool currently contains 28,646 / 29,501 verified Apptainer SIF images. Shared environments are stored once. Each published SIF passed apptainer inspect, a contained shell smoke check, and remote SHA256 verification. These checks do not constitute a full task evaluation. Download apptainer/downloadapptainer.py and run with Python 3.11+ and…
Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. One config per transcription system, so their output stays separable. The combined config interleaves several transcription systems within each shard, so this is the breakdown across the whole dataset. The systems differ substantially, so treat them as separate sources when training. Each row carries the system that produced it in its provider and model columns; the labels below are withheld aliases for the same systems, in…
A massive, highly-diverse dataset of Fused Deposition Modeling (FDM) G-codes generated directly from Printables. This dataset is specifically designed for training machine learning models on raw 3D printing manufacturing instructions (G-code). It can be used for tasks like G-code generation, print failure prediction, semantic analysis of toolpaths, and printer-agnostic slice classification. To ensure a highly robust and diverse set of training data, every original 3D model from the source dataset has been iteratively sliced into multiple variants (default: 10 per model) using PrusaSlicer. The pipeline supports and natively slices.stl,.obj,.3mf,.step, and.amf files. For each variant, a…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.

