SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
Snapshot of public HTTPS GET sources (USGS, EONET, ISS, NWS, GDACS, NOAA SWPC, Open-Meteo, OpenSky sample, launches, CoinGecko, Celestrak, lattice mirrors). - Every point is RESOURCE - Failed sources stay named SHADOW (no invented coordinates) - Not private intelligence. Not a live Star Chart write.
Per-checkpoint mechanistic metrics for Beetle language models, tracking how the induction circuit forms during training. The files here have six different schemas, so they are exposed as separate configs. Loading the directory as a single table fails with a cast error — that is why the configs above exist, not a bug. repeatdependence is the column that matters for deciding whether a high-PS head is really an induction head: candidates score ~0.7, every other head ~0.0003. A head that attends "somewhere earlier" can beat chance on PS alone without needing the repeat. A phase transition. PS sits at the chance floor (0.028, which is exactly uniform causal attention at the mean query position)…
Fac-similés d'éditions critiques de textes grecs anciens, rassemblés pour le projet TLG libre, qui reconstitue en TEI ouvert le corpus du Thesaurus Linguae Graecae. Chaque volume archivé ici porte une ou plusieurs œuvres du Canon TLG, et le C'est la clé d'entrée du dépôt. Une ligne par volume archivé: Un volume de la Patrologia Graeca ou des Fragmente der griechischen Historiker peut porter des dizaines d'œuvres: tlgids les liste toutes. Les fac-similés proviennent de dépôts publics: Internet Archive, Wikimedia Commons, Gallica, ANEMI, Diogeneia, la Bayerische Staatsbibliothek, Persée, l'Universitätsbibliothek Heidelberg et divers dépôts institutionnels DSpace. Le statut de droits déclaré…
Continuous knowledge datasets from the IDA Dataset Factory. - {name}.csv · {name}.jsonl · {name}.parquet (when available) · README.md See repository LICENSE. Public knowledge from trusted open sources.
Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judgeverdict (PERFECT/ERROR), judgereason, judgestatus, and judgeerror. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judgeerror; they are not OCR ERRORs. Direct Meta API, muse-spark-1.3-contributor, low reasoning effort. Reasons and verdicts are model judgments, not verified ground truth. The browser renderer is the same as in the first-three-shard reasons run, including its known readability limitations. No enrichment is…
The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day. Large data never lives in GitHub — it lives here, partitioned by sport / date / layer. - Raw is append-only and immutable — full provenance back to the source API. - Bronze/Silver normalize to one canonical schema with stable IDs and de-duplication. - Gold is analytics/ML-ready: a feature store plus train/validation/test splits. - ML data is partitioned Parquet (by sport/date/layer) — never one giant CSV. Every feature is computed leakage-safely — for a given fixture, only matches with date < fixture.matchdate are used: The prediction label is the realized match outcome (home/draw/away or 2-way per…
End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible with livekit/eot-bench-data. Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered silencespans. Per the eot-bench convention the last silence span is the true end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored). For every data type, all shards except the last form the train base; that type's last shard is split 50/25/25 into extra train / validation / test (shuffled, seed 42). - all (default) — every language + Malaysian subset. - 13 languages: en, fr, de, it, ja, ko, zh, pl, pt, ru, es, th, tr. - 5 Malaysian subsets (word-level, no…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.

