SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Snapshot of public HTTPS GET sources (USGS, EONET, ISS, NWS, GDACS, NOAA SWPC, Open-Meteo, OpenSky sample, launches, CoinGecko, Celestrak, lattice mirrors). - Every point is RESOURCE - Failed sources stay named SHADOW (no invented coordinates) - Not private intelligence. Not a live Star Chart write.

Publicly accessible mit

Per-checkpoint mechanistic metrics for Beetle language models, tracking how the induction circuit forms during training. The files here have six different schemas, so they are exposed as separate configs. Loading the directory as a single table fails with a cast error — that is why the configs above exist, not a bug. repeatdependence is the column that matters for deciding whether a high-PS head is really an induction head: candidates score ~0.7, every other head ~0.0003. A head that attends "somewhere earlier" can beat chance on PS alone without needing the repeat. A phase transition. PS sits at the chance floor (0.028, which is exactly uniform causal attention at the mean query position)…

Publicly accessible apache-2.0

Fac-similés d'éditions critiques de textes grecs anciens, rassemblés pour le projet TLG libre, qui reconstitue en TEI ouvert le corpus du Thesaurus Linguae Graecae. Chaque volume archivé ici porte une ou plusieurs œuvres du Canon TLG, et le C'est la clé d'entrée du dépôt. Une ligne par volume archivé: Un volume de la Patrologia Graeca ou des Fragmente der griechischen Historiker peut porter des dizaines d'œuvres: tlgids les liste toutes. Les fac-similés proviennent de dépôts publics: Internet Archive, Wikimedia Commons, Gallica, ANEMI, Diogeneia, la Bayerische Staatsbibliothek, Persée, l'Universitätsbibliothek Heidelberg et divers dépôts institutionnels DSpace. Le statut de droits déclaré…

Publicly accessible other

Dataset · Text classification

ida-dataset

Ary Wibowo

Continuous knowledge datasets from the IDA Dataset Factory. - {name}.csv · {name}.jsonl · {name}.parquet (when available) · README.md See repository LICENSE. Public knowledge from trusted open sources.

Publicly accessible other 1K<n<10K

Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judgeverdict (PERFECT/ERROR), judgereason, judgestatus, and judgeerror. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judgeerror; they are not OCR ERRORs. Direct Meta API, muse-spark-1.3-contributor, low reasoning effort. Reasons and verdicts are model judgments, not verified ground truth. The browser renderer is the same as in the first-three-shard reasons run, including its known readability limitations. No enrichment is…

Publicly accessible

Dataset · Tabular classification

sports-trends-dataset

Ruslan Magana Vsevolodovna

The data backbone of Ruslan Magana Sports Intelligence — refreshed automatically every day. Large data never lives in GitHub — it lives here, partitioned by sport / date / layer. - Raw is append-only and immutable — full provenance back to the source API. - Bronze/Silver normalize to one canonical schema with stable IDs and de-duplication. - Gold is analytics/ML-ready: a feature store plus train/validation/test splits. - ML data is partitioned Parquet (by sport/date/layer) — never one giant CSV. Every feature is computed leakage-safely — for a given fixture, only matches with date < fixture.matchdate are used: The prediction label is the realized match outcome (home/draw/away or 2-way per…

Publicly accessible mit 10K<n<100K

Dataset · Voice activity detection

semantic-vad-eot

Scicom (MSC) BHD

End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible with livekit/eot-bench-data. Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered silencespans. Per the eot-bench convention the last silence span is the true end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored). For every data type, all shards except the last form the train base; that type's last shard is split 50/25/25 into extra train / validation / test (shuffled, seed 42). - all (default) — every language + Malaysian subset. - 13 languages: en, fr, de, it, ja, ko, zh, pl, pt, ru, es, th, tr. - 5 Malaysian subsets (word-level, no…

Publicly accessible cc-by-4.0

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.