SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
Barcode to product name, description, product image, retailer and safety data sheet link, harvested from Danish web shops. Product text is in Danish. The images are stored as bytes under images/, not as links. A URL is a promise: the retailer can change CDN tomorrow, and the dataset would be empty without looking empty. The originals, untouched. Every image is stored byte for byte as the retailer sent it — same format, same resolution. We do not re-encode and we do label text, a downscaled JPEG is too little. The format column says what the row actually is (jpg, png, webp, avif, gif). Pillow is opened ONLY to decide whether the file is an image at all, and if so which kind — the content…
Calibration parameters from IBM Quantum Heron processors joined to ambient and space-weather conditions at the time of each measurement. Designed for time-series forecasting of qubit drift and for studying environmental coupling to superconducting calibrations. A GitHub Action (source) polls backend.properties() on every available Heron device every 30 minutes and appends new calibration events keyed on (backend, property, qubita, qubitb, calibratedtime). All backends are at IBM's Yorktown Heights, NY facility (lat=41.27, lon=-73.78). Coherence. T1, T2 per qubit, in microseconds. Single-qubit gate errors. sxerror, xerror, iderror, rxerror, rzerror. Reported separately by the IBM API; on a…
Датасет собирается инкрементально в течение нескольких сессий бесплатного Google Colab. - Qwen2.5-1.5B-Instruct - Qwen2.5-3B-Instruct - SmolLM2-1.7B-Instruct Qwen/Qwen2.5-3B-Instruct questionid, question, modela/answera, modelb/answerb, preferred (A/B), preferredmodel, preferencereason Датасет пополняется по мере запуска сборочного ноутбука. Чекпоинты сырых данных лежат в папке raw/ этого репозитория. При каждом новом запуске ноутбук подтягивает уже накопленные данные и продолжает сбор с недостающих вопросов и пар ответов, не дублируя уже сделанную работу.
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Every alignment number in this dataset is only as meaningful as the brain RDMs it was computed against. So before any model result, the same pipeline is asked whether anything stimulus-driven correlates with those RDMs — stimulus duration, intensity, word length, frequency, phoneme and syllable counts, an acoustic model of the audio where the stimuli are audio, and the study's own condition contrast — each tested by a permutation test that shuffles stimulus identity. correction — not the acoustic model of the audio the children actually heard, not the study's own experimental…
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for long-term preservation and reproducibility. Layout mirrors the verified backup trees they were copied from: - tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz, 000.json) list the original source paths. - taurus/data1/ /: per-run backup roots with their own progress and receipt JSON files. - mac/critic-hack-recovery-20260916/: incremental archives with sibling manifest JSON. MANIFEST.jsonl at…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.

