This dataset contains Norwegian speech data from NRK TV broadcasts (norge-rundt, supernytt), processed for automatic speech recognition (ASR) evaluation and research. - id: Unique chunk identifier - audio: Audio data (WAV format, 16kHz) - text: Original transcription text - textnormalized: Normalized text (lowercase, standardized) - durationseconds: Audio duration - chunktype: Type of chunk (subtitlealigned or fixedduration) - episodeid: Source episode identifier - programname: NRK program name - starttime: Start time in source video (seconds) - endtime: End time in source video (seconds) TBD - Please verify NRK terms of use before distribution. If you use this dataset, please cite: For…
SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-aiexpense. This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai.
Verified recovery SFT trajectories for visual web agents. - 2000 multi-turn conversations in policysft.jsonl - Benchmarks (as even as the accepted pool allowed): MiniWoB 639, TimeWarp 639, WebArena-Verified 362, VisualWebArena 360 Image files referenced by path live on the collection machine under /root/SFT200020260916/batches/. This JSONL is the training text; pair it with those PNGs if you train a vision model.
This dataset was created using LeRobot.
Teleoperated demonstration dataset for training a grasping policy on an SO-ARM101 arm with an AmazingHand dexterous hand (imitation learning). Pick up the cube with the dexterous hand — grasp a cube on the table using the dexterous hand. Collection flow: an operator moves the leader arm → the follower tracks it while its gripper proportionally drives the hand's open/close → joint angles and both camera streams are recorded synchronously. Per-episode flow: start from a fixed initial pose → perform one complete grasp (approach → grasp → lift → move → place) → end. Position coverage: the cube was placed at 4 different table positions, 5 episodes each, to teach positional generalization.…
This public pilot contains packed uint16 token IDs for pretraining the Tiny Base 300M, 2,048-context language model. It contains 50,000,000 training, 1,000,000 validation, and 1,000,000 test tokens. It does not redistribute raw documents. train.bin, validation.bin, and test.bin are little-endian uint16 streams. Use tokenizer.json with the Hugging Face tokenizers library. manifest.json records exact token counts, split policy, source revisions, mixture, and SHA-256 checksums; settings.json records preparation settings. The deterministic split is SHA-256 of normalized text modulo 1,000: 0–9 test, 10–19 validation, and 20–999 training. The mixture is 70% FineWeb-Edu, 25% English Wikipedia, and…
This public dataset contains 65,131 rendered counterfactual cases organized into 8,491 same-history action/future groups. Each case has 4 historical and 8 future CAMF0 frames plus WorldEngine metadata. Archives preserve complete pair groups. The audit bundle contains group JSON manifests, receipts, validation reports and the completion contract. Large generator-intermediate subset PKLs are intentionally excluded because the CAST paired-cache builder consumes the group JSON and rendered frames only. Download all files and run bash extractall.sh DESTINATION. The extraction script verifies every archive against SHA256SUMS first.
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.
