SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
This dataset contains a collection of files with randomized directory and file names. File bytes are preserved unchanged. Names and metadata embedded inside file contents or archives are not removed. Original file extensions are preserved, including.tar for archive shards. Archive contents and internal sample names are unchanged; existing WebDataset shards retain their format. Access requires manual approval by the repository owner. The original-path mapping is held locally by the owner and is not distributed in this repository. Ask the owner for provenance, format guidance, and applicable source licenses before reuse. This repository does not grant additional rights to the underlying…
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors. If you utilize this open intelligence dataset in your academic research, industrial benchmarking, or LLM training pipelines, please cite our repository: Maintained automatically by Sino Open Intelligence Network.
Daily meteorological observations from Brazilian observatories, 1883–1890, transcribed from printed nineteenth-century tables by an open 2B model running offline, with every row carrying its provenance and a quality verdict. readings a day), air temperature (mean, max, min, in-shelter and unsheltered), vapour tension, relative humidity, wind direction and force, cloudiness, rainfall, evaporation, ozone, and for Cuyabá the height of the Rio Cuyabá. Each row is one printed line: valuesasprinted (what the page shows, no silent correction), values (the publication's conventions undone in code, e.g. the barometer's elided leading digit restored), markers (words the page prints instead of a…
Part of a dataset collection on Hugging Face. Complete Rocket Lab launch manifest — past and upcoming missions flown or planned by Peter Beck's small-launch company, sourced from The Space Devs Launch Library 2 API. Covers Electron (small-lift two-stage rocket using Rutherford 3D-printed engines, operational since 2017 from Mahia Peninsula in New Zealand and from LC-2 at Wallops Island, Virginia) and the forthcoming Neutron medium-lift partially reusable rocket. Each row captures mission identifier, vehicle, launch time, status, pad location, target orbit, and a free-text mission description. Includes the full Electron history from the May 2017 'It's a Test' debut through the current…
This dataset was created using LeRobot.
A community-driven dataset for the Pa'O (blk) language, designed to support artificial intelligence, natural language processing (NLP), large language models (LLMs), conversational AI, and language technology research. The dataset focuses on three core Pa'O language resources: The primary goal is to build high-quality, reusable Pa'O language data for AI research, language technology, digital language preservation, and future Pa'O-capable AI systems. pao-ai-qa-conversation-instruction/ ├── data/ │ │ └── qa.jsonl │ ├── conversation/ │ │ └── conversations.jsonl │ └── instruction/ │ └── instructions.jsonl ├── README.md Question-and-answer pairs for Pa'O language understanding and…
Copyright (c) 2026 Toukir Studio. All Rights Reserved. These image assets are proprietary and created specifically for the LexiCore dictionary project. - Unauthorized copying, redistribution, scraping, modification, or republication of these assets is strictly prohibited. - These images may not be used for commercial purposes or included in any third-party training datasets without explicit written permission.
Research snapshot of official national legislation from Nepal Law Commission (lawcommission.gov.np). Not legal advice. The official gazette / authentic source prevails over this corpus. - data/laws.parquet — one row per instrument. - data/articles.parquet — article/section rows for the same snapshot. - scrapers/collectnp.py — research collector used to build this snapshot. Official Nepal Law Commission texts; authentic text prevails; not legal advice. This dataset is not legal advice. The official source prevails. - https://www.lawcommission.gov.np/ Thin deepen 2026-09-18 PT: Hub-restore tip 52 from endomorphosis/ipfsnepallaws@8ebcab98cea47106485663d85b5556c9845fd91a then offline…
MatrAIx Demo Application Data Evaluation data is grouped first by product surface (Type), then by task. Each task owns its persona profiles and one or more model/provider artifact folders. The canonical task definitions and implementation source are maintained in the MatrAIx GitHub application/tasks directory. Tasks Task Type Domain Folder GitHub Task Candy Land price sensitivity Survey Commerce View folder View task Annual checkup habits Survey
Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. Hallucination is Incorrectness OR Misgroundedness. Incompleteness and Substantial Correctness are independent of Hallucination. A response may be both Substantially Correct and Hallucinated. Hallucination × Substantial Correctness: H yes / SC yes = 269; H yes / SC no = 178; H no / SC yes = 130; H no / SC no = 35. Each cell is count (percentage; Wilson…
This dataset contains the public Gravitational Wave Open Science Center O4a strain release at 16,384 Hz for H1, L1. Each detector is represented independently. A contiguous source span produces three Parquet files: Strain, DQmask, and Injmask. This research has made use of data or software obtained from the Gravitational Wave Open Science Center (gwosc.org), a service of the LIGO Scientific Collaboration, the Virgo Collaboration, and KAGRA. This material is based upon work supported by NSF's LIGO Laboratory which is a major facility fully funded by the National Science Foundation, as well as the Science and Technology Facilities Council (STFC) of the United Kingdom, the Max-Planck-Society…
Daily-updated measurements of internet round-trip latency toward submarine cable landing regions, and detected routing anomalies (detours, path changes), collected by GeoCables - a live atlas of the world's submarine cable infrastructure with its own distributed measurement network. - geocables-latency-daily.csv - per-day latency aggregates (median / p90 / min RTT in milliseconds) for measured city-to-target routes. - geocables-route-events.csv - detected routing events: detours, reroutes and recoveries, with distance and RTT deltas (IP addresses scrubbed). Measurements are performed continuously by GeoCables' own distributed probe network; values are aggregated daily. Event detection…
This dataset was created using LeRobot.
Every forecast the OpenThomas weather desk read or produced, as-of. Pushed daily from the trading box, so each commit is a timestamped record of what was known when — the leak-free substrate behind the paper's replay experiments (docs/EXPERIMENTS.md in the replay/ holds the frozen decision-time rows (one JSONL per window, digest in the filename) and the experiment results scored on them. A row carries the settlement station and day, the baseline probability at the snapshot, the Kalshi bid/ask at that snapshot, and the outcome — never anything the agent could not have known at decision time. - replay/e3-high9.txt - replay/e3-low1.txt - replay/e3-low2.txt - replay/e3e4.log - replay/e3e4.sh…
Co-registered Sentinel-1 / Sentinel-2 pairs for self-supervised pre-training in Earth observation. Generated 2026-08-09. Every entry is a pair: one Sentinel-2 optical tile and one Sentinel-1 radar tile over the same ground, written to the same grid — identical bounding box, identical CRS (the tile's local UTM zone) and 10 m pixels — so the two rasters correspond pixel for pixel with no resampling on your side. The two acquisitions are at most three days apart; the realised gap is recorded per pair. Each tile is 512 × 512 pixels at 10 m, so it covers 5.12 × 5.12 km. The remaining Sentinel bands are not shipped, to keep the corpus trainable on ordinary hardware. They remain recoverable: every…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.
