SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
Independent, daily-updated trust & safety data for the x402 agent-payment economy (HTTP 402 + stablecoins on Base) and the MCP server ecosystem — by PulseFeed. AI agents increasingly pay for APIs autonomously over x402 and connect to MCP servers that can run code on install. But 20% of listed x402 endpoints are dead or invalid, and "live" is not the same claim as "payable": of 24131 endpoints that return a valid 402, only 22426 across 1904 operators are things an agent can actually buy — the rest are documentation placeholders, testnet sandboxes priced in dollars, or payment schemes outside the x402 spec. Both figures are in ecosystem.json as healthyEndpoints and payableEndpoints; compare…
Private, append-only research data collected by hourlycryptoquant. - futurewindowl21s: contracts, snapshots - polymarketl21s: contracts, featuresnapshots - polymarketl25s: contracts, featuresnapshots - polymarketrtds: priceticks, structuralsnapshots - externalvenuel21s: venueseconds Rows are exported from SQLite as immutable Parquet shards. The repository is updated approximately every four hours. Paths are partitioned by source, table, and UTC date. rowid is the source-database row identifier used for incremental checkpointing. This dataset intentionally excludes environment files, credentials, wallet and account information, live orders, live executions, paper portfolios, logs, and…
Part of a dataset collection on Hugging Face. Daily health snapshots of the AST SpaceMobile BlueBird direct-to-cell satellite constellation, derived from CelesTrak GP (General Perturbations) data. Tracks the BlueWalker 3 prototype and all production BlueBird satellites as the constellation is built out. AST SpaceMobile is developing a unique LEO constellation designed to provide broadband connectivity directly to ordinary unmodified smartphones, without special satellite-phone hardware. The BlueBird satellites use unusually large phased-array antennas (up to 700 m2) to close the link budget with handset-sized antennas on the ground. The BlueWalker 3 prototype launched in 2022, the first…
One row per player, game, team, and statistic group. Regular-season games at AA, High-A, and Single-A, subject to configured coverage. Source: MLB Stats API. Data is extracted from box-score stats, never seasonStats. Missing statistics are null. pitchingIPstr uses baseball notation; use pitchingouts for arithmetic. Files are grouped by season, historical league ID, and team ID. catalog.json maps IDs to names and table paths. The train split is a loading convention; no training/test split has been applied. state/index.json records season completion; partially backfilled seasons can be visible. state/SEASON.json records the games successfully stored. Current-season games are refreshed…
MLEBench Goal and Flame Chase 6h Campaign Backup Private operational backup of the 2026-09-03 FlowBench-isolated MLEBench campaigns. It preserves the historical Goal and Flame Chase evidence and the fresh 75-task Kimi K3 run, 75 GLM-5.3 Goal cells, and 75 fixed-time GPT-5.6 Sol/Kimi K3 Flame Chase cells. Historical Kimi records do not count toward the fresh continuation result. At this snapshot, the
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.