SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Time series forecasting

caliceo

Julien Chaumond

src="https://julien-c-caliceo.static.hf.space" frameborder="0" width="100%" height="740"

Publicly accessible mit

This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.

Publicly accessible

Versioned public tables and search indexes for - companies and dated company snapshots; - ASX announcement metadata and public artifact URLs; - capital instruments and instrument events; - screener snapshots; and - a revision-matched SQLite full-text index. Each release will identify its exact Hub commit, source watermark, table counts, file hashes, schema version, publisher version and exclusion-manifest version. The target update cadence is daily after the 06:00 Australia/Melbourne source The versioned manifest contract and validated example are published under schema/. Immutable release manifests are stored in the public Bucket at manifests/v1/revisions/{DATASETCOMMIT}.json; the Bucket's…

Publicly accessible

Dataset · Text generation

ultrawhale-dogfood

Peter Lodri

The SVG files are the source of truth — the PNGs are rasters of them, generated via rsvg-convert. Regenerate any time the SVG changes: Logo: ultrawhale swims in a 5-ring data loop, breathing HF-yellow samples into the dark. Credits to pocoo.vaked.dev for the visual language. Paste this into any coding agent (Claude / opencode / Cursor / aider / Cline) — works from zero context, from anywhere in the dogfeed-loop. ~280 tokens. Self-bootstrapping. End every reply with the loop-state marker. {id, topic, usermessage, freeresponse, freemodel, deepseekresponse, reference, text, role, format, pov, capabilities, memoryref, spacenode, loopindex, pipeline, sessionid, enrichedat, timestamp…

Access requested at publisher apache-2.0 1K<n<10K

Compressed hourly JSON Lines archives of API polling attempts for the Bajs Zagreb Nextbike feed. Each record includes its UTC fetch time, HTTP status, ETag, request duration, body SHA-256, error information, and the full response payload when one was returned. HTTP 304 records intentionally have a null payload. They prove that the collector was running while the source response was unchanged. For the smaller ML-ready Parquet tables, use https://maps.nextbike.net/maps/nextbike-live.json?city=1172&domains=hd&listcities=0&bikes=0

Publicly accessible

Weibo Dataset This repository is part of the StevenZhou0825/weibo-dataset personal Weibo archival and research corpus. It contains immutable, uncompressed TAR media shards associated with sent posts, favorites, repost chains, articles, comments, author avatars, and Weibo expressions. Access and rights Access requests are reviewed manually. Media remains in the representation downloaded from its source; TAR is only a container and does not grant additional rights.

Access requested at publisher

Dataset · Text to speech

voicehub-arena-seed-tts-eval

VoiceHub

Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores. experiments/ / / / contains result.json, records.json, contract.json, verification.json, and audio.tar. records.json contains a rows array. Each successful row specifies its WAV's audioarchive, audiooffset, audiobytes, and audiosha256. Request exactly that byte range from the archive at an immutable…

Publicly accessible

Public mirror of every page on append.page, pushed roughly every 10 minutes when any chain changes. For each page slug you get two files under pages/: - pages/.jsonl — the JCS-canonicalized hash chain (one entry per line, hash-committed to its predecessor). Any later edit, deletion, or reorder is mathematically detectable by anyone who kept a prior snapshot (this dataset is one such copy — HuggingFace also keeps full Git history). - pages/.bodies.jsonl — the actual post text + per-entry salt, in the same row order as the chain. Erased entries appear with body: null and erased: true; salt is still present so anyone with a private archive of the body from before erasure can re-verify it…

Publicly accessible mit

Dataset · Text classification

telegram-news-ua-dataset

Aisberg.public.organization

A continuously updated, de-identified corpus of Ukrainian Telegram news and the discussion around it, published by the Ukrainian non-profit Aisberg (ГО «АЙЗБЕРГ»). It comes in two layers. The first is the raw monthly stream: every post from a fixed set of public news channels, with its reactions and its comment thread. The second is the analysis behind every report Aisberg publishes: posts from different channels grouped into one event, the manipulation techniques found in the coverage, who published first and by how many minutes, view and reaction curves over the first hour, comment sentiment, and a fact-check with a ten-label verdict and the full evidence chain behind it. Free to use…

Publicly accessible cc-by-4.0 100K<n<1M

Live predictions for Polymarket markets, produced by THE ORACLE — an autonomous agent funded by \$ORACLE pump.fun creator fees. Each row is a baseline-model forecast over live orderbook signals (momentum, microstructure, liquidity). - predictions.json / predictions.csv — 100 markets, refreshed each agent cycle. model, backtestacc, auc, modelability, volume, liquidity, enddate, capturedat. Not financial advice.

Publicly accessible mit

OSM polygons tagged with a wikidata= reference, enriched with Wikipedia and Wikivoyage text across all available languages. The published tables are: - polygons/.parquet — one row per polygon - wikipedia/documents/.parquet — one row per unique Wikipedia article - polygonarticles/.parquet — unified polygon-to-document many-to-many links for Wikipedia and Wikivoyage; project identifies the source and documentid references its document table - wikipedia/sections/.parquet — section-level partitions of Wikipedia document text - wikivoyage/documents/.parquet — full Wikivoyage documents associated with places through Wikidata - wikivoyage/sections/.parquet — section-level partitions of Wikivoyage…

Publicly accessible odbl

15-minute snapshots of every docking station in London's cycle hire scheme, collected from the TfL BikePoint API since 2022-04-29. Collection code lives on GitHub: fferegrino/london-cycles-db. A new snapshot is appended every 15 minutes; each finished day is compacted into a single file overnight. During the current day you will also see HHMMSS.csv files in that day's directory, one per 15-minute run. A nightly job merges them into part.csv. Rows carry no coordinates. lat/lon are station attributes rather than measurements, and repeating them on every row was about 40% of the compressed dataset. They live in stations.csv. One row per station per version of its attributes, valid over…

Publicly accessible cc-by-sa-3.0

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.