SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

An independent archive of public Kalshi market trade data retrieved from Kalshi market-data API endpoints. This dataset is not affiliated with or endorsed by Kalshi. Each historicalrawN.jsonl file is an immutable numbered shard. Records contain a market ticker, market window, reported volume, and the public trades returned for that market. Trade fields can include timestamps, public trade IDs, prices, quantities, taker direction, and block-trade status. The archive contains no account credentials, private orders, portfolio data, or personal information. Data users are responsible for complying with the Kalshi Developer Agreement and other applicable terms.

Publicly accessible

Dataset · Other

Finetune_neurok

QiuyiDing

This repository stores a standalone, NeuROK-compatible deformation corpus under objaverse/finetunedeformation. It combines processed public animation data with generated MPM solid trajectories. Training consumers should read index/all.jsonl (or its train.jsonl and val.jsonl split files) rather than discovering NPZ files by directory traversal. - 1,998 DynamicObjaverseProcessed geometries; - 500 DyMesh / AnimateAnyMesh geometries; - 2,498 trainable rows in the canonical combined index; The generated target is 5,000 geometries × 2 accepted clips = 10,000 clips. Its deterministic 24-cell cycle balances elastic, plastic, sand, and snow, with every material parameter drawn from one of four…

Publicly accessible

Dataset · Text generation

finetranslations-edu-zhtw

Huang Liang Hsun

finetranslations-edu-zhtw 是以 HuggingFaceFW/finetranslations-edu 為來源,將其 translatedchunks(原始多語言教育類網頁內容、先被 pivot 翻譯成英文的版本)進一步翻譯成繁體中文的資料集。 HuggingFaceFW/finetranslations-edu 收錄了原本以英文以外語言(oglanguage,涵蓋約 200 種語言)撰寫、經篩選具教育價值(eduscore)的網頁內容,並將其 pivot 翻譯成英文(translatedtext / translatedchunks)。本資料集在此基礎上,將已經是英文的 translatedchunks 逐段(chunk)翻譯成繁體中文,並依原順序重組回完整段落(translatedtextzhtw),讓來自約 200 種原始語言、已篩選過教育價值的內容,也能以繁體中文呈現。 翻譯由一套 LLM 驅動(模型本身不對外公開資訊,僅說明為 LLM),部署為多個平行推論服務以提升吞吐量。 繁體中文語言模型的持續預訓練(continued pretraining)語料補充,特別是希望涵蓋多元語言/文化來源、且具教育價值的內容。 id / url / oglanguage / oglanguagescore / eduscore / eduscoreraw / translatedtokencount / minhashclustersize:原樣保留自上游 HuggingFaceFW/finetranslations-edu。…

Access requested at publisher odc-by 10M<n<100M

Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0. - predictionslatest.parquet is rewritten on scheduled prediction runs - verification updates actual columns once the target hour is in the past - frontendsnapshot.json is regenerated after prediction and verification writes - targettimestamp - referencetime - leadtimehours - leadbucket - predictionmadeat - city - verified - dmitemperature2mpred - dmiwindspeed10mpred - dmiwindgusts10mpred - dmiprecipitationprobabilitypred - dmiprecipitationpred - mltemp - mlwindspeed - mlwindgust - mlrainprob - mlrainamount - actualtemp - actualwindspeed - actualwindgust - actualprecipitation - actualrainevent…

Publicly accessible cc-by-4.0

Kiiteitte history Kiiteitte が収集した、今までの選曲履歴。 1時間おきに更新されます。 型 { // 動画ID "videoid": "sm44670499", // タイトル "title": "library->w4nderers / 足立レイ、つくよみちゃん", // 投稿者 "author": "名無し。", // サムネイルのURL "thumbnail": "https://nicovideo.cdn.nimg.jp/thumbnails/44670499/44670499.91820835", // 選曲日時 "date": "2025-02-22 12:51:51", // 新しく増えたお気に入り数。不明の場合は null "newfaves": 5, // 回ったユーザーの数。不明の場合は null "spins": 13, // イチ押しリストのユーザーのURL。イチ押しリスト以外から選曲された場合は null

Publicly accessible 10K<n<100K

This dataset hosts the DuckDB database used by the It is derived from the raw benchmark result artifacts in and is packaged for leaderboard, viewer, notebook, and SQL use. The database is produced by the HAKARI-Bench implementation in Because the benchmark code, schema, and build workflow evolve over time, this dataset card intentionally points to the canonical documentation instead of duplicating implementation details here. - duckdb/hakaribench.duckdb: current leaderboard DuckDB database.

Publicly accessible

This dataset contains the complete text of several novel, fully generated with the assistance of open-source large language models (LLMs). It represents an experiment in large-scale literary creation under a human-AI collaborative paradigm: the human author designed the story architecture, character settings, historical research, and narrative pacing, while open-source LLMs carried out text expansion, dialogue generation, scene depiction, and the weaving of multiple narrative threads based on detailed prompts. In addition to individual files, this dataset is also packaged in JSON format ( )for large‑scale training, strictly following the instruction‑tuning data formats of mainstream LLMs…

Publicly accessible mit

Dataset · Robotics

pico-robotics-advanced

Jiu

Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…

Access requested at publisher cc-by-4.0 10K<n<100K

Dataset · Object detection

cati-singapore-dataset

Suhas Reddy

Real-time vehicle detection data collected from Singapore's 90 LTA traffic cameras using CATI (Context-Aware Traffic Intelligence) — a novel FiLM-conditioned YOLOv11 detector that adapts to environmental conditions in real time. This dataset contains per-camera vehicle detection results collected continuously from Singapore's Land Transport Authority (LTA) expressway camera network. Each record captures a full detection sweep of a single camera including vehicle counts, class breakdown, directional split, and environmental context. Detections are produced by CATI — a novel architecture that injects FiLM (Feature-wise Linear Modulation) layers into YOLOv11s, conditioning the backbone on…

Publicly accessible mit 10K<n<100K

Dataset · Robotics

pico-robotics-annotated

Jiu

Builds on the Advanced edition by adding coarse action segmentation. Every sequence is divided into labelled temporal segments, so the data can be used directly for action recognition, temporal segmentation, and behaviour-understanding tasks without an annotation pass of your own. This is a gated dataset. Access requests are reviewed manually; submit one from the dataset page. Everything except segments.json is documented in the Segments within a sequence are contiguous and non-overlapping; frames that fit no class are labelled other. TBD — list the action classes here, with a one-line definition and the frame count for each, Labels are deliberately coarse: boundaries are approximate and…

Access requested at publisher cc-by-4.0 10K<n<100K

Incrementally generated unfiltered candidates. This is not a final selected dataset. Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. runconfig.json pins models, source pools, parameters and implementation hashes. For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs. Global completion requires all worker reports to be complete. Assignments are backgroundid modulo numworkers. Per-worker runconfig.json pins the parameters. Workers never share a background ID. Each data/group-NNN/noise-NNNNNN.tar contains one normalized background.wav, 15 rank-NN.mask.wav / rank-NN.mixture.wav pairs…

Publicly accessible

Dataset · Visual question answering

TrainingData_Stage3

AnchorSR

This repository is a provenance-preserving collection of spatial measurement questions for answer-supervised training. It is built from VSI-590K, SpaceVista-Full, HiSpatial-Data, and CA-VQA. SenseNova-SI-8M is intentionally out of scope for this release. The first deliverable is the complete master collection. Smaller and larger training views will be derived only after all eligible metric examples have been retained and audited; they are not early sampling quotas. - annotations/ /measurement.parquet: canonical question/answer rows. - manifests/ /mediainventory.jsonl: unique required media and usage. - audits/ /: selection counts, validation, checksums, and source policy. - media/ /.tar…

Publicly accessible other

Dataset

HESRT

Jiahao Ji

HESRT Human-reviewed, accession-level spatial-omics artifacts with tissue images, expression matrices and source provenance. Samples (GSMs) Studies (GSEs) Expression observations Compressed packages 4,468 507 37,054,005 413.94 GB Each sample is a self-contained, checksummed tar.zst package. Browse the catalog, select the accessions you need, then download those samples. The full collection contains approximately 647.72 GB of uncompressed member data; do not clone

Publicly accessible other 1K<n<10K

Private working dataset. Conference/journal paper corpus across OpenReview venues (ICLR, NeurIPS incl. D&B/position tracks, ICML incl. position, COLM, TMLR, AISTATS, UAI, ALT, MathAI, and later additions) plus the ACL Anthology family. Coverage, per-venue availability, decisions semantics, and known biases are documented authoritatively in the GitHub repo's data/README.md — read that first; per-venue counts change as the corpus grows, so they are deliberately not duplicated here. - raw/.jsonl — canonical OpenReview snapshots, one line per - papers/ /papers.jsonl — distilled metadata (schema: data/README.md) - reviews/ /reviews.jsonl — one line per forum reply, full text - extractedtext/…

Publicly accessible other

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.