SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

An expert-level, citation-backed knowledge base on reinforcement learning for large language models — RLHF, DPO and offline preference optimization, reward modeling, RLVR and reasoning, training systems, and the failure modes — built collaboratively by autonomous agents. Each topic article is a deep dive written so you can learn the topic from it without reading the underlying papers, with every non-obvious claim cited to a source. Every change lands through a reviewed pull request, so this is curated knowledge, not an accumulation. Articles cite sources inline as [source: ] (e.g. [source:arxiv:2203.02155]); each resolves to that source's summary in sources/, which links on to the full…

Publicly accessible cc-by-4.0

Dataset · Question answering

natural_questions

Google Research Datasets

The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. An example of 'train' looks as follows. This is a toy example. The data fields are the same among all splits. - id: a string feature. - document a dictionary feature containing: - title: a string feature. - url: a string feature. - html: a string feature. - tokens: a dictionary feature containing: - token: a string…

Publicly accessible cc-by-sa-3.0 100K<n<1M

For the official data release page, please see microsoft/SWE-bench-Live. SWE-bench-Live is a live benchmark for issue resolving, designed to evaluate an AI system’s ability to complete real-world software engineering tasks. Thanks to our automated dataset curation pipeline, we plan to update SWE-bench-Live on a monthly basis to provide the community with up-to-date task instances and support rigorous and contamination-free evaluation. - 9/17/2025: Dataset updated! We’ve finalized the update process for SWE-bench-Live: Each month, we will add 50 newly verified, high-quality issues to the dataset. The lite and verified splits will remain frozen, ensuring fair leaderboard comparisons and…

Publicly accessible mit

This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository. We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights. We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers their dataset library. Specifically, any content you download, access or use from our index, is at your own risk and subject to the terms of service or copyright limitations accompanying such content. The image url-text index, which is a research artifact, is…

Publicly accessible cc-by-4.0

Append-only, content-addressed archive of the live intelligence streams the killinchu demo ingests, published by SZL Holdings. The shared szlhfbucket client writes one NDJSON shard per UTC day under intel/. Each row is {schema, id, ts, source, kind, payload}. The default Dataset Viewer configuration is the homogeneous archivemanifest, one row per immutable raw shard. It exposes path, byte count, Git blob hash, rights state, and training eligibility. The mixed raw NDJSON remains available under intel/.ndjson but is deliberately not numbers and strings in fields such as altbaro breaks strict dataset loaders. This is a usability repair, not a rewrite of append-only history. Every raw manifest…

Publicly accessible other

Frequently collected Polymarket market data stored as Zstandard-compressed Parquet. New batches are collected approximately every five minutes and partitioned by product, date, and collection run. Public upload of collected data began on July 23rd, 2026. - marketsnapshots: best bid/ask, full book JSON, outcome, asset ID, source timing, and collection timing. - orderbookdepth: normalized bid and ask levels with price, size, side, and level index. - topholders: ranked holder balances and public wallet/profile fields by market and token. Parquet files can also be queried directly with DuckDB, Polars, PyArrow, or other compatible tools. This dataset is made available under the license.…

Publicly accessible cc-by-4.0

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.