SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Original dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, year={2020}

Publicly accessible cc-by-4.0

Retargeted AMASS for Robotics Project Overview This project aims to retarget motion data from the AMASS dataset to various robot models and open-source the retargeted data to facilitate research and applications in robotics and human-robot interaction. AMASS (Archive of Motion Capture as Surface Shapes) is a high-quality human motion capture dataset, and the SMPL-X model is a powerful tool for generating realistic human motion data. By adapting the motion data from AMASS

Publicly accessible cc-by-4.0 10K<n<100K

Dataset · Summarization

Scientific-Summaries

LAION eV

22 million LLM-generated structured summaries of scientific papers, enriched with OpenAlex scholarly metadata. Each paper has an 18-field structured summary covering methodology, key results, claims, limitations, and more. This public dataset includes full paper text for ~5.3 million papers where open-access status has been confirmed -- either through OpenAlex metadata or because the paper originates from a permissively licensed source such as the arXiv preprint server, bioRxiv, medRxiv, or ChemRxiv. Full text (textsanitized, textraw) is included in this public dataset when either of the following is true: 1. The paper originates from a permissively licensed source: arXiv preprint server…

Publicly accessible cc-by-4.0 10M<n<100M

Dataset

scitail

Ai2

The SciTail dataset is an entailment dataset created from multiple-choice science exams and web sentences. Each question and the correct answer choice are converted into an assertive statement to form the hypothesis. We use information retrieval to obtain relevant text from a large text corpus of web sentences, and use these sentences as a premise P. We crowdsource the annotation of such premise-hypothesis pair as supports (entails) or not (neutral), in order to create the SciTail dataset. The dataset contains 27,026 examples with 10,101 examples with entails label and 16,925 examples with neutral label An example of 'train' looks as follows. An example of 'validation' looks as follows. An…

Publicly accessible

Dataset · Text generation

fineweb-2

FineData

FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular FineWeb dataset, bringing high quality pretraining data to over 1000 languages. The FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions

Publicly accessible odc-by n>1T

Dataset · Question answering

squad_v2

Pranav R

Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable. SQuAD 2.0 combines the 100,000 questions in SQuAD1.1 with over 50,000 unanswerable questions written adversarially by crowdworkers to look similar to answerable ones. To do well on SQuAD2.0, systems must not only answer questions when possible, but also determine when no answer is supported by the paragraph and abstain from answering. Question Answering. English (en). An example of…

Publicly accessible cc-by-sa-4.0 100K<n<1M

This dataset is comprised of forecasts from the German Weather Service's (DWD) ICON-Global model from March 2023 to the present with all variables included. Each forecast runs up to 4 days into the future, and the model is ran 4 times per day. This data is an archive of the publicly available data at https://opendata.dwd.de/weather/nwp/, converted to Zarr format with Xarray. No other processing of the data is performed. Note: The raw files are deleted after 24 hours, and there is no long-term archive available publicly. This data is intended for use in renewable energy forecasting, weather forecasting, and anything that can use high-quality weather forecasts over Europe. The dataset is…

Access requested at publisher cc-by-4.0 1K<n<10K

Dataset · Image feature extraction

FOMO260K

UCPH DIKU Foundation Models for Brain MRI

FOMO260K: Brain MRI Dataset for Large-Scale Self-Supervised Learning with Clinical Data Dataset paper preprint: A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning. https://arxiv.org/pdf/2506.14432. Description FOMO260K is a large-scale dataset of brain MRI scans, including both clinical and research-grade scans. The dataset includes a wide range of sequences, including T1, MPRAGE, T2, T2, FLAIR, SWI, T1c, PD, DWI

Publicly accessible cc-by-nc-sa-4.0 100K<n<1M

Dataset · Question answering

TOFU

Locus Lab

The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set. - Website: The landing page for TOFU - arXiv Paper: Detailed information about the TOFU dataset and its significance in unlearning tasks. - GitHub Repository: Access the source code, fine-tuning scripts, and additional resources for the TOFU dataset. - Dataset on Hugging Face: Direct link to download the…

Publicly accessible mit 1K<n<10K

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.