SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Video text to text

EgoLife

LMMs-Lab

Data cleaning, stay tuned! Please refer to https://egolife-ai.github.io/ first for general info. Checkout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information. Code: https://github.com/egolife-ai/EgoLife

Publicly accessible mit

Dataset · Visual question answering

IntuitivePhysics

WorldBench

WorldBench is a new benchmark designed to evaluate the physical understanding and prediction of modern world models and vision-language models. There are two components: The video based benchmark can be found in /scenes. There are 4 high-level categories for different physics concepts being tested. Within each, there are 3-5 scenes each with 25-50 variations. The text based benchmark is in /textualquestions. There are 4 JSON files, one per category. Code to run the evaluation for this benchmark along with instructions can be found here: https://drive.google.com/file/d/1TNHfV-mKiidl1eFWJyctBOodWJnCajA/view?usp=sharing

Publicly accessible n<1K

The primary source is hosted on GitHub, please open issues and pull requests there, not here. Terminal-Bench is now a continuous benchmark: new versions are released periodically as tags on the source repo instead of one-off snapshots. This dataset mirrors that model on the Hub: instead of a separate terminal-bench-X.Y repo per release, one repo, tagged per version. main always tracks the latest published version; each release is additionally available as an immutable Hub tag matching the source repo's GitHub tag. The official published dataset and leaderboard are hosted on the Harbor Hub. Usage e.g. harbor run -d terminal-bench/[email protected]. This repo is a mirror of…

Publicly accessible apache-2.0

Dataset

PADBench

Sunzhen

Vincent-HKUSTGZ/PADBench This repository contains the complete dataset from the finishtransferrank folder, including Llama2-7B models fine-tuned with PEFTGuard for different datasets. Content Structure This dataset includes the following subdirectories:.cache/: Contains model files for.cache dataset chatglm6btoxicbackdoorshardrank256qv/: Contains model files for chatglm6btoxicbackdoorshardrank256qv dataset flant5xltoxicbackdoorshardrank256qv/

Publicly accessible

Flood Detection Dataset Introduction This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide. Key features: Cloud-penetrating SAR data for consistent

Publicly accessible mit

Dataset · Question answering

PubMedQA

Qiao Jin

The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts. The official leaderboard is available at: https://pubmedqa.github.io/. 500 questions in the pqalabeled are used as the test set. They can be found at https://github.com/pubmedqa/pubmedqa. English Thanks to @tuner007 for adding this dataset.

Publicly accessible mit 100K<n<1M

Foursquare OS Places is now a gated dataset on Hugging Face. Read more about why we are making this change here: https://medium.com/@foursquare/evolving-fsq-os-places-fa7a3f5197cd With Foursquare’s Open Source Places, you can access free data to accelerate geospatial innovation and insights. View the Places OS Data Schemas for a full list of available attributes. In order to access Foursquare's Open Source Places data, it is recommended to use Spark. Here is how to load the Places data in Spark from Hugging Face. - For Spark 3, you can use the readparquet helper function from the HF Spark documentation. It provides an easy API to load a Spark Dataframe from Hugging Face, without having to…

Access requested at publisher apache-2.0

DL3DV Benchmark Download Instructions This repo contains 140 scenes in the DL3DV-benchmark, which are sampled from DL3DV-10K. The repo includes a README, License, colmaps/images (compatible to nerfstudio and 3D gaussian splatting), scene labels and the performances of methods reported in the paper (ZipNeRF, 3DGS, MipNeRF-360, nerfacto, Instant-NGP). The benchmark preview page can be found here https://dl3dv-10k.github.io/DL3DV-Benchmark-Preview/. Download As the whole

Access requested at publisher n>1T

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.