SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019. An example of 'train' looks as follows. The data fields are the same among all splits. - ind: a int32 feature. - activitylabel: a string feature. - ctxa: a string feature. - ctxb: a string feature. - ctx: a string feature. - endings: a list of string features. - sourceid: a string feature. - split: a string feature. - splittype: a string feature. - label: a string feature. MIT https://github.com/rowanz/hellaswag/blob/master/LICENSE Thanks to @albertvillanova, @mariamabarham, @thomwolf, @patrickvonplaten, @lewtun for adding this dataset.

Publicly accessible

Dataset · Reinforcement learning

jat-dataset

Jack of All Trades project

The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. The following table presents a comparative analysis of scores across various domains and tasks. The scores highlight the performance difference between a random agent and the episodes recorded in our dataset. - text: a string feature - images: a image feature - imageobservations: a Sequence(image) feature - textobservations: a Sequence(string) feature - discreteobservations: a Sequence(Sequence(int64)) feature…

Publicly accessible apache-2.0

Further cleaning done. Please look through the dataset and ensure that I didn't miss anything. Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c - Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPTVicunaunfiltered/blob/main/ShareGPTV3unfilteredcleanedsplitnoimsorry.json - Has instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPTVicunaunfiltered/blob/main/ShareGPTV3unfilteredcleanedsplit.json The choice is yours. The first dataset may go to far and remove valuable data. The second is better for when the AI asks for…

Publicly accessible apache-2.0

This catalog is developed for use with the Siril 1.4 series as a public reference database. Hugging Face is one of several mirrors used to distribute the data. This database is provided for both offline download and also for online access. This dataset is provided for scientific and reproducibility purposes. This is an extract of the Gaia DR3 catalog optimized for spectrophotometric color calibration. The catalog is indexed at HEALpix level 8 and selects up to the 127 brightest sources in each level 8 HEALpixel, though in many HEALpixels the number is limited by the Gaia sources that have xpsampled records available. However this strategy ensures an even density of stars, providing…

Publicly accessible gpl-3.0

Dataset · Text generation

fineweb-edu

FineData

FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We then used this classifier to retain only the most educational web pages. FineWeb-Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data. The Dataset Curation section details the process for creating the dataset. You can find a deduplicated version of FineWeb-edu in SmolLM-Corpus. We find that the deduplication of this dataset doesn't have…

Publicly accessible odc-by n>1T

Dataset · Text generation

fineweb

FineData

The FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the datatrove library, our large scale data processing library. FineWeb was originally meant to be a fully open replication of RefinedWeb, with a release of the full dataset under the ODC-By 1.0 license. However, by carefully adding additional filtering steps, we managed to push the performance of FineWeb well above that of the original RefinedWeb, and models trained on our dataset also outperform models trained on other commonly used high quality web datasets (like C4…

Publicly accessible odc-by n>1T

Dataset · Text to 3d

SAGE-10k

NVIDIA

To download the full dataset, you can use the following code. If you encounter any issues, please refer to the official Hugging Face documentation. You can use kits in kits/examples.sh to generate glb, usd files, as well as render video with the generated camera trajectory and load into IsaacSim. This dataset is purely agentic-driven generated from SAGE without any manual filtering. The quality of every scene might be varied. This dataset is released under the Apache License 2.0. You are free to use, modify, and distribute this dataset for both commercial and non-commercial purposes, provided that proper attribution is given.

Publicly accessible apache-2.0 10K<n<100K

Dataset

winogrande

Ai2

WinoGrande is a new collection of 44k problems, inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern 2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires commonsense reasoning. An example of 'train' looks as follows. An example of 'validation' looks as follows. An example of 'validation' looks as follows. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - sentence: a string feature. - option1: a string feature. - option2…

Publicly accessible

Dataset

cc

Symato Team

What is Symato CC? To download all WARC data from Common Crawl then filter out Vietnamese in Markdown and Plaintext format. There is 1% of Vietnamse in CC, extract all of them out should be a lot (~10TB of plaintext). Main contributors https://huggingface.co/nampdn-ai https://huggingface.co/binhvq https://huggingface.co/th1nhng0 https://huggingface.co/iambestfeed Simple quality filters To make use of raw data from common crawl, you need to do filtering

Publicly accessible mit 1K<n<10K

Dataset · Time series forecasting

GiftEvalPretrain

Salesforce AI Research

Pretraining dataset aligned with GIFT-Eval that has 71 univariate and 17 multivariate datasets, spanning seven domains and 13 frequencies, totaling 4.5 million time series and 230 billion data points. Notably this collection of data has no leakage issue with the train/test split and can be used to pretrain foundation models that can be fairly evaluated on GIFT-Eval. This release is for research purposes only in support of an academic paper. Our models, datasets, and code are not specifically designed or evaluated for all downstream purposes. We strongly recommend users evaluate and address potential concerns related to accuracy, safety, and fairness before deploying this model. We encourage…

Publicly accessible apache-2.0 1M<n<10M

Dataset · Question answering

sciq

Ai2

The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided. An example of 'train' looks as follows. The data fields are the same among all splits. - question: a string feature. - distractor3: a string feature. - distractor1: a string feature. - distractor2: a string feature. - correctanswer: a string feature. - support: a string feature. The dataset is licensed under the Creative Commons Attribution-NonCommercial 3.0 Unported License. Thanks…

Publicly accessible cc-by-nc-3.0 10K<n<100K

RT-Pose introduces a human pose estimation (HPE) dataset and benchmark by integrating a unique combination of calibrated radar ADC data, 4D radar tensors, stereo RGB images, and LiDAR point clouds. This integration marks a significant advancement in studying human pose analysis through multi-modality datasets. The data collection hardware system comprises two RGB cameras, a non-repetitive horizontal scanning LiDAR, and a cascade imaging radar module. We collect the dataset in 40 scenes with indoor and outdoor environments. The dataset comprises 72,000 frames distributed across 240 sequences. The structured organization ensures a realistic distribution of human motions, which is crucial for…

Publicly accessible cc-by-nc-sa-4.0 1K<n<10K

Dataset

Syn4D

Zeren Jiang

Syn4D: A Multiview Synthetic 4D Dataset Syn4D is a synthetic 4D dataset with multi-view RGB videos, depth, masks, tracking geometry, and supporting object mesh metadata. Layout data/ syn4dv1stride1/ # Syn4D V1, every frame can be a tracking reference frame syn4dv1stride1attachments/ # Additional per-scene attachments for Syn4D V1 stride-1 syn4dv1stride5/ # Syn4D V1, every 5th frame can be a tracking

Publicly accessible cc-by-4.0

Dataset · Text generation

IFEval

Google

This dataset contains the prompts used in the Instruction-Following Eval (IFEval) benchmark for large language models. It contains around 500 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times" which can be verified by heuristics. To load the dataset, run: The IFEval dataset is designed for evaluating chat or instruction fine-tuned language models and is one of the core benchmarks used in the Open LLM Leaderboard. The data in IFEval are in English (BCP-47 en). An example of the train split looks as follows: The data fields are as follows: key: A unique ID for the prompt. prompt: Describes the task the model should perform.…

Publicly accessible apache-2.0

Dataset · Text generation

finephrase

FineData

Synthetic data generated by DataTrove: The finalized run produced 1,354,044,711 (≈1.35B) samples and generated 486,367,076,933 (≈486.4B) completion tokens. Final counts were computed from generated parquet outputs using examples/inference/countcompletiontokens.py and the runs in projects/datatrove/finephrasetokencounts//slurm/stats.json. Each sample includes standard fields such as: - text (source input text from FineWeb-Edu, not the generated output) - rolloutresults (list of generation result objects; one per rollout) - finishreason - text (generated transformed output; for single-rollout runs this is in rolloutresults[0].text) - usage - completiontokens - prompttokens…

Publicly accessible odc-by n>1M

PrimeBot Household Bimanual Manipulation Challenge Dataset 中文 | English 中文 目录 关于我们 更新日志 真机遥操作数据 训练集说明 验证集说明 数据集字段说明 URDF 图像 语言指令 本体感知与动作 机器人推理接口 UMI数据 数据概览 目录结构 数据集字段说明 图像 本体感知与动作 索引字段 标注与 IMU 关于我们 我们来自上纬新材-启元研究院,我们的使命是加速个人机器人时代到来,加速家用机器人时代到来。我们开源高质量面向家庭操作的双臂操作数据集,同时开放机器人硬件描述以供可视化、可复现研究。 如果本数据集对您的工作有帮助,感谢引用: @misc{xu2026scalingbimanualhouseholdmanipulation, title={Scaling Bimanual Household Manipulation from 1,500

Publicly accessible cc-by-sa-4.0

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.