SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019. An example of 'train' looks as follows. The data fields are the same among all splits. - ind: a int32 feature. - activitylabel: a string feature. - ctxa: a string feature. - ctxb: a string feature. - ctx: a string feature. - endings: a list of string features. - sourceid: a string feature. - split: a string feature. - splittype: a string feature. - label: a string feature. MIT https://github.com/rowanz/hellaswag/blob/master/LICENSE Thanks to @albertvillanova, @mariamabarham, @thomwolf, @patrickvonplaten, @lewtun for adding this dataset.
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. The following table presents a comparative analysis of scores across various domains and tasks. The scores highlight the performance difference between a random agent and the episodes recorded in our dataset. - text: a string feature - images: a image feature - imageobservations: a Sequence(image) feature - textobservations: a Sequence(string) feature - discreteobservations: a Sequence(Sequence(int64)) feature…
Further cleaning done. Please look through the dataset and ensure that I didn't miss anything. Update: Confirmed working method for training the model: https://huggingface.co/AlekseyKorshuk/vicuna-7b/discussions/4#64346c08ef6d5abefe42c12c - Removes instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPTVicunaunfiltered/blob/main/ShareGPTV3unfilteredcleanedsplitnoimsorry.json - Has instances of "I'm sorry, but": https://huggingface.co/datasets/anon8231489123/ShareGPTVicunaunfiltered/blob/main/ShareGPTV3unfilteredcleanedsplit.json The choice is yours. The first dataset may go to far and remove valuable data. The second is better for when the AI asks for…
This catalog is developed for use with the Siril 1.4 series as a public reference database. Hugging Face is one of several mirrors used to distribute the data. This database is provided for both offline download and also for online access. This dataset is provided for scientific and reproducibility purposes. This is an extract of the Gaia DR3 catalog optimized for spectrophotometric color calibration. The catalog is indexed at HEALpix level 8 and selects up to the 127 brightest sources in each level 8 HEALpixel, though in many HEALpixels the number is limited by the Gaia sources that have xpsampled records available. However this strategy ensures an even density of stars, providing…
FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We then used this classifier to retain only the most educational web pages. FineWeb-Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data. The Dataset Curation section details the process for creating the dataset. You can find a deduplicated version of FineWeb-edu in SmolLM-Corpus. We find that the deduplication of this dataset doesn't have…
The FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the datatrove library, our large scale data processing library. FineWeb was originally meant to be a fully open replication of RefinedWeb, with a release of the full dataset under the ODC-By 1.0 license. However, by carefully adding additional filtering steps, we managed to push the performance of FineWeb well above that of the original RefinedWeb, and models trained on our dataset also outperform models trained on other commonly used high quality web datasets (like C4…
To download the full dataset, you can use the following code. If you encounter any issues, please refer to the official Hugging Face documentation. You can use kits in kits/examples.sh to generate glb, usd files, as well as render video with the generated camera trajectory and load into IsaacSim. This dataset is purely agentic-driven generated from SAGE without any manual filtering. The quality of every scene might be varied. This dataset is released under the Apache License 2.0. You are free to use, modify, and distribute this dataset for both commercial and non-commercial purposes, provided that proper attribution is given.
WinoGrande is a new collection of 44k problems, inspired by Winograd Schema Challenge (Levesque, Davis, and Morgenstern 2011), but adjusted to improve the scale and robustness against the dataset-specific bias. Formulated as a fill-in-a-blank task with binary options, the goal is to choose the right option for a given sentence which requires commonsense reasoning. An example of 'train' looks as follows. An example of 'validation' looks as follows. An example of 'validation' looks as follows. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - sentence: a string feature. - option1: a string feature. - option2…
What is Symato CC? To download all WARC data from Common Crawl then filter out Vietnamese in Markdown and Plaintext format. There is 1% of Vietnamse in CC, extract all of them out should be a lot (~10TB of plaintext). Main contributors https://huggingface.co/nampdn-ai https://huggingface.co/binhvq https://huggingface.co/th1nhng0 https://huggingface.co/iambestfeed Simple quality filters To make use of raw data from common crawl, you need to do filtering
Pretraining dataset aligned with GIFT-Eval that has 71 univariate and 17 multivariate datasets, spanning seven domains and 13 frequencies, totaling 4.5 million time series and 230 billion data points. Notably this collection of data has no leakage issue with the train/test split and can be used to pretrain foundation models that can be fairly evaluated on GIFT-Eval. This release is for research purposes only in support of an academic paper. Our models, datasets, and code are not specifically designed or evaluated for all downstream purposes. We strongly recommend users evaluate and address potential concerns related to accuracy, safety, and fairness before deploying this model. We encourage…
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided. An example of 'train' looks as follows. The data fields are the same among all splits. - question: a string feature. - distractor3: a string feature. - distractor1: a string feature. - distractor2: a string feature. - correctanswer: a string feature. - support: a string feature. The dataset is licensed under the Creative Commons Attribution-NonCommercial 3.0 Unported License. Thanks…
RT-Pose introduces a human pose estimation (HPE) dataset and benchmark by integrating a unique combination of calibrated radar ADC data, 4D radar tensors, stereo RGB images, and LiDAR point clouds. This integration marks a significant advancement in studying human pose analysis through multi-modality datasets. The data collection hardware system comprises two RGB cameras, a non-repetitive horizontal scanning LiDAR, and a cascade imaging radar module. We collect the dataset in 40 scenes with indoor and outdoor environments. The dataset comprises 72,000 frames distributed across 240 sequences. The structured organization ensures a realistic distribution of human motions, which is crucial for…
Syn4D: A Multiview Synthetic 4D Dataset Syn4D is a synthetic 4D dataset with multi-view RGB videos, depth, masks, tracking geometry, and supporting object mesh metadata. Layout data/ syn4dv1stride1/ # Syn4D V1, every frame can be a tracking reference frame syn4dv1stride1attachments/ # Additional per-scene attachments for Syn4D V1 stride-1 syn4dv1stride5/ # Syn4D V1, every 5th frame can be a tracking
This dataset contains the prompts used in the Instruction-Following Eval (IFEval) benchmark for large language models. It contains around 500 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times" which can be verified by heuristics. To load the dataset, run: The IFEval dataset is designed for evaluating chat or instruction fine-tuned language models and is one of the core benchmarks used in the Open LLM Leaderboard. The data in IFEval are in English (BCP-47 en). An example of the train split looks as follows: The data fields are as follows: key: A unique ID for the prompt. prompt: Describes the task the model should perform.…
This dataset was created using LeRobot.
This dataset was created using LeRobot.
Synthetic data generated by DataTrove: The finalized run produced 1,354,044,711 (≈1.35B) samples and generated 486,367,076,933 (≈486.4B) completion tokens. Final counts were computed from generated parquet outputs using examples/inference/countcompletiontokens.py and the runs in projects/datatrove/finephrasetokencounts//slurm/stats.json. Each sample includes standard fields such as: - text (source input text from FineWeb-Edu, not the generated output) - rolloutresults (list of generation result objects; one per rollout) - finishreason - text (generated transformed output; for single-rollout runs this is in rolloutresults[0].text) - usage - completiontokens - prompttokens…
PrimeBot Household Bimanual Manipulation Challenge Dataset 中文 | English 中文 目录 关于我们 更新日志 真机遥操作数据 训练集说明 验证集说明 数据集字段说明 URDF 图像 语言指令 本体感知与动作 机器人推理接口 UMI数据 数据概览 目录结构 数据集字段说明 图像 本体感知与动作 索引字段 标注与 IMU 关于我们 我们来自上纬新材-启元研究院,我们的使命是加速个人机器人时代到来,加速家用机器人时代到来。我们开源高质量面向家庭操作的双臂操作数据集,同时开放机器人硬件描述以供可视化、可复现研究。 如果本数据集对您的工作有帮助,感谢引用: @misc{xu2026scalingbimanualhouseholdmanipulation, title={Scaling Bimanual Household Manipulation from 1,500
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.



