SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

A set of prerender 2D images from GSO datasets, including png, mask, camera calibration and poses. When training Deep Learning models for Novel View Synthesis (NVS) or 3D-to-2D Representation Learning, loading 3D meshes and rendering views on-the-fly inside PyTorch DataLoaders creates massive bottlenecks. This script solves critical problems

Publicly accessible mit

Dataset · Text to video

PAWBench-Results

YuandongPu

This dataset package defines the clean HF repository layout and contains the public metadata needed to run the PAWBench Full31 coverage and Table1 PAW-Cal PAWEval rerun inputs. In the local no-copy handoff, videos are not duplicated under this directory; they are uploaded into sharded videos/bymediaid/ subdirectories from the private upload plan after explicit human approval. - videos/bymediaid/ /.mp4: one physical video per mediaid after human upload from the private upload plan. - manifests/sharedmediaregistry.jsonl: public media identity registry. - manifests/full31pawevalmanifest.jsonl: Full31 logical PAWEval rows. - manifests/table1pawevalrerunrows.jsonl: Table1 PAW-Cal rerun rows.…

Access requested at publisher cc-by-4.0

Dataset · Image to image

Noob2EDIT

AnimeNexa LaxCore

HF 仓库:AnimeNexaLaxCore/Noob2EDIT,公开、人工审核开放。仓库内 data/ 下是分片、旁路表和 manifest.json,demo/ 是加载器适配用的演示包。 来源:ModelScope niangao233/NOOB2EDITGAMECG15M(hitomi gamecg 画廊,按相似度聚类成组)。 本目录是清洗、配对、VLM 标注后的训练用导出,格式为 WebDataset 分片,一个样本 = 一个差分组。 demo/ 下是给数据加载器适配用的小规模演示包,格式与正式分片完全一致。 - family:cg 游戏 CG 差分;sprite 立绘(原图带透明通道,已合成白底)。两套分片分开,训练时按比例混合。 - split:train / val,按画廊 id 的 sha1 哈希切分(val 占 2%),同一游戏的组不会跨集。 - 分片目标大小 1 GiB(正式),demo/ 为 300 MB。分片内的组顺序按画廊哈希打乱,各年份、各游戏天然混合。 - 分片文件名与 index 单调递增,manifest 中记录每个分片的字节数、sha256、样本数、边数和旁路表 sha256。 一个组在 tar 中是连续的若干成员,key 为 {galleryid}g{组号}: - CG 分支保留原始 AVIF 字节,未重编码。 - 立绘和任何带真实透明像素的图,在写入时合成到纯白背景并重编码为 AVIF q90(images[].reencoded = true)。训练数据中不存在 alpha 通道。 - 组内所有图尺寸一致(width /…

Access requested at publisher other 1M<n<10M

수집 시각 · 2026-09-18 23:26:58 0개 (GPU 0장) · 대기 3개 클러스터 유휴 GPU 0장 종료 코드가 124 라 sacct 가 FAILED 로 적지만 정상 동작입니다. 진짜 실패는 경과 시간이 한도보다 훨씬 짧거나 iter 체크포인트가 늘지 않은 경우입니다. 계정마다 접속하지 않고 한 계정에서 전부 조회합니다. 대기 사유·예상 시작 시각은 sacct 로는 얻을 수 없습니다.

Publicly accessible

Dataset

Longitudinal-CT

L

Mirror of Longitudinal-CT v2 (DOI: 10.57754/FDAT.75kj1-64747; Each patient zip usually contains baseline/follow-up CT NIfTI and lesion masks. Gatidis et al., A longitudinal whole-body CT dataset with manually annotated tumor lesions.

Publicly accessible other 10M<n<100M

Dataset

rx-eu

Numix

Archives de produits de télédétection, mises à jour en continu. Source des données: EUMETNET OPERA, via le service Open Radar Data du projet RODEO (https://rodeo-project.eu/) — CC-BY 4.0. Attribution: « Données EUMETNET OPERA ».

Access requested at publisher cc-by-4.0

Dataset · Video classification

jais

Ng Thien An

jais Synthetic multi-object rigid-body videos generated with Kubric (PyBullet physics + Blender/Cycles rendering, GPU-rendered). Every scene has 4-6 fixed-mass bodies on a table; exactly one body (the subject) receives an initial velocity along a single heading and nothing else is actuated. Some bodies sit in the subject's corridor and get struck, others are bystanders that are never touched. Spheres roll, boxes slide, and a sphere launched without spin slides first and then

Access requested at publisher apache-2.0 1K<n<10K

An independent copy of United States federal public datasets that have been removed, discontinued, degraded, or formally proposed for elimination. Everything here was already public and is United States Government work, which is not subject to copyright under Nothing in this mirror is modified. No values are corrected, no records are dropped, no columns are renamed, no formats are converted. Each file is stored exactly as the agency published it, alongside the date it was retrieved and a SHA-256 checksum, so that any copy can be verified against any other copy. This is not affiliated with, endorsed by, or connected to NOAA, the CDC, the EPA, or any other agency or organisation. Two…

Publicly accessible other 1B<n<10B

Dataset

koskari

Koskari

Verified Koskari language examples exported from the Koskari conformance corpus. Each row traces to a corpus record with machine-checked expectations at one or more language phases (read, expand, eval, error, invalid). Every corpus record. One row per record regardless of which verify stages are present. Missing stages are null. User documentation pages (Markdown) from the source repository's docs/reference/, docs/start/, docs/guide/, docs/patterns/, and docs/agents/ trees. One row per page. The path column matches the strings in the corpus config's references column verbatim, so a record's references resolve directly to their documentation content. The agents section is written for AI…

Publicly accessible gpl-3.0

Dataset · Text generation

style-dpo

Hub

taskcategories: - text-generation sizecategories

Publicly accessible apache-2.0 100M<n<1B

paper-with-me 서비스용 SQLite 스냅샷. Papers with Code 2025-07 아카이브(pwc-archive, CC-BY-SA 4.0) + arXiv/HF Daily Papers/GitHub 일일 수집분.

Publicly accessible cc-by-sa-4.0

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.