SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Image segmentation

GTCrack

Guoxi Liu

Automated detection of expressway pavement cracks is critically important for infrastructure inspection and maintenance. In this task, range information provides a key geometric cue for distinguishing genuine cracks from pavement textures, shadows, stains, and other visual interference. However, industrial-grade pavement data acquisition equipment is expensive, and pixel-wise annotation is labor intensive. Consequently, most existing pavement crack datasets contain only visual imagery and cannot precisely represent the geometric morphology of the pavement surface. GTCrack was developed to alleviate this data bottleneck. A survey vehicle equipped with two laser line profilers was used to…

Publicly accessible

This repo holds every shape of the released dataset described under "The data", packed as tar shards per representation. Each split's uuids are sorted and cut into shards of 1,000 shapes. A tar holds one representation's files for one shard (atlas: atlas.npz; voxels: the two.vxz files; multiview: multiview/; thumbnail: thumbnail.png) at paths /, as in the tree under "The data". Shard k holds the same shapes in every representation. The atlas.npz files here are stored compressed; the arrays are the same and np.load reads them the same way. Emissive 3D shapes from TexVerse with their material and emission maps in three representations: a UV atlas, sparse voxels, and six rendered views. One…

Publicly accessible other 10K<n<100K

Dataset · Question answering

MS-MARCO-0G-120M

0G.AI

A long-context retrieval and question-answering benchmark with 1,143,371 documents and 10,000 questions, including 50 questions about 0G. It extends the ms100M bank from MSA-RAG-BENCHMARKS with 20 million additional text tokens and 655 new questions grounded in the added documents. 120M is the benchmark's nominal size. The actual corpus contains 125,708,601 raw tokens under the MSA-4B tokenizer: 105,708,601 original tokens plus exactly 20,000,000 added tokens. Counts exclude prompts, wrappers, and special tokens. This is a derived benchmark, not an official Microsoft MS MARCO release. The original documents occupy positions 0:961686; the original questions occupy positions 0:9345. Their…

Publicly accessible other 1M<n<10M

Dataset · Image and text to text

MemGUI-3K-Verl

Anonymous

RL prompt corpus of the ICLR 2027 submission MemGUI-RL (project page: https://memgui-rl-anonymous.github.io/). Every record is one annotated state of a MemGUI-3K trajectory in the ConAct conversation format consumed by the FARPO trainer (https://github.com/memgui-rl-anonymous/MemGUI-RL): Each record has taskid, stepnumber, ispositive, badstep, rawresponse (the annotated ConAct response:,,,, ), conversations (system prompt, user prompt with the folded history / UI memory / recent step record and the screenshot, reference assistant turn) and metadata (impact / reasonableness labels of the step, trajectory ids). Screenshots are not duplicated here: image references have the form…

Publicly accessible apache-2.0

Dataset · Robotics

MesaTask-CTRC-100-shuffled

Yifanwin

This dataset holds the starting state of a scene restoration task: compared with the target state, 1-3 objects in each scene have been moved to random positions and need to be put back. source.scene on each object identifies its scene. - position is the model origin (not the bounding-box center), in centimeters - size is the object's target extent; the model is already scaled to match it. (x, y) are the horizontal extents and z is the height. A model's own local up-axis may not align with z, so orient it using size as the reference - rotation is the complete absolute orientation as a quaternion in [x, y, z, w] order. Apply it directly, per object, and do not compose any additional rotation…

Publicly accessible cc-by-nc-4.0

Dataset · Robotics

MesaTask-CTRC-100-target

Yifanwin

This dataset holds the target state of a scene restoration task: each scene should be restored to the arrangement given here. source.scene on each object identifies its scene and can be used to pair with other data from the same batch. The dataset is self-contained: every referenced model lives under assets/, there are no absolute paths, and no external asset library is required. - position is the model origin (not the bounding-box center), in centimeters - size is the object's target extent; the model is already scaled to match it. (x, y) are the horizontal extents and z is the height. A model's own local up-axis may not align with z, so orient it using size as the reference - rotation is…

Publicly accessible cc-by-nc-4.0

两位抖音博主的视频数据集 + SAM 2.1 语义分割结果,按博主(抖音号/UID)分别打包。 视频文件名为 序号作品ID.mp4;metadata.json 中 awemeid 与作品 ID 对应,含标题/点赞/评论等元数据。 tar 均为未压缩归档(纯打包,字节无损),解压即得与原始一致的目录树。 - frameXXXseg.png:分割叠加图(原始帧 + 彩色实例掩码) - masks.npz:键 frameXXX 对应 (Nmasks, H, W) bool 掩码数组 - report.json:每帧掩码数量与耗时统计 掩码为类无关实例分割;需要类别标签可自行接 CLIP / Grounding DINO。 recon.npz 内含每帧 pred{i}pts(点云)与 view{i}img(RGB)等,可用于点云可视化/深度/位姿估计。 四个 tar 相互独立、内部路径均相对 data/,可并行下载解压。

Publicly accessible

Dataset · Text generation

amazon-c11-nothink-distillation

Anikait Singh

This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the complete pinned C11 stopping-objective corpus. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and test splits. - asingh15/amazon-c11-distillation at 3f7302f2eb78cfa8a90110bd370237fd11e1638c (combined; catalog faff332d2bf465d37c7d6e9c0bb068522d80e3ae1e5f777688ec6c607f99f1d2)

Publicly accessible 100K<n<1M

Dataset · Text generation

amazon-c2-distillation

Anikait Singh

Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the complete pinned C11 stopping-objective corpus. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and test splits. - asingh15/amazon-c11-distillation at 3f7302f2eb78cfa8a90110bd370237fd11e1638c (combined-and-full-judge; catalog faff332d2bf465d37c7d6e9c0bb068522d80e3ae1e5f777688ec6c607f99f1d2)

Publicly accessible 100K<n<1M

Dataset · Text generation

amazon-c2-distillation-filtered

Anikait Singh

Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/ /; it is not exposed as another dataset configuration. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and test splits. - asingh15/amazon-c11-distillation-filtered at f3cbaf3082ca95ba94db218810b8b3e4fa076e81 (membership-and-combined; catalog…

Publicly accessible 10K<n<100K

Dataset · Text generation

chrono-2021-quality-harvest

Lima

Clean, cross-source-deduplicated training text available no later than 2021-12-31. Target: 15B exact anacoluthe89/chrono-2015 tokens. The build rejects any row that would exceed 15B; 20B is an external safety ceiling, not a collection goal. No lane can exceed 26.7% of the target. News and Stack Exchange additionally use per-domain/site token caps. - availability date must be on or before 2021-12-31; - publication/event date must be on or before 2021-12-31; - chunks mentioning years after 2021 are rejected; - exact normalized text is deduplicated globally across every lane; - token quotas use the submitted model family's fast tokenizer, not chars/4; - low-alpha, repetitive, boilerplate…

Access requested at publisher other

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.