Automated detection of expressway pavement cracks is critically important for infrastructure inspection and maintenance. In this task, range information provides a key geometric cue for distinguishing genuine cracks from pavement textures, shadows, stains, and other visual interference. However, industrial-grade pavement data acquisition equipment is expensive, and pixel-wise annotation is labor intensive. Consequently, most existing pavement crack datasets contain only visual imagery and cannot precisely represent the geometric morphology of the pavement surface. GTCrack was developed to alleviate this data bottleneck. A survey vehicle equipped with two laser line profilers was used to…
SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
This repo holds every shape of the released dataset described under "The data", packed as tar shards per representation. Each split's uuids are sorted and cut into shards of 1,000 shapes. A tar holds one representation's files for one shard (atlas: atlas.npz; voxels: the two.vxz files; multiview: multiview/; thumbnail: thumbnail.png) at paths /, as in the tree under "The data". Shard k holds the same shapes in every representation. The atlas.npz files here are stored compressed; the arrays are the same and np.load reads them the same way. Emissive 3D shapes from TexVerse with their material and emission maps in three representations: a UV atlas, sparse voxels, and six rendered views. One…
A long-context retrieval and question-answering benchmark with 1,143,371 documents and 10,000 questions, including 50 questions about 0G. It extends the ms100M bank from MSA-RAG-BENCHMARKS with 20 million additional text tokens and 655 new questions grounded in the added documents. 120M is the benchmark's nominal size. The actual corpus contains 125,708,601 raw tokens under the MSA-4B tokenizer: 105,708,601 original tokens plus exactly 20,000,000 added tokens. Counts exclude prompts, wrappers, and special tokens. This is a derived benchmark, not an official Microsoft MS MARCO release. The original documents occupy positions 0:961686; the original questions occupy positions 0:9345. Their…
RL prompt corpus of the ICLR 2027 submission MemGUI-RL (project page: https://memgui-rl-anonymous.github.io/). Every record is one annotated state of a MemGUI-3K trajectory in the ConAct conversation format consumed by the FARPO trainer (https://github.com/memgui-rl-anonymous/MemGUI-RL): Each record has taskid, stepnumber, ispositive, badstep, rawresponse (the annotated ConAct response:,,,, ), conversations (system prompt, user prompt with the folded history / UI memory / recent step record and the screenshot, reference assistant turn) and metadata (impact / reasonableness labels of the step, trajectory ids). Screenshots are not duplicated here: image references have the form…
This dataset holds the starting state of a scene restoration task: compared with the target state, 1-3 objects in each scene have been moved to random positions and need to be put back. source.scene on each object identifies its scene. - position is the model origin (not the bounding-box center), in centimeters - size is the object's target extent; the model is already scaled to match it. (x, y) are the horizontal extents and z is the height. A model's own local up-axis may not align with z, so orient it using size as the reference - rotation is the complete absolute orientation as a quaternion in [x, y, z, w] order. Apply it directly, per object, and do not compose any additional rotation…
This dataset holds the target state of a scene restoration task: each scene should be restored to the arrangement given here. source.scene on each object identifies its scene and can be used to pair with other data from the same batch. The dataset is self-contained: every referenced model lives under assets/, there are no absolute paths, and no external asset library is required. - position is the model origin (not the bounding-box center), in centimeters - size is the object's target extent; the model is already scaled to match it. (x, y) are the horizontal extents and z is the height. A model's own local up-axis may not align with z, so orient it using size as the reference - rotation is…
两位抖音博主的视频数据集 + SAM 2.1 语义分割结果,按博主(抖音号/UID)分别打包。 视频文件名为 序号作品ID.mp4;metadata.json 中 awemeid 与作品 ID 对应,含标题/点赞/评论等元数据。 tar 均为未压缩归档(纯打包,字节无损),解压即得与原始一致的目录树。 - frameXXXseg.png:分割叠加图(原始帧 + 彩色实例掩码) - masks.npz:键 frameXXX 对应 (Nmasks, H, W) bool 掩码数组 - report.json:每帧掩码数量与耗时统计 掩码为类无关实例分割;需要类别标签可自行接 CLIP / Grounding DINO。 recon.npz 内含每帧 pred{i}pts(点云)与 view{i}img(RGB)等,可用于点云可视化/深度/位姿估计。 四个 tar 相互独立、内部路径均相对 data/,可并行下载解压。
Small dataset for experiments related to project Aide. - Fabien Allemand ([email protected])
This preserves the selected C11 writer/criterion-judge examples and their row geometry while discarding all teacher scratch reasoning. Membership follows the complete pinned C11 stopping-objective corpus. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and test splits. - asingh15/amazon-c11-distillation at 3f7302f2eb78cfa8a90110bd370237fd11e1638c (combined; catalog faff332d2bf465d37c7d6e9c0bb068522d80e3ae1e5f777688ec6c607f99f1d2)
Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the complete pinned C11 stopping-objective corpus. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and test splits. - asingh15/amazon-c11-distillation at 3f7302f2eb78cfa8a90110bd370237fd11e1638c (combined-and-full-judge; catalog faff332d2bf465d37c7d6e9c0bb068522d80e3ae1e5f777688ec6c607f99f1d2)
Each retained trajectory contributes one direct rubric-writer target and one full-rubric listwise judge target over the original 40 candidates. Teacher scratch reasoning is discarded. Membership follows the pinned C11 quality-filter policy. Signed aggregate filter provenance is under quality/ /; it is not exposed as another dataset configuration. The six configurations cross the two candidate variants with the three frozen stopping objectives. Every configuration exposes only its combined training view and preserves the native train, validation, and test splits. - asingh15/amazon-c11-distillation-filtered at f3cbaf3082ca95ba94db218810b8b3e4fa076e81 (membership-and-combined; catalog…
Clean, cross-source-deduplicated training text available no later than 2021-12-31. Target: 15B exact anacoluthe89/chrono-2015 tokens. The build rejects any row that would exceed 15B; 20B is an external safety ceiling, not a collection goal. No lane can exceed 26.7% of the target. News and Stack Exchange additionally use per-domain/site token caps. - availability date must be on or before 2021-12-31; - publication/event date must be on or before 2021-12-31; - chunks mentioning years after 2021 are rejected; - exact normalized text is deduplicated globally across every lane; - token quotas use the submitted model family's fast tokenizer, not chars/4; - low-alpha, repetitive, boilerplate…
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.
