SAVRN Model Hub
AI Training Datasets
Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.
Updated 2026-09-18 · How the library is built
859 datasets, sorted by most downloaded.
A fully human-reviewed, re-annotated version of the classic Raccoon object detection dataset. Every one of the 193 images was manually reviewed; loose, missing, and incorrect bounding boxes were corrected. Every image was compared with the original VOC annotations using greedy IoU matching (matched ≥ 0.5, unchanged ≥ 0.95): Key observations - The original boxes were often loose (e.g., raccoon-1: box (81,88)-(522,408) → tightened to (80,105)-(529,405), IoU 0.91). The average IoU of matched boxes is 0.810, i.e. the corrected boxes bound the raccoon noticeably more tightly. - Occluded / partially visible raccoons were kept with tight boxes; several missed instances were added. - 7 exact/near…
Reproduction package for the framework paper: frozen-gate machinery, the deployment/long-horizon/domain-shift evaluations, the agentic tiers, and the shared v3 validation campaign. See INDEX.md for the claim→artifact map, MANIFEST.sha256 for per-file hashes, and v3-campaign/EVIDENCE-RELEASE.md for the verification recipe. PENDING.md lists what is not yet packaged and why. Code is Apache-2.0; third-party corpora and model assets retain their upstream licences. This is a partial evidence release; consult PENDING.md before using it. MANIFEST.sha256 covers this snapshot. External evidence gaps remain listed in PENDING.md. See REPRODUCE.md for download and verification instructions. P6 audit…
Part of the AI4Manufacturing FORGE corpus (Category C, task T-C1), and the corpus's first arc-welding dataset. Each record is one labelled stretch of a pulsed gas-metal arc weld, drawn as the heat input of every current pulse along that stretch with the heat-input window laid over it: a thin grey trace through the individual pulses, a rolling-median line, and a shaded band with dashed limits. The only thing to read off it is how much of the trace leaves the band. That is not the usual Category-C picture, and it is not an oversight. Every other signal dataset here asks which fault line is present; welding has none. A welder judges a stretch of weld by how much heat went into it and whether…
Annoy: This should be a the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
Annoy: This should be a the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
Annoy: This should be a the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
Annoy: This should be a the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
Annoy: This should be a the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
Annoy: This should be a the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
Annoy: This should be a the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
This is the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Please also check the raw data after our processing if you are interested: toolathlon5/Annoy-PyEdu-Rs-Raw. Models Introduction While having full executable code theoretically allows us to generate reliable execution trajectories as responses, two challenges arise: 1) Obtaining a deterministic reverse function for input prediction is impractical; 2) Automatically constructed trajectories are constrained by pre-designed templates and lack the expressiveness and generalizability of free-form natural language reasoning. Thus, we adopt a fully LLM-based approach…
the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
the resource page of the our resources collection on Huggingface, we highlight your currect position with a blue block. Dataset Dataset Link Annoy-PythonEdu-Rs Please also check the raw data after our processing
Model Collections
Hand-picked starting points, each with the reason it exists.
Collection · 4 entries
Embedding models for retrieval
Sentence and document embedding models used to build retrieval systems. Dimension and sequence length matter more than size here, and both come from the publisher.
Collection · 4 entries
Models that fit on one accelerator
Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.
Collection · 6 entries
Open-weight text models worth knowing
Widely used open-weight language models, chosen because each one is a distinct family rather than a variant of the one above it. Selection, not a ranking.
Collection · 3 entries
Speech and audio models
Recognition and synthesis models, grouped so the two directions are easy to compare.
Related SAVRN Research
The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.
SAVRN Index
What open models cost to run
The same open-weight model priced by every host that serves it, per million tokens.
Research Hub
Data center trackers and maps
Moratoriums, permits, power, water and capital behind the facilities that run these models.
Method
How the Model Hub is built
Sources, evidence labels, refresh behaviour, and the limits of every comparison here.
