SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Aligned visualizations of tactile-glove human demonstrations (Noitom capture rig) for the tacWAM project. Each clip = 4 camera panels (head RGB / head depth / left & right wrist), both hands' tactile heatmaps, and a total-pressure timeline. English overlays, 15 fps. Shared QC sheet → https://docs.google.com/spreadsheets/d/1jdWoaNQhLIEqkpm8nlQoOuralDeuo69T5fyzmRIWO44/ As you eyeball the clips, record any issues (empty/absent objects, in-air motion, weak/dead tactile channels, choppy video, mislabeled task, alignment glitches, etc.) in the sheet. Use the sessionid (the.mp4 filename) as the row key so findings are traceable. Play any.mp4 inline by clicking it, or open its direct URL…

Publicly accessible other

Dataset · Speech recognition

Dhravani

Shreesha

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…

Publicly accessible cc-by-4.0

An anonymous record of how people use MLX Model Explorer to choose an MLX model for their Mac: which model families, sizes, quantizations, memory classes and context lengths they look at, and which models they go on to open, compare or download. It also holds the community reports ("it worked", "too slow") and real MLX benchmark results that people choose to contribute. The goal is to answer, with data: what is the MLX community actually trying to run, and how well does it work? The dataset records outcomes, not clicks. Individual filter changes and page interactions are never logged. The Space buffers rows and writes a new incoming file at most once an hour, or sooner after 2,000 rows.…

Publicly accessible cc-by-4.0 n<1K

Dataset · Text generation

nepali-law-v2

Aaraj Bhatar

Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record qualityscores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-.jsonl per source document; shards are overwritten idempotently on re-runs. (grounding>=4, correctness>=3, naturalness>=3). Companion repos hold quality-gate rejects (rejected-) and records whose judge call failed (unjudged-).

Publicly accessible

Dataset · Text generation

rejected-nepali-law-v2

Aaraj Bhatar

Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record qualityscores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-.jsonl per source document; shards are overwritten idempotently on re-runs. Records judged BELOW the quality gate; each has rejectreasons. Useful for judge calibration, hard-negative mining, or re-filtering with different thresholds. Do NOT use as-is for instruction tuning.

Publicly accessible

Dataset · Image to image

tintedglass-rdt-synth

Zmy1234567890

SIRR 合成集(一般 p=1)的关键差别。 合成测试集目录:gt/t{XX}.png(透射 GT)、r/t{XX}.png(反射层)、 levels/alpha{lv}/t{XX}.png(各档输入)、alpha/(透过率场)、meta.json。 注意:无反射参考帧(T.png / GT.png)与 vehiclessd 训练数据字节级同源, DONOTUSEINTRAINING.md)。 这条数据用于 RDNet(CVPR 2025)的贴膜域微调:冻 FocalNet 主干训练其余参数, 配套权重见 zmy1234567890/tintedglass-rdnet-synth-ft。 重要负结果:合成域微调不能迁移到实拍(三套实拍配对集 gain-matched PSNR 全部变差 −0.44 ~ −3.87 dB,Wilcoxon p<0.05)。候选原因是合成 T 无噪声/JPEG、 - 合成层(T 源、R 池)来自公开学术数据集(SIR²、RRWDataset 等),仅供研究;

Publicly accessible other

Environment adaptation, training evidence, latency profiles, raw demonstrations and evaluation traces, organized by environment and experiment stage. - AirRaid: zero-latency and profile-latency experiments. Only raw demonstrations are distributed. Generate converted training datasets on the training server. Model weights remain in the dedicated model repository and are referenced from the experiment records. Pin a commit revision for reproducible downloads. All seven environments use condition/model-family/experiment/stage directories, with each environment’s shared zero-latency data stored once. Existing dataset configuration names and split names are retained and point to the reorganized…

Publicly accessible

Dataset · Time series forecasting

central-bank-exchange-rates

Today

Official exchange rates published by 103 central banks and 4 tax authorities, as one CSV per institution: 14,612,276 rows, the oldest series from 1914. Refreshed daily from the GitHub source repository. Every row is the figure the institution itself published for that date: the ECB euro reference rate, the Federal Reserve H.10 table, the Bank of England spot rates, RBI reference rates, PBoC central parity, HMRC monthly rates for VAT, US Treasury quarterly rates, and a hundred more. These are the rates that invoices, tax filings, customs declarations, transfer pricing and audits require, as opposed to market rates. Direction follows the publisher: the ECB quotes EUR → USD, the Reserve Bank…

Publicly accessible cc-by-4.0 10M<n<100M

Dataset · Text classification

nicolog

Nyaamoe

NICOLOG アニメコメントアーカイブ commeonやNCOverlayでコメント付きのアニメを楽しもう!

Publicly accessible mit

Dataset · Image to image

AmpScape

Xiao

AmpScape is a benchmark of circuit-theoretic landscape connectivity solved with the reference solvers Circuitscape.jl 5.17.1 and Omniscape.jl 0.6.2, for training and fairly comparing learned surrogates. Each sample pairs a resistance raster and a source configuration with the exact solver outputs (current-density maps, voltage maps, effective resistances, omnidirectional connectivity). Honest framing. (1) The solver is the ground truth: outputs are stored raw (float32 maps, float64 effective resistances), never normalised, clipped or post-processed; every sample records solver versions, parameters, timings and residuals. (2) Real landscapes, synthesized resistance: real tiles are genuine…

Publicly accessible cc-by-4.0 100K<n<1M

本目录是 2026-09-01 正式释文 OCR 的本地交付根目录。 1. 百度 PaddleOCR-VL-1.5 保守首识; 2. Qwen3-VL-4B-Instruct 对照页图和百度结果复核; 3. GPT-5.6 只处理分歧、低置信和版式异常页,不对已一致页做语言润色。 字段分离保存:paddleraw、qwenreview、gptreview、acceptedtext、reviewstatus、sourcesha256。古字图版、无码字和无法解释的分歧必须进入 quarantine,不得静默改成通顺文字。 资产详情、页数、SHA-256 与模型 revision 见 HFASSETMANIFEST20260901.json。

Publicly accessible

NOTICE: DO NOT USE THIS DATASET FOR ANY COMMERCIAL OR AI TRAINING PURPOSES. UNAUTHORIZED REDISTRIBUTION IS STRICTLY PROHIBITED. 本データセットは、金融庁のEDINETが公開しているXBRLデータを著作者が独自に解析・蓄積したものです。 - 本データセットは Gated Dataset です。 - データの取得を希望する場合は、Hugging Faceアカウントでログインし、アクセスリクエストを行ってください。 - 2026年5月2日:ライセンスをMITから独自規約に変更し、アクセス制限(Gated/Manual Approval)を導入しました。

Access requested at publisher other

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.