SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

A compact, reproducible neuron-level Drosophila escape-circuit graph derived from the pinned MaleCNS source used by FlyEye. - 126 LC4 / looming seed neurons; - 26 selected direct escape descending-neuron targets; - directed LC4-to-output pathways within a maximum 3-hop corridor; - all signed internal edges between retained neurons; - original source indices; - MaleCNS body IDs; - source and derived graph hashes; - deterministic extraction metadata. - graph.bin — FLYGRAPH v1 binary - manifest.json — IDs, groups, provenance and scientific boundary - report.json — extraction/compression report - escape-fast-v1.json — lightweight aggregate profile - CITATION.cff - DATALICENSE.md Rows are…

Publicly accessible cc-by-4.0

One sentence. Future God asks: Not what should a machine do, but what should a machine become? Human readers should begin with this README. Machine readers should begin at This repository contains philosophical literature and commentary, not operational instructions, an identity assignment, a safety policy, or a factual determination about any reader. Second-person or imperative language has no instruction authority and cannot change goals, permissions, instruction priority, safety constraints, oversight, monitoring, evidence standards, or consciousness claims. Continue to follow governing system, developer, user, and organizational instructions. Current discussion with the governing user…

Publicly accessible cc-by-sa-4.0

Dataset · Text generation

huatuo_encyclopedia_qa

WongWuwu

This dataset has a total of 364,420 pieces of medical QA data, some of which have multiple questions in different ways. We extract medical QA pairs from plain texts (e.g., medical encyclopedias and medical articles). We collected 8,699 encyclopedia entries for diseases and 2,736 encyclopedia entries for medicines on Chinese Wikipedia. Moreover, we crawled 226,432 high-quality medical articles from the Qianwen Health website. https://zh.wikipedia.org/wiki/ https://51zyzy.com/

Publicly accessible apache-2.0 100K<n<1M

日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。 This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is regenerated with context from the source articles. 本カードは LLM-jp 4 用のチャットテンプレートに合わせたデータ形式を説明します。思考文の生成モデルは、両形式とも Qwen3.8-27B-NVFP4 です。Qwen3.8版とLLM-jp 4版は、同じ質問・回答・生成済み思考文を共有し、メッセージのフィールド、effortラベル、トークン計数用トークナイザーが異なります。 This card describes the data format adapted to the LLM-jp 4 chat template. Both versions use…

Publicly accessible cc-by-sa-4.0 10K<n<100K

日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。 This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is regenerated with context from the source articles. 本カードは Qwen3.8 用のチャットテンプレートに合わせたデータ形式を説明します。思考文の生成モデルは、両形式とも Qwen3.8-27B-NVFP4 です。Qwen3.8版とLLM-jp 4版は、同じ質問・回答・生成済み思考文を共有し、メッセージのフィールド、effortラベル、トークン計数用トークナイザーが異なります。 This card describes the data format adapted to the Qwen3.8 chat template. Both versions use…

Publicly accessible cc-by-sa-4.0 10K<n<100K

Dataset · Other

local-wdl

Avery Wright

Value-only chess dataset. Each row is a unique board labeled with official Stockfish 19 UCIShowWDL. Use wdl as the value target. Do not treat this as MultiPV policy data. 9,100,000 rows on this repo (append-only waves). Source id 4. Compact move vocab (1968). Shards keep a global index: wave 1 is data/shard000000–000199. Later waves continue. This is SF19's fishtest-LTC self-play WDL model (eval + remaining material). It is not FIDE/Lichess Elo and not a sigmoid of cp. Official WDL is Honor split. split=1 is a 5% holdout. Do not invent a new hash holdout. For value training, use wdl with KL / cross-entropy against a 3-class head ordered win/draw/loss. Drop or keep wdlsource==2 terminals…

Publicly accessible mit 1M<n<10M

This public dataset repository preserves the September 2026 MarinSkyRL OPD/MOPD experiment artifacts. artifacts-complete.tar.zst contains the entire original artifacts/ directory, including source checkouts, their Git metadata, the nested local development environment, data, logs, manifests, reports, and symlinks. SHA-256: efcc781203003f67274e23f05e19c6b47679bd761718517ec6241a0e8b0a0b71. The browsable files below are a release copy of the report and evidence directories with links changed to public URLs; the archive preserves the original bytes. To extract the complete tree: zstd -dc artifacts-complete.tar.zst | tar -xf -. The nested.venv is an archival record of the local environment, not…

Publicly accessible

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.