SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Text retrieval

ipfs_qatar_laws

Benjamin J Barber

Research snapshot of official national legislation from Al-Meezan (Qatari Legal Portal / Ministry of Justice). Not legal advice. The official gazette / authentic source prevails over this corpus. - data/laws.parquet — one row per instrument (LAWFLAT, zstd). - data/articles.parquet — article/section rows for the same snapshot. - scrapers/collectqa.py — research collector used to build this snapshot. Official Al-Meezan texts; authentic Official Gazette prevails; not legal advice. This dataset is not legal advice. The official source prevails. - https://www.almeezan.qa/ National laws in Arabic from almeezan.qa. Constitution included. In-force national laws from official Al-Meezan…

Publicly accessible other 10K<n<100K

Mirrored from https://councilof.ai at 2026-09-18T03:43:32Z. councilof.ai is authoritative. This mirror exists for one measured reason: the origin answers HTTP 403 (Cloudflare error 1010) to a plain Python or Perl client, on the API as well as the site, so a machine consumer following our own published instructions is refused. Hugging Face serves those same clients. Every file carries the URL it came from and the sha256 of the bytes as fetched, in MIRROR-MANIFEST.json. If a file here disagrees with the origin, the origin wins and this mirror is stale. Nothing here is a certification. We measure and we never certify, and verification is free.

Publicly accessible cc-by-4.0

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Fresh job listings for AI, ML, and Software Engineering positions across the United States. Includes canonical job posting URLs, hiring company names, job titles, locations, raw posting age, and estimated posting dates. Generated via Apify Actor captainhandsome/linkedin-public-jobs-search. - title: (e.g. AI Engineer) - company: (e.g. P-1 AI) - location: (e.g. San Francisco Bay Area) - url: (e.g. https://www.linkedin.com/jobs/view/ai-engineer-at-p-1-ai-443)…

Publicly accessible mit n<1K

Dataset · Text retrieval

ipfs_solomonislands_laws

Benjamin J Barber

Research snapshot of official national legislation from AGC Legislation Portal — Acts currently in force (attorneygenerals.gov.sb). Not legal advice. The official gazette / authentic source prevails over this corpus. - data/laws.parquet — one row per instrument. - data/articles.parquet — article/section rows for the same snapshot. - scrapers/collectsb.py — research collector used to build this snapshot. Official Solomon Islands AGC texts; authentic text prevails; not legal advice. This dataset is not legal advice. The official source prevails. - https://attorneygenerals.gov.sb/legislation/legislation-portal/ Download Monitor Acts currently in force via official AGC portal. Polite ≤1 req/s.…

Publicly accessible other 1K<n<10K

Evaluation logs of every model row reported in the ICLR 2027 submission MemGUI-RL (project page: https://memgui-rl-anonymous.github.io/). One zip archive per evaluation session: memguibench/.zip - per-task ConAct trajectories (traj.json, three attempts where applicable), MemGUI-Eval judge outputs (results.csv, per-attempt judgments, failure labels, information-retention analyses) and run metadata. Screenshots are omitted for size; the paper's case-study tasks are included under casestudies/ with down-scaled screenshots. mobileworld/.zip - per-task trajectories and the MobileWorld GUI-only evaluation report. manifest.json maps every archive to the row of the paper it supports.…

Publicly accessible apache-2.0

Dataset · Feature extraction

epstein-index

Robb Doering

Machine-readable text of the publicly released Jeffrey Epstein–related document collections: the DOJ "EFTA" disclosures (Data Sets 1–12), House Oversight releases, court records (Maxwell, Doe v. Epstein/Indyke, Florida v. Epstein, USVI v. JPMorgan, and others), and FOIA productions (FBI, BOP, CBP, Florida). All content is public-record material released by the U.S. Department of Justice, federal/state courts, and FOIA respondents. This is a living dataset: this snapshot publishes 32,172 documents — including the full House Oversight estate releases of Oct 17 (8,718 pages) and Nov 12 (22,903 pages); roughly 33,000 source PDFs are downloaded and mid-OCR at snapshot time, with over 1.3M…

Publicly accessible other 10K<n<100K

Dream.exe Benchmark Code · DVD checkpoints This release preserves the original 101 benchmark cases. The Dream.exe authors constructed the task instances, frozen scenes, camera settings, generation inputs, and reference data using RoboCasa data and environments. Download and reproduce After completing the code repository's installation guide, run from the code repository root: hf download kaimingyang/Dream.exe --repo-type dataset \ --include 'bench/'

Publicly accessible

Dataset · Speech recognition

dhravani-mit-test

Shrikant Nayak

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…

Publicly accessible cc-by-4.0

Dataset · Speech recognition

dhravani-iitpatna-test

Shrikant Nayak

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…

Publicly accessible cc-by-4.0

TableVerse -> Gaussian splats Real Gaussian splats (.spz) converted from TableVerse-100K (ByteDance, CC BY 4.0) via this project's own analytic, training-free mesh-to-splat converter (splataverse.cli mesh-to-splat --split components, CPU/torch backend). Source attribution Dataset: TableVerse-100K, ByteDance (https://huggingface.co/datasets/ByteDance/TableVerse) Paper: Wang, Boyuan et al., "TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded

Access requested at publisher cc-by-4.0

Dataset · Speech recognition

dhravani-IIT_Guwahati-test

Shrikant Nayak

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…

Publicly accessible cc-by-4.0

Dataset · Speech recognition

dhravani-IGDTUW_Delhi-test

Shrikant Nayak

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…

Publicly accessible cc-by-4.0

Dataset · Tabular classification

ula-launches

Julien Simon

Part of a dataset collection on Hugging Face. Complete United Launch Alliance (ULA) launch manifest — past and upcoming flights of Atlas V, Delta II, Delta IV, Delta IV Heavy, and Vulcan Centaur — sourced from The Space Devs Launch Library 2 API. ULA is the Boeing-Lockheed Martin joint venture formed in 2006 to consolidate US national security space launches under a single EELV (Evolved Expendable Launch Vehicle) provider. For nearly a decade ULA enjoyed a monopoly on high-value national security payloads before SpaceX's Falcon 9 broke into the market. ULA retired its Delta family in favor of Vulcan Centaur, which flew its first flight in 2024 and is now ramping to take over both national…

Publicly accessible other n<1K

Dataset · Speech recognition

dhravani-iitdelhi-test

Shrikant Nayak

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…

Publicly accessible cc-by-4.0

An anonymous record of what people try to run locally, gathered by Local Model Explorer: the hardware they plan for (GPU memory, number of cards, system or unified memory), the models and context lengths they look at, which GGUF quants they open and copy commands for, and the llama-bench results and reports they choose to share. The question it answers: what hardware do local LLM users have, what do they try to run on it, and how fast does it actually go? The dataset records outcomes, not clicks. Every row has every column; columns that don't apply to an event type are null. Flagged rows are kept so filters stay transparent; the Space leaves them out of its statistics. Early each month the…

Publicly accessible cc-by-4.0 n<1K

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.