SAVRN
Search Contact SAVRN

SAVRN Model Hub

AI Training Datasets

Each dataset with its card, its structure (every split, row count and column type), its files, its license, and the models that disclose training on it.

2,760Models
859Datasets
254Papers
1,692Publishers
5,040Sourced relationships

Updated 2026-09-18 · How the library is built

859 datasets, sorted by most downloaded.

Dataset · Other

sp500-sec-edgar-filings

Joey

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Real verified extract of corporate annual (10-K) and quarterly/material filings from the SEC EDGAR system. Includes company names, CIK numbers, tickers, form types, filing dates, accession numbers, and direct SEC EDGAR URLs. Generated via Apify Actor captainhandsome/sec-edgar-filings-search. - cik: (e.g. 0000789019) - companyname: (e.g. MICROSOFT CORP) - ticker: (e.g. MSFT) - tickers: (e.g. ['MSFT']) - exchanges: (e.g. ['Nasdaq']) - sic: (e.g. 7372)…

Publicly accessible mit n<1K

Dataset · Tabular classification

california-licensed-contractors

Joey

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead generation, labor market intelligence, and compliance verification. Public registry extract of verified California licensed specialty and general building contractors from the California Contractors State License Board (CSLB). Includes business names, license numbers, current license status, issue dates, expiration dates, and classifications. Generated via Apify Actor captainhandsome/ca-contractor-license-search. - contractorname: (e.g. smith jay r) - nametype: (e.g. previous) - licensenumber: (e.g. 1018539) - city: (e.g. berkeley)…

Publicly accessible mit n<1K

Dataset · Speech recognition

dhravani-iiitdelhi-test

Shrikant Nayak

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…

Publicly accessible cc-by-4.0

Dataset · Tabular classification

austin-software-engineer-jobs

Joey

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead generation, labor market intelligence, and compliance verification. Clean structured extract of active software engineering, full stack, and AI developer job listings in the Austin, Texas metro area. Includes canonical job URLs, company names, employer ratings, location tags, salary estimates, raw posting age (e.g. 24h, 3d), and normalized estimated posting dates. Generated via Apify Actor captainhandsome/glassdoor-jobs-scraper. - jobtitle: (e.g. Entry-level Software Developer) - joburl: (e.g.…

Publicly accessible mit n<1K

Dataset · Other

MADBench-full

Anonymous

MADBench-Full contains 5,200 fully labeled execution traces from five LLMs solving synthetic escape-room tasks. The updated release adds paired tool-enabled and no-tool runs over two clue domains, making it possible to study tool selection, argument construction, tool execution failures, error recovery, and downstream error propagation in a controlled multi-agent system. Four models have 1,200 traces each: 2 clue domains × 2 tool modes × 3 temperatures × 100 rooms. claude-opus-4-8 has 400 traces at temperature 0.0 only. Each room embeds one or more reasoning problems inside escape-room clues: - gsm-hard uses challenging arithmetic word problems. - livecodebench uses chained Python programs.…

Publicly accessible mit 1K<n<10K

Dataset · Other

afdb

N

Predicted monomer structures for 292,804 unique protein sequences: 38,934 from a Gene Ontology molecular-function coupling set and 254,182 audited ancestral sequence reconstructions (312 sequences occur in both). The predictions serve as teacher labels for 1,421,494 training pairs; one structure is reused by every pair referencing that exact sequence. Sequence IDs are seq, so a prediction can always be matched back to its exact sequence. Stock open-source AlphaFold 3 with default settings — no protocol Both MSAs AlphaFold 3 builds are retained: the unpaired MSA and the paired (UniProt) MSA. The paired MSA is kept even though every input is a single chain, because AlphaFold 3 featurises it…

Publicly accessible other 100K<n<1M

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Real-time snapshot of top live streaming channels on Twitch. Includes channel display names, game/category names, stream titles, concurrent viewer counts, broadcaster languages, stream start timestamps, and thumbnail URLs. Generated via Apify Actor captainhandsome/twitch-live-streams-scraper. - streamlabel: (e.g. 1 million subscriber celebration - happyhappygal) - channelurl: (e.g. https://www.twitch.tv/happyhappygal) - channelname: (e.g. happyhappygal)…

Publicly accessible mit n<1K

GR00T-N1.7 LIBERO-X backbone features — 90-task fine-tune (LEVEL1-3) Aligned rollouts of a LIBERO-X fine-tune of GR00T-N1.7 (rohansiva/gr00t-libero-x-90task) on the LIBERO-X simulator, over the exact 90 tasks that checkpoint was fine-tuned on (30 tasks × 3 difficulty levels, LEVEL1–LEVEL3). Train and eval task sets are identical by design, so this is an in-distribution dataset for the checkpoint. 90 tasks × 20 rollouts = 1,800 episodes (600 per level). Every GR00T inference

Access requested at publisher 1K<n<10K

Dataset · Tabular classification

phoenix-hvac-contractor-leads

Joey

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead generation, labor market intelligence, and compliance verification. Real verified extract of HVAC repair, installation, and commercial contractor businesses across Phoenix, Arizona including business names, phone numbers, full addresses, ratings, review counts, Google Maps URLs, and website domains. Generated via Apify Actor captainhandsome/google-maps-business-search. - name: (e.g. Ken Muncy Air Conditioning) - placeurl: (e.g. https://www.google.com/maps/place/Ken+Muncy+Air+Conditioning) - placeid: (e.g.…

Publicly accessible mit n<1K

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Structured user review and sentiment corpus extracted from the Google Play Store. Includes app package IDs, reviewer ratings (1-5 stars), review text, user thumbs-up vote counts, review submission timestamps, and developer responses. Generated via Apify Actor captainhandsome/google-play-reviews-scraper. - appid: (e.g. com.google.android.youtube) - reviewid: (e.g. c0338d5d-4dd2-4288-ba3d-f146055e4a99) - username: (e.g. mohd Arif Arif) - userimage: (e.g.…

Publicly accessible mit n<1K

This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Public procurement dataset of prime federal awards, defense contracts, and AI grant obligations from USAspending.gov. Includes recipient names, awarding agencies, funding offices, award amounts, action dates, and award descriptions. Generated via Apify Actor captainhandsome/usaspending-federal-awards. - awardfamily: (e.g. contracts) - awardid: (e.g. DEAC0494AL85000) - recipientname: (e.g. LOCKHEED MARTIN CORP) - awardtype: (e.g. DEFINITIVE CONTRACT) - amount…

Publicly accessible mit n<1K

Original, deterministic English fixtures for Adam Pippert's personal Granite Decisions project. The original default config has 162 examples: 54 train, 54 calibration, and 54 test. These exercise the pipeline; they are not a representative quality benchmark. The source is the project's original template generator, published here as makesmokedata.py, from commit 543345ea033370484ca226424afd73164d48ca35. All dataset content and the generator are MIT licensed; see LICENSE. No third-party dataset or model output was used to generate labels. Jev was used separately for evaluation, never as a source of training labels. Each JSONL record has an id, state.request containing an English software-work…

Publicly accessible mit 10K<n<100K

Model Collections

Hand-picked starting points, each with the reason it exists.

Collection · 4 entries

Models that fit on one accelerator

Models whose publisher-reported parameter count puts them within reach of a single accelerator at common precisions. Memory needed depends on precision and serving configuration, so treat the parameter count as the starting point, not the answer.

Related SAVRN Research

The hub sits beside SAVRN's market data and infrastructure research: what models cost to run, and what it takes to run them.