This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Real verified extract of corporate annual (10-K) and quarterly/material filings from the SEC EDGAR system. Includes company names, CIK numbers, tickers, form types, filing dates, accession numbers, and direct SEC EDGAR URLs. Generated via Apify Actor captainhandsome/sec-edgar-filings-search. - cik: (e.g. 0000789019) - companyname: (e.g. MICROSOFT CORP) - ticker: (e.g. MSFT) - tickers: (e.g. ['MSFT']) - exchanges: (e.g. ['Nasdaq']) - sic: (e.g. 7372)…
Publicly accessible
mit
n<1K
Dataset · Tabular classification
Joey
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead generation, labor market intelligence, and compliance verification. Public registry extract of verified California licensed specialty and general building contractors from the California Contractors State License Board (CSLB). Includes business names, license numbers, current license status, issue dates, expiration dates, and classifications. Generated via Apify Actor captainhandsome/ca-contractor-license-search. - contractorname: (e.g. smith jay r) - nametype: (e.g. previous) - licensenumber: (e.g. 1018539) - city: (e.g. berkeley)…
Publicly accessible
mit
n<1K
Publicly accessible
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. - ⌨ Keyboard shortcuts for efficiency 1. Create a transcript CSV file with your content: 2. Start the Flask application: 3. Access the interface: 1. Authentication 2. Session Setup - Click "Start Session" 3. Recording - Use on-screen controls or keyboard shortcuts: - R: Start recording / Stop recording - Space: Play recording - Enter: Save…
Publicly accessible
cc-by-4.0
Publicly accessible
Dataset · Tabular classification
Joey
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead generation, labor market intelligence, and compliance verification. Clean structured extract of active software engineering, full stack, and AI developer job listings in the Austin, Texas metro area. Includes canonical job URLs, company names, employer ratings, location tags, salary estimates, raw posting age (e.g. 24h, 3d), and normalized estimated posting dates. Generated via Apify Actor captainhandsome/glassdoor-jobs-scraper. - jobtitle: (e.g. Entry-level Software Developer) - joburl: (e.g.…
Publicly accessible
mit
n<1K
Publicly accessible
MADBench-Full contains 5,200 fully labeled execution traces from five LLMs solving synthetic escape-room tasks. The updated release adds paired tool-enabled and no-tool runs over two clue domains, making it possible to study tool selection, argument construction, tool execution failures, error recovery, and downstream error propagation in a controlled multi-agent system. Four models have 1,200 traces each: 2 clue domains × 2 tool modes × 3 temperatures × 100 rooms. claude-opus-4-8 has 400 traces at temperature 0.0 only. Each room embeds one or more reasoning problems inside escape-room clues: - gsm-hard uses challenging arithmetic word problems. - livecodebench uses chained Python programs.…
Publicly accessible
mit
1K<n<10K
Predicted monomer structures for 292,804 unique protein sequences: 38,934 from a Gene Ontology molecular-function coupling set and 254,182 audited ancestral sequence reconstructions (312 sequences occur in both). The predictions serve as teacher labels for 1,421,494 training pairs; one structure is reused by every pair referencing that exact sequence. Sequence IDs are seq, so a prediction can always be matched back to its exact sequence. Stock open-source AlphaFold 3 with default settings — no protocol Both MSAs AlphaFold 3 builds are retained: the unpaired MSA and the paired (UniProt) MSA. The paired MSA is kept even though every input is a single chain, because AlphaFold 3 featurises it…
Publicly accessible
other
100K<n<1M
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Real-time snapshot of top live streaming channels on Twitch. Includes channel display names, game/category names, stream titles, concurrent viewer counts, broadcaster languages, stream start timestamps, and thumbnail URLs. Generated via Apify Actor captainhandsome/twitch-live-streams-scraper. - streamlabel: (e.g. 1 million subscriber celebration - happyhappygal) - channelurl: (e.g. https://www.twitch.tv/happyhappygal) - channelname: (e.g. happyhappygal)…
Publicly accessible
mit
n<1K
GR00T-N1.7 LIBERO-X backbone features — 90-task fine-tune (LEVEL1-3) Aligned rollouts of a LIBERO-X fine-tune of GR00T-N1.7 (rohansiva/gr00t-libero-x-90task) on the LIBERO-X simulator, over the exact 90 tasks that checkpoint was fine-tuned on (30 tasks × 3 difficulty levels, LEVEL1–LEVEL3). Train and eval task sets are identical by design, so this is an in-distribution dataset for the checkpoint. 90 tasks × 20 rollouts = 1,800 episodes (600 per level). Every GR00T inference
Access requested at publisher
1K<n<10K
Dataset · Tabular classification
Joey
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead generation, labor market intelligence, and compliance verification. Real verified extract of HVAC repair, installation, and commercial contractor businesses across Phoenix, Arizona including business names, phone numbers, full addresses, ratings, review counts, Google Maps URLs, and website domains. Generated via Apify Actor captainhandsome/google-maps-business-search. - name: (e.g. Ken Muncy Air Conditioning) - placeurl: (e.g. https://www.google.com/maps/place/Ken+Muncy+Air+Conditioning) - placeid: (e.g.…
Publicly accessible
mit
n<1K
Access requested at publisher
Publicly accessible
Publicly accessible
Publicly accessible
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Structured user review and sentiment corpus extracted from the Google Play Store. Includes app package IDs, reviewer ratings (1-5 stars), review text, user thumbs-up vote counts, review submission timestamps, and developer responses. Generated via Apify Actor captainhandsome/google-play-reviews-scraper. - appid: (e.g. com.google.android.youtube) - reviewid: (e.g. c0338d5d-4dd2-4288-ba3d-f146055e4a99) - username: (e.g. mohd Arif Arif) - userimage: (e.g.…
Publicly accessible
mit
n<1K
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Public procurement dataset of prime federal awards, defense contracts, and AI grant obligations from USAspending.gov. Includes recipient names, awarding agencies, funding offices, award amounts, action dates, and award descriptions. Generated via Apify Actor captainhandsome/usaspending-federal-awards. - awardfamily: (e.g. contracts) - awardid: (e.g. DEAC0494AL85000) - recipientname: (e.g. LOCKHEED MARTIN CORP) - awardtype: (e.g. DEFINITIVE CONTRACT) - amount…
Publicly accessible
mit
n<1K
Publicly accessible
Publicly accessible
Original, deterministic English fixtures for Adam Pippert's personal Granite Decisions project. The original default config has 162 examples: 54 train, 54 calibration, and 54 test. These exercise the pipeline; they are not a representative quality benchmark. The source is the project's original template generator, published here as makesmokedata.py, from commit 543345ea033370484ca226424afd73164d48ca35. All dataset content and the generator are MIT licensed; see LICENSE. No third-party dataset or model output was used to generate labels. Jev was used separately for evaluation, never as a source of training labels. Each JSONL record has an id, state.request containing an English software-work…
Publicly accessible
mit
10K<n<100K
Publicly accessible
Publicly accessible
Publicly accessible