P3 (Public Pool of Prompts) is a collection of prompted English datasets covering a diverse set of NLP tasks. A prompt is the combination of an input template and a target template. The templates are functions mapping a data example into natural language for the input and target sequences. For example, in the case of an NLI dataset, the data example would include fields for Premise, Hypothesis, Label. An input template would be If {Premise} is true, is it also true that {Hypothesis}?, whereas a target template can be defined with the label choices Choices[label]. Here Choices is prompt-specific metadata that consists of the options yes, maybe, no corresponding to label being entailment (0)…
Publicly accessible
apache-2.0
100M<n<1B
1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours. - 100+ real-world scenarios across 5 target domains. - 200+ types of tasks: - 87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc. Your browser does not support the video tag. Your browser does not support the video tag. Your browser does not support the video tag. - [2025/3/1] AgiBot World Beta released. - [ ] AgiBot World Colosseum:Comprehensive platform (expected release date: 2025) - [ ] 2025 AgiBot World Challenge (expected release date: 2025) To download the full dataset, you can use the following code. If you encounter any issues, please refer to the official Hugging Face documentation. If…
Access requested at publisher
n>1T
Publicly accessible
100M<n<1B
ngii-map-full-light Light point/line extract from NGII 1/1000 topographic data for Korea. Not for shipping into GitHub — use this Hugging Face dataset instead. CRS Korea2000CentralBelt2010 projected meters [x, y] Layers (per region under byregion/ /) Layer Description C023 poles (전주/통신주) C022 lights (가로등·보안등) A002 roads (도로 중심선) B001tiny building footprints <25 m² as centroids B002 lines (구분/재질 라인) Also
Publicly accessible
other
1M<n<10M
InternData-A1 InternData-A1 is a hybrid synthetic-real manipulation dataset containing over 630k trajectories and 7,433 hours across 4 embodiments, 18 skills, 70 tasks, and 227 scenes, covering rigid, articulated, deformable, and fluid-object manipulation. Your browser does not support the video tag. Your browser does not support the video tag.
Access requested at publisher
n>1T
SWE-rebench is a large-scale dataset designed to support training and evaluation of LLM-based software engineering (SWE) agents, building upon and expanding our earlier release, SWE-bench-extra. It is constructed using a fully automated pipeline that continuously extracts real-world interactive SWE tasks from GitHub repositories at scale, as detailed in our paper SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. The dataset currently comprises over 21,000 issue–pull request pairs from 3,400+ Python repositories, each validated for correctness through automated environment setup and test execution. A curated subset of these…
Publicly accessible
cc-by-4.0
Patrick Rim, Kevin Harris, Braden Copple, Shangchen Han, Xu Xie, Ivan Shugurov, Sizhe An, He Wen, Alex Wong, Tomas Hodan, and Kun He CVPR 2026; https://arxiv.org/abs/2603.28760 SHOW3D is a large-scale multi-view dataset of hand–object interactions captured in the wild. It is intended to advance research on egocentric 3D hand–object interaction understanding, and generalization of perception models to real-world
Publicly accessible
cc-by-nc-4.0
1K<n<10K
Per-checkpoint mechanistic metrics for Beetle language models, tracking how the induction circuit forms during training. The files here have six different schemas, so they are exposed as separate configs. Loading the directory as a single table fails with a cast error — that is why the configs above exist, not a bug. repeatdependence is the column that matters for deciding whether a high-PS head is really an induction head: candidates score ~0.7, every other head ~0.0003. A head that attends "somewhere earlier" can beat chance on PS alone without needing the repeat. A phase transition. PS sits at the chance floor (0.028, which is exactly uniform causal attention at the mean query position)…
Publicly accessible
apache-2.0
This repository stores a standalone, NeuROK-compatible deformation corpus under objaverse/finetunedeformation. It combines processed public animation data with generated MPM solid trajectories. Training consumers should read index/all.jsonl (or its train.jsonl and val.jsonl split files) rather than discovering NPZ files by directory traversal. - 1,998 DynamicObjaverseProcessed geometries; - 500 DyMesh / AnimateAnyMesh geometries; - 2,498 trainable rows in the canonical combined index; The generated target is 5,000 geometries × 2 accepted clips = 10,000 clips. Its deterministic 24-cell cycle balances elastic, plastic, sand, and snow, with every material parameter drawn from one of four…
Publicly accessible
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Fresh job listings for AI, ML, and Software Engineering positions across the United States. Includes canonical job posting URLs, hiring company names, job titles, locations, raw posting age, and estimated posting dates. Generated via Apify Actor captainhandsome/linkedin-public-jobs-search. - title: (e.g. AI Engineer) - company: (e.g. P-1 AI) - location: (e.g. San Francisco Bay Area) - url: (e.g. https://www.linkedin.com/jobs/view/ai-engineer-at-p-1-ai-443)…
Publicly accessible
mit
n<1K
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Real verified extract of corporate annual (10-K) and quarterly/material filings from the SEC EDGAR system. Includes company names, CIK numbers, tickers, form types, filing dates, accession numbers, and direct SEC EDGAR URLs. Generated via Apify Actor captainhandsome/sec-edgar-filings-search. - cik: (e.g. 0000789019) - companyname: (e.g. MICROSOFT CORP) - ticker: (e.g. MSFT) - tickers: (e.g. ['MSFT']) - exchanges: (e.g. ['Nasdaq']) - sic: (e.g. 7372)…
Publicly accessible
mit
n<1K
MADBench-Full contains 5,200 fully labeled execution traces from five LLMs solving synthetic escape-room tasks. The updated release adds paired tool-enabled and no-tool runs over two clue domains, making it possible to study tool selection, argument construction, tool execution failures, error recovery, and downstream error propagation in a controlled multi-agent system. Four models have 1,200 traces each: 2 clue domains × 2 tool modes × 3 temperatures × 100 rooms. claude-opus-4-8 has 400 traces at temperature 0.0 only. Each room embeds one or more reasoning problems inside escape-room clues: - gsm-hard uses challenging arithmetic word problems. - livecodebench uses chained Python programs.…
Publicly accessible
mit
1K<n<10K
Predicted monomer structures for 292,804 unique protein sequences: 38,934 from a Gene Ontology molecular-function coupling set and 254,182 audited ancestral sequence reconstructions (312 sequences occur in both). The predictions serve as teacher labels for 1,421,494 training pairs; one structure is reused by every pair referencing that exact sequence. Sequence IDs are seq, so a prediction can always be matched back to its exact sequence. Stock open-source AlphaFold 3 with default settings — no protocol Both MSAs AlphaFold 3 builds are retained: the unpaired MSA and the paired (UniProt) MSA. The paired MSA is kept even though every input is a single chain, because AlphaFold 3 featurises it…
Publicly accessible
other
100K<n<1M
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Real-time snapshot of top live streaming channels on Twitch. Includes channel display names, game/category names, stream titles, concurrent viewer counts, broadcaster languages, stream start timestamps, and thumbnail URLs. Generated via Apify Actor captainhandsome/twitch-live-streams-scraper. - streamlabel: (e.g. 1 million subscriber celebration - happyhappygal) - channelurl: (e.g. https://www.twitch.tv/happyhappygal) - channelname: (e.g. happyhappygal)…
Publicly accessible
mit
n<1K
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Structured user review and sentiment corpus extracted from the Google Play Store. Includes app package IDs, reviewer ratings (1-5 stars), review text, user thumbs-up vote counts, review submission timestamps, and developer responses. Generated via Apify Actor captainhandsome/google-play-reviews-scraper. - appid: (e.g. com.google.android.youtube) - reviewid: (e.g. c0338d5d-4dd2-4288-ba3d-f146055e4a99) - username: (e.g. mohd Arif Arif) - userimage: (e.g.…
Publicly accessible
mit
n<1K
This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Public procurement dataset of prime federal awards, defense contracts, and AI grant obligations from USAspending.gov. Includes recipient names, awarding agencies, funding offices, award amounts, action dates, and award descriptions. Generated via Apify Actor captainhandsome/usaspending-federal-awards. - awardfamily: (e.g. contracts) - awardid: (e.g. DEAC0494AL85000) - recipientname: (e.g. LOCKHEED MARTIN CORP) - awardtype: (e.g. DEFINITIVE CONTRACT) - amount…
Publicly accessible
mit
n<1K
Value-only chess dataset. Each row is a unique board labeled with official Stockfish 19 UCIShowWDL. Use wdl as the value target. Do not treat this as MultiPV policy data. 9,100,000 rows on this repo (append-only waves). Source id 4. Compact move vocab (1968). Shards keep a global index: wave 1 is data/shard000000–000199. Later waves continue. This is SF19's fishtest-LTC self-play WDL model (eval + remaining material). It is not FIDE/Lichess Elo and not a sigmoid of cp. Official WDL is Honor split. split=1 is a 5% holdout. Do not invent a new hash holdout. For value training, use wdl with KL / cross-entropy against a 3-class head ordered win/draw/loss. Drop or keep wdlsource==2 terminals…
Publicly accessible
mit
1M<n<10M
This public dataset contains 65,131 rendered counterfactual cases organized into 8,491 same-history action/future groups. Each case has 4 historical and 8 future CAMF0 frames plus WorldEngine metadata. Archives preserve complete pair groups. The audit bundle contains group JSON manifests, receipts, validation reports and the completion contract. Large generator-intermediate subset PKLs are intentionally excluded because the CAST paired-cache builder consumes the group JSON and rendered frames only. Download all files and run bash extractall.sh DESTINATION. The extraction script verifies every archive against SHA256SUMS first.
Publicly accessible
other