The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE). Natural Language Inference (NLI), also known as Recognizing Textual Entailment (RTE), is the task of determining the inference relation between two (short, ordered) texts: entailment, contradiction, or neutral (MacCartney and Manning 2008). See the corpus webpage for a list of published results. The language in the dataset is English as spoken by users of the website Flickr and as spoken by…
Publicly accessible
cc-by-sa-4.0
100K<n<1M
HotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) the questions require finding and reasoning over multiple supporting documents to answer; (2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas; (3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason with strong supervision and explain the predictions; (4) we offer a new type of factoid comparison questions to test QA systems’ ability to extract relevant facts and perform necessary comparison. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields…
Publicly accessible
cc-by-sa-4.0
100K<n<1M
1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours. - 100+ real-world scenarios across 5 target domains. - 200+ types of tasks: - 87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc. Your browser does not support the video tag. Your browser does not support the video tag. Your browser does not support the video tag. - [2025/3/1] AgiBot World Beta released. - [ ] AgiBot World Colosseum:Comprehensive platform (expected release date: 2025) - [ ] 2025 AgiBot World Challenge (expected release date: 2025) To download the full dataset, you can use the following code. If you encounter any issues, please refer to the official Hugging Face documentation. If…
Access requested at publisher
n>1T
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973 Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each part. Pixabay has also been compressed into multiple parts. After extracting them, all videos should be placed into a single folder. For SAM data, please download from the official link. After downloading 1000 compressed files, extract all the…
Publicly accessible
mit
Publicly accessible
Dataset · Speech recognition
Google
Universal Representations of Speech](https://arxiv.org/abs/2205.12446) Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven geographical areas: The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The dataset can be downloaded and prepared…
Publicly accessible
cc-by-4.0
10K<n<100K
Dataset · Video classification
Ropedia
Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified. Interactive Intelligence from Human Xperience Xperience-10M Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial
Access requested at publisher
other
1M<n<10M
Publicly accessible
MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 103.8 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with multiple vision-language models, and is further augmented with synthetically generated samples. Each image is released with pre-computed embeddings, structured annotations and pre-encoded VAE latents to accelerate downstream use. A 4B-parameter…
Publicly accessible
apache-2.0
100M<n<1B
SWE-bench Lite is subset of SWE-bench, a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of SWE-bench: Can Language Models Resolve Real-World GitHub Issues? This dataset only contains the problemstatement (i.e. issue text) and the basecommit which can represents the state of the codebase before the issue has been resolved. If you want to run inference using the "Oracle" or BM25 retrieval settings mentioned in the paper, consider the following datasets.…
Publicly accessible
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2024. AIME is a prestigious high school mathematics competition known for its challenging mathematical problems. Each record contains the following fields: - ID: Problem identifier (e.g., "2024-I-1" represents Problem 1 from 2024 Contest I) - Problem: Problem statement - Solution: Detailed solution process - Answer: Final numerical answer This dataset is primarily used for: 1. Evaluating Large Language Models' (LLMs) mathematical reasoning capabilities 2. Testing models' problem-solving abilities on complex mathematical problems 3. Researching AI performance on structured mathematical tasks - Covers…
Publicly accessible
mit
n<1K
Publicly accessible
English | Ultra-FineWeb is a large-scale, high-quality, and efficiently-filtered dataset. We use the proposed efficient verification-based high-quality filtering pipeline to the FineWeb and Chinese FineWeb datasets (source data from Chinese FineWeb-edu-v2, which includes IndustryCorpus2, MiChao, WuDao, SkyPile, WanJuan, ChineseWebText, TeleChat, and CCI3), resulting in the creation of higher-quality Ultra-FineWeb-en with approximately 1T tokens, and Ultra-FineWeb-zh datasets with approximately 120B tokens, collectively referred to as Ultra-FineWeb. Ultra-FineWeb serves as a core pre-training web dataset for the MiniCPM4 Series and MiniCPM5 Series models. - Ultra-FineWeb-L1: L1 filtered data…
Publicly accessible
apache-2.0
n>1T
Publicly accessible
If you use the AIME25 dataset in your research, please consider citing it as follows
Publicly accessible
apache-2.0
Publicly accessible
AI-MO Olympiad Reference Dataset This dataset contains a structured collection of Olympiad problems and their solutions, organized by competition. Contains high quality data, prioritizing "official" solutions to problems. Structure / # Problems and solutions from the International Mathematical Olympiad ├── raw/ # Raw problem/solution statements (.pdf) │ ├── file1.pdf │ ├── file2.pdf ├── downloadscript/ # the scripts used to
Publicly accessible
C
Dataset · Text generation
Chi
Publicly accessible
mit
n<1K
Publicly accessible
Publicly accessible
Falcon RefinedWeb is a massive English web dataset built by TII and released under an ODC-By 1.0 license. See the paper on arXiv for more details. RefinedWeb is built through stringent filtering and large-scale deduplication of CommonCrawl; we found models trained on RefinedWeb to achieve performance in-line or better than models trained on curated datasets, while only relying on web data. RefinedWeb is also "multimodal-friendly": it contains links and alt texts for images in processed samples. This public extract should contain 500-650GT depending on the tokenizer you use, and can be enhanced with the curated corpora of your choosing. This public extract is about ~500GB to download…
Publicly accessible
odc-by
100B<n<1T
Publicly accessible
Dataset · Question answering
NVIDIA
Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. The high-quality trajectories in Terminal-Corpus enable models of various sizes to achieve performance that rivals or exceeds much larger frontier models on the Terminal-Bench 2.0 benchmark. Training on Terminal-Corpus yields substantial gains across the Qwen3 model family: The Nemotron-Terminal-32B (27.4%) outperforms the 480B-parameter Qwen3-Coder (23.9%) and…
Publicly accessible
cc-by-4.0
100K<n<1M
Publicly accessible