A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences relevant to the task, and an implementation of three neural baseline models for this dataset. We pose ARC as a challenge to the community. An example of 'train' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - id: a string…
SAVRN Model Hub · Datasets by Task
Question Answering Datasets
25 open-weight question answering datasets in the SAVRN Model Hub, with Ai2, Pranav R and NVIDIA publishing the most.
SAVRN's Take
Twenty-five datasets sit under question answering, and the ones people pull most are built to score a model, not to train one. ARC holds 7,787 grade-school science questions in a Challenge Set and an Easy Set and leads at 879,812 downloads; MMLU, from the Center for AI Safety, sits at 779,659 and spreads multiple-choice questions across branches of knowledge. OpenBookQA asks for multi-step reasoning against an open book of facts, SciQ carries 13,679 science exam questions with four options each, and SQuAD tests reading comprehension, where the answer is a span of a Wikipedia passage or the question cannot be answered.
Storage is not the issue: ARC is 1.2 MB, SQuAD 16 MB, and only MMLU at 270 MB registers on disk, so the cost of an evaluation is the inference through every question, a token bill rather than a storage purchase. Licenses are the thing to read. Seven of the 25 carry cc-by-sa-4.0, ARC and SQuAD among them, five are MIT including MMLU, five are cc-by-4.0, and three are unknown, OpenBookQA being one. SciQ is the lone cc-by-nc-3.0 entry, and that nc is the clause to check before it goes near a paid product. Ai2 publishes five of the 25, more than anyone else. Take ARC and MMLU first; the download counts say so.
Most Downloaded
| Dataset | Publisher | License | Monthly downloads |
|---|---|---|---|
| ai2_arc | Ai2 | cc-by-sa-4.0 | 879.8k |
| mmlu | Center for AI Safety | mit | 779.7k |
| openbookqa | Ai2 | unknown | 487.6k |
| sciq | Ai2 | cc-by-nc-3.0 | 392.5k |
| commonsense_qa | Tel Aviv University | mit | 308.4k |
| squad | Pranav R | cc-by-sa-4.0 | 263.8k |
| MMLU-Pro | TIGER-Lab | mit | 246k |
| medmcqa | Open Life Science AI | apache-2.0 | 204.5k |
| trivia_qa | Mandar Joshi | unknown | 156.7k |
| OpenMathInstruct-2 | NVIDIA | cc-by-4.0 | 146k |
Licenses
| License | Datasets | Commercial use |
|---|---|---|
| cc-by-sa-4.0 | 7 | Yes |
| mit | 5 | Yes |
| cc-by-4.0 | 5 | Yes |
| unknown | 3 | Read the license |
| other | 1 | Read the license |
| not stated | 1 | Not stated |
| cc-by-sa-3.0 | 1 | Read the license |
| cc-by-nc-3.0 | 1 | Read the license |
Who Publishes Them
| Publisher | Datasets |
|---|---|
| Ai2 | 5 |
| Pranav R | 2 |
| NVIDIA | 2 |
| Ikedachin | 2 |
| Yonatan Bisk | 1 |
| TIGER-Lab | 1 |
All 25 Datasets
Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021). This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability. A complete list of tasks: ['abstractalgebra', 'anatomy', 'astronomy'…
OpenBookQA aims to promote research in advanced question-answering, probing a deeper understanding of both the topic (with salient facts summarized as an open book, also provided with the dataset) and the language it is expressed in. In particular, it contains questions that require multi-step reasoning, use of additional common and commonsense knowledge, and rich text comprehension. OpenBookQA is a new kind of question-answering dataset modeled after open book exams for assessing human understanding of a subject. An example of 'train' looks as follows: An example of 'train' looks as follows: The data fields are the same among all splits. - id: a string feature. - questionstem: a string…
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others. The questions are in multiple-choice format with 4 answer options each. For the majority of the questions, an additional paragraph with supporting evidence for the correct answer is provided. An example of 'train' looks as follows. The data fields are the same among all splits. - question: a string feature. - distractor3: a string feature. - distractor1: a string feature. - distractor2: a string feature. - correctanswer: a string feature. - support: a string feature. The dataset is licensed under the Creative Commons Attribution-NonCommercial 3.0 Unported License. Thanks…
CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers. It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details. The dataset is in English (en). An example of 'train' looks as follows: The data fields are the same among all splits. - id (str): Unique ID. - question: a string feature. - questionconcept (str): ConceptNet concept associated to the question. - choices: a dictionary feature…
Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable. SQuAD 1.1 contains 100,000+ question-answer pairs on 500+ articles. Question Answering. English (en). An example of 'train' looks as follows. The data fields are the same among all splits. - id: a string feature. - title: a string feature. - context: a string feature. - question: a string feature. - answers: a dictionary feature containing: - text: a string feature. - answerstart: a int32…
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. - \[2026.03.11\] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among others. Stay tuned for updated rankings and analysis. - \[2026.01.18\] Fixed leading space issue in answer options (affected chemistry, physics, and other STEM subsets). This formatting inconsistency could have been exploited as a shortcut. Thanks to @giffmana and @fujikanaeda for identifying…
MedMCQA is a large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions. MedMCQA has more than 194k high-quality AIIMS & NEET PG entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average token length of 12.77 and high topical diversity. Each sample contains a question, correct answer(s), and other options which require a deeper language understanding as it tests the 10+ reasoning abilities of a model across a wide range of medical subjects & topics. A detailed explanation of the solution, along with the above information, is provided in this study. MedMCQA provides an…
TriviaqQA is a reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaqQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence documents, six per question on average, that provide high quality distant supervision for answering the questions. English. An example of 'train' looks as follows. An example of 'train' looks as follows. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - question: a string feature. - questionid: a string feature. - questionsource: a string feature. - entitypages: a dictionary feature containing…
OpenMathInstruct-2 is a math instruction tuning dataset with 14M problem-solution pairs generated using the Llama3.1-405B-Instruct model. The training set problems of GSM8K and MATH are used for constructing the dataset in the following ways: OpenMathInstruct-2 dataset contains the following fields: - generatedsolution: Synthetically generated solution. - expectedanswer: For problems in the training set, it is the ground-truth answer provided in the datasets. For augmented problems, it is the majority-voting answer. - problemsource: Whether the problem is taken directly from GSM8K or MATH or is an augmented version derived from either dataset. We also release the 1M, 2M, and 5M…
QuaRTz is a crowdsourced dataset of 3864 multiple-choice questions about open domain qualitative relationships. Each question is paired with one of 405 different background sentences (sometimes short paragraphs). The QuaRTz dataset V1 contains 3864 questions about open domain qualitative relationships. Each question is paired with one of 405 different background sentences (sometimes short paragraphs). The dataset is split into train (2696), dev (384) and test (784). A background sentence will only appear in a single split. An example of 'train' looks as follows. The data fields are the same among all splits. - id: a string feature. - question: a string feature. - choices: a dictionary…
QASC is a question-answering dataset with a focus on sentence composition. It consists of 9,980 8-way multiple-choice questions about grade school science (8,134 train, 926 dev, 920 test), and comes with a corpus of 17M sentences. An example of 'validation' looks as follows. The data fields are the same among all splits. - id: a string feature. - question: a string feature. - choices: a dictionary feature containing: - text: a string feature. - label: a string feature. - answerKey: a string feature. - fact1: a string feature. - fact2: a string feature. - combinedfact: a string feature. - formattedquestion: a string feature. The dataset is released under CC BY 4.0 license. Thanks to…
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model training corpora. We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who…
To apply eyeshadow without a brush, should I use a cotton swab or a toothpick? Questions requiring this kind of physical commonsense pose a challenge to state-of-the-art natural language understanding systems. The PIQA dataset introduces the task of physical commonsense reasoning and a corresponding benchmark dataset Physical Interaction: Question Answering or PIQA. Physical commonsense knowledge is a major challenge on the road to true AI-completeness, including robots that interact with the world and understand natural language. PIQA focuses on everyday situations with a preference for atypical solutions. The dataset is inspired by instructables.com, which provides users with instructions…
HotpotQA is a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) the questions require finding and reasoning over multiple supporting documents to answer; (2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas; (3) we provide sentence-level supporting facts required for reasoning, allowingQA systems to reason with strong supervision and explain the predictions; (4) we offer a new type of factoid comparison questions to test QA systems’ ability to extract relevant facts and perform necessary comparison. An example of 'validation' looks as follows. An example of 'train' looks as follows. The data fields…
Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. The high-quality trajectories in Terminal-Corpus enable models of various sizes to achieve performance that rivals or exceeds much larger frontier models on the Terminal-Bench 2.0 benchmark. Training on Terminal-Corpus yields substantial gains across the Qwen3 model family: The Nemotron-Terminal-32B (27.4%) outperforms the 480B-parameter Qwen3-Coder (23.9%) and…
Stanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable. SQuAD 2.0 combines the 100,000 questions in SQuAD1.1 with over 50,000 unanswerable questions written adversarially by crowdworkers to look similar to answerable ones. To do well on SQuAD2.0, systems must not only answer questions when possible, but also determine when no answer is supported by the paragraph and abstain from answering. Question Answering. English (en). An example of…
The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial fibrillation after coronary artery bypass grafting?) using the corresponding abstracts. The official leaderboard is available at: https://pubmedqa.github.io/. 500 questions in the pqalabeled are used as the test set. They can be found at https://github.com/pubmedqa/pubmedqa. English Thanks to @tuner007 for adding this dataset.
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set. - Website: The landing page for TOFU - arXiv Paper: Detailed information about the TOFU dataset and its significance in unlearning tasks. - GitHub Repository: Access the source code, fine-tuning scripts, and additional resources for the TOFU dataset. - Dataset on Hugging Face: Direct link to download the…
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. This dataset enables the evaluation of mono- and multi-lingual models in high-, medium-, and low-resource languages. Each question has four multiple-choice answers and is linked to a short passage from the FLORES-200 dataset. The human annotation procedure was carefully curated to create questions that discriminate between different levels of generalizable language comprehension and is reinforced by extensive quality checks. While all questions directly relate to the passage, the English dataset on its own proves difficult enough to challenge state-of-the-art language models. Being…
The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. An example of 'train' looks as follows. This is a toy example. The data fields are the same among all splits. - id: a string feature. - document a dictionary feature containing: - title: a string feature. - url: a string feature. - html: a string feature. - tokens: a dictionary feature containing: - token: a string…
A community-driven dataset for the Pa'O (blk) language, designed to support artificial intelligence, natural language processing (NLP), large language models (LLMs), conversational AI, and language technology research. The dataset focuses on three core Pa'O language resources: The primary goal is to build high-quality, reusable Pa'O language data for AI research, language technology, digital language preservation, and future Pa'O-capable AI systems. pao-ai-qa-conversation-instruction/ ├── data/ │ │ └── qa.jsonl │ ├── conversation/ │ │ └── conversations.jsonl │ └── instruction/ │ └── instructions.jsonl ├── README.md Question-and-answer pairs for Pa'O language understanding and…
日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。 This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is regenerated with context from the source articles. 本カードは LLM-jp 4 用のチャットテンプレートに合わせたデータ形式を説明します。思考文の生成モデルは、両形式とも Qwen3.8-27B-NVFP4 です。Qwen3.8版とLLM-jp 4版は、同じ質問・回答・生成済み思考文を共有し、メッセージのフィールド、effortラベル、トークン計数用トークナイザーが異なります。 This card describes the data format adapted to the LLM-jp 4 chat template. Both versions use…
日本語・今治弁のQAを用いて、reasoning effort に応じた思考文の生成を学習するための教師ありファインチューニング(SFT)用データセットです。Imabari Wiki QA v4 Validated の質問と回答を保持し、元記事の文脈を参照して思考文を再生成しています。 This dataset supports supervised fine-tuning (SFT) of reasoning-effort-conditioned explanations using Japanese QA with Imabari dialect expressions. Questions and answers from Imabari Wiki QA v4 Validated are preserved, while reasoning text is regenerated with context from the source articles. 本カードは Qwen3.8 用のチャットテンプレートに合わせたデータ形式を説明します。思考文の生成モデルは、両形式とも Qwen3.8-27B-NVFP4 です。Qwen3.8版とLLM-jp 4版は、同じ質問・回答・生成済み思考文を共有し、メッセージのフィールド、effortラベル、トークン計数用トークナイザーが異なります。 This card describes the data format adapted to the Qwen3.8 chat template. Both versions use…
A long-context retrieval and question-answering benchmark with 1,143,371 documents and 10,000 questions, including 50 questions about 0G. It extends the ms100M bank from MSA-RAG-BENCHMARKS with 20 million additional text tokens and 655 new questions grounded in the added documents. 120M is the benchmark's nominal size. The actual corpus contains 125,708,601 raw tokens under the MSA-4B tokenizer: 105,708,601 original tokens plus exactly 20,000,000 added tokens. Counts exclude prompts, wrappers, and special tokens. This is a derived benchmark, not an official Microsoft MS MARCO release. The original documents occupy positions 0:961686; the original questions occupy positions 0:9345. Their…


