The data construction workflow can be summarized as follows: 1. Deduplicate: The FineWeb dataset is deduplicated using exact deduplication and MinHash techniques to remove redundant data. 2. URL Labeling: Root URLs from FineWeb are counted, and the top 1 million URLs are labeled using GPT-4. This step generates DoI (Domain-of-Interest) Coarse-Grained URLs and DoNI (Domain-of-Non-Interest) Coarse-Grained URLs as seed data sources. 3. Coarse Recall: a. Based on the labeled root URLs, data is sampled for each domain. b. The sampled data is labeled using Qwen2-7B-Instruct, producing 500K DoI Positive Data and 500K DoI Negative Data (note that for N>1 iterations, each 500K samples are composed…
Publicly accessible
apache-2.0
n>1T
HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/.
Publicly accessible
cc-by-nc-sa-4.0
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far larger vocabulary and retains the original case, punctuation and numbers - all of which are removed in PTB. As it is composed of full articles, the dataset is well suited for models that can take advantage of long term dependencies. Each subset comes in two different variants…
Publicly accessible
cc-by-sa-3.0
1M<n<10M
これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。 しかし、旧 API が廃止されてしまったため過去ログを API 経由で取得することができなくなり、またアーカイブされた過去ログから見たい範囲のログを探す場合も、アーカイブのサイズが合計約 150GB もあることから、とても以前のように手軽に過去ログに触れることはできなくなってしまいました。 このデータセットでは、ニコニコ実況のすべての過去ログを後世に残すべく、Nekopanda 氏が配布されていた旧ニコニコ実況の 2020/12/15 までのすべての過去ログに加え、コミュニティでの実況番組も含めた新ニコニコ実況、さらに 2024/06/10 からは実況用代替コメントサーバーである NX-Jikkyo の当日分の過去ログを5分に1回収集し、随時反映しています。
Publicly accessible
mit
和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报
Publicly accessible
n>1T
Publicly accessible
UniOcc: A Unified Benchmark for Occupancy Forecasting and Prediction in researchers, have you ever been bothered by the fact that popular datasets all have their different formats, and standardizing them is a pain? Have you ever been frustrated by the difficulty of just understanding the file semantics? This challenge is even worse in the occupancy domain. But, UniOcc is here to help. UniOcc is a unified
Publicly accessible
mit
Dataset · Text generation
Ai2
A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: - en.noclean: 2.3TB - en.noblocklist: 380GB - realnewslike: 15GB - multilingual (mC4): 9.7TB (108 subsets, one per language) The en.noblocklist variant is exactly the same as the en variant, except we turned off the so-called "badwords filter", which removes all documents that contain words from the lists at…
Publicly accessible
odc-by
n<1K
We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks. Users can download a specific subset of data by specifying the dataset name. 1. Option 1: with huggingface-cli Replace gr1armsonly.CanSort/ with the dataset name you want to download. 2. Option 2: Github LFS Replace the "gr1armswaist.CupToDrawer" with the intended dataset folder
Publicly accessible
cc-by-4.0
Publicly accessible
mit
This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: metadata subset containing extra columns used in dataset construction: - problem - groundtruthsolution - deepseekreasoning - deepseeksolution - domain - source - testcases (code only) - startercode(code only) The numbers reported in the tables below are evaluated with our open-source tool Evalchemy. We are fully open-source. Our model weights, datasets, data…
Publicly accessible
Dataset · Text generation
OpenAI
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. - These problems take between 2 and 8 steps to solve. - Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to reach the final answer. - A bright middle school student should be able to solve every problem: from the paper, "Problems require no concepts beyond the level of early Algebra, and the vast majority of problems can be solved without explicitly defining a variable."…
Publicly accessible
mit
1K<n<10K
This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently accessible for research and development purposes. The files are organized by application categories, mirroring the structure of the original OSWorld evaluation examples: Each application folder contains subfolders named after specific evaluation scenarios, with…
Publicly accessible
apache-2.0
Publicly accessible
Publicly accessible
Based on Hersbach et al. 2020 with data exposed through Copernicus C3S API 26 variable subset of data as described in Table 3 of Bonev et al. 2023. Each file contains all 26 variables sampled every 6 hours (starting with 00:00:00) for an entire month in a given year.
Publicly accessible
Dataset · Question answering
Ai2
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences relevant to the task, and an implementation of three neural baseline models for this dataset. We pose ARC as a challenge to the community. An example of 'train' looks as follows. An example of 'train' looks as follows. The data fields are the same among all splits. - id: a string…
Publicly accessible
cc-by-sa-4.0
1K<n<10K
Production-ready pipeline (Python package videovec2wav2tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio Speech recognition — faster-whisper, CPU & CUDA, automatic language detection, word-level timestamps. Segmentation — cut audio by transcript timestamps into dataset/audio/000001.wav …. Dataset generation — metadata.csv, dataset.jsonl, ttsmetadata.csv. Feature extraction (optional) — streaming features/train.bin + train.dat with float32 samples, mel spectrograms, duration and sample rate. Statistics…
Publicly accessible
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: A manually-curated evaluation dataset for fine-grained analysis of system performance on a broad range of linguistic phenomena. This dataset evaluates sentence understanding through Natural Language Inference (NLI) problems. Use a model trained on MulitNLI to produce predictions for this dataset. The Corpus of Linguistic Acceptability consists of English acceptability judgments drawn from…
Publicly accessible
other
10K<n<100K
Publicly accessible
apache-2.0
Measuring Massive Multitask Language Understanding by Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt (ICLR 2021). This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn. This covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability. A complete list of tasks: ['abstractalgebra', 'anatomy', 'astronomy'…
Publicly accessible
mit
10K<n<100K
This is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider repo.
Publicly accessible
mit
Publicly accessible
If you find LLaVA-One-Vision-1.5-Mid-Training-85M useful in your research, please consider to cite the following related papers
Publicly accessible
apache-2.0