The data construction workflow can be summarized as follows: 1. Deduplicate: The FineWeb dataset is deduplicated using exact deduplication and MinHash techniques to remove redundant data. 2. URL Labeling: Root URLs from FineWeb are counted, and the top 1 million URLs are labeled using GPT-4. This step generates DoI (Domain-of-Interest) Coarse-Grained URLs and DoNI (Domain-of-Non-Interest) Coarse-Grained URLs as seed data sources. 3. Coarse Recall: a. Based on the labeled root URLs, data is sampled for each domain. b. The sampled data is labeled using Qwen2-7B-Instruct, producing 500K DoI Positive Data and 500K DoI Negative Data (note that for N>1 iterations, each 500K samples are composed…
Publicly accessible
apache-2.0
n>1T
A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents" This dataset contains around 200M samples in PIN format, with around 312 TB storage. News [ 2025.09.22 ]!NEW! We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it. [ 2024.12.06 ]!NEW! We have updated the quality signals, enabling a swift assessment of whether a sample meets the required specifications based on our quality indicators. Further detailed descriptions will be provided in the forthcoming formal publication. (Aside from the Chinese-Markdown subset, there are unresolved issues that are currently being addressed.) This dataset…
Publicly accessible
apache-2.0
100M<n<1B