Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump, and you can load them per date and language this way: Click the Nomic Atlas map below to visualize the 6.4 million samples in the 20231101.en split. The dataset is generally used for Language Modeling. You can find the list of languages here…
Organization
Wikimedia
wikimedia
Models in Library0
Datasets in Library1
Models on Hugging Face—
Followers545