SAVRN
Search Contact SAVRN

Organization

HPLT

HPLT

Web as a corpus, Large Language Models, Machine Translation, Language Technologies, Natural Language Processing, Internet Archive, CommonCrawl

Models in Library0
Datasets in Library1
Models on Hugging Face678
Followers101

Datasets

Dataset · Fill mask

HPLT2.0_cleaned

HPLT

We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. This is the variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here. The original JSONL files (which take ~4x fewer disk space than this HF version) and the larger non-cleaned version can be found at https://hplt-project.org/datasets/v2.0. We conducted the FineWeb-style…

Publicly accessible cc0-1.0 n>1T