SAVRN
Search Contact SAVRN

fineweb-edu-translated · Dataset Card

fineweb-edu-translated: Dataset Card

Written by Helsinki-NLP Research Group, published under odc-by, revision d82b075a4a81, read 2026-09-18. Shown as written; SAVRN's own facts about this dataset are on its page.

fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages.

In the v1.1 release, additional translations are added for Czech (ces), Ukrainian (ukr) and Finnish (fin). For Czech and Ukrainian, this release doubles the data and for Finnish, we include translations for the entire fineweb-edu data set with its 350B token release.

More information about how the data has been produced can be found on https://github.com/Helsinki-NLP/translate-fineweb. The corpus is also available with aligned sentences in synOPUS: https://opus.nlpl.eu/synthetic/transweb-edu.php

Supported Languages

Langid Language
bos Bosnian
bul Bulgarian
cat Catalan
ces Czech
dan Danish
deu German
ell Modern Greek
eng English
est Estonian
eus Basque
fin Finnish
fra French
gle Irish
glg Galician
hrv Croatian
hun Hungarian
isl Icelandic
ita Italian
kat Georgian
lav Latvian
lit Lithuanian
mkd Macedonian
mlt Maltese
nld Dutch
nno Norwegian Nynorsk
nob Norwegian Bokmål
pol Polish
por Portuguese
ron Romanian
slk Slovak
slv Slovenian
spa Spanish
sqi Albanian
srp_Cyrl Serbian (cyrillic script)
swe Swedish
tur Turkish
ukr Ukrainian

Citation Information

Please acknowledge the source when using the data and, please, cite the following article if you use any part of this corpus in your own work:

@article{tiedemann2023democratizing,
  title={Democratizing neural machine translation with {OPUS-MT}},
  author={Tiedemann, J{\"o}rg and Aulamo, Mikko and Bakshandaeva, Daria and Boggia, Michele and Gr{\"o}nroos, Stig-Arne and Nieminen, Tommi and Raganato, Alessandro and Scherrer, Yves and Vazquez, Raul and Virpioja, Sami},
  journal={Language Resources and Evaluation},
  number={58},
  pages={713--755},
  year={2023},
  publisher={Springer Nature},
  issn={1574-0218},
  doi={10.1007/s10579-023-09704-w}
}

Translation Models

The following translation models have been used for creating the data:

Langid translation model
bos Tatoeba-MT-models/eng-hbs/opus+bt-2021-04-20 download
bul Tatoeba-MT-models/eng-bul/opusTCv20210807+bt_transformer-big_2022-02-25 download
cat Tatoeba-MT-models/deu+eng+fra+por+spa-roa/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
ces Tatoeba-MT-models/eng-ces+slk/opusTCv20210807+bt_transformer-big_2022-03-13 download
dan Tatoeba-MT-models/eng-gmq/opusTCv20210807+bt_transformer-big_2022-03-17 download
deu Tatoeba-MT-models/eng-deu/opusTCv20210807+bt-2021-12-08 download
ell Tatoeba-MT-models/eng-ell/opusTCv20210807+bt_transformer-big_2022-03-13 download
est Tatoeba-MT-models/deu+eng+fra+por+spa-urj/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
eus HPLT-MT-models/en-eu/translate-en-eu-v1.0-hplt_opus download
fin Tatoeba-MT-models/eng-fin/opusTCv20210807+bt_transformer-big_2022-03-09 download
fra Tatoeba-MT-models/eng-fra/opusTCv20210807+bt_transformer-big_2022-03-09 download
gle HPLT-MT-models/en-ga/translate-en-ga-v1.0-hplt_opus download
glg Tatoeba-MT-models/deu+eng+fra+por+spa-itc/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
hrv Tatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
hun Tatoeba-MT-models/eng-hun/opusTCv20210807+bt_transformer-big_2022-02-25 download
isl HPLT-MT-models/en-is/translate-en-is-v1.0-hplt_opus download
ita Tatoeba-MT-models/deu+eng+fra+por+spa-itc/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
kat Tatoeba-MT-models/deu+eng+fra+por+spa-cau/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
lav Tatoeba-MT-models/eng-lav/opusTCv20210807+bt_transformer-big_2022-03-13 download
lit Tatoeba-MT-models/deu+eng+fra+por+spa-bat/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
mkd Tatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
mlt HPLT-MT-models/en-mt/translate-en-mt-v1.0-hplt_opus download
nld Tatoeba-MT-models/deu+eng+fra+por+spa-gmw/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
nno Tatoeba-MT-models/deu+eng+fra+por+spa-gmq/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
nob Tatoeba-MT-models/deu+eng+fra+por+spa-gmq/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
pol Tatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
por Tatoeba-MT-models/gem-fra+ita+por+spa/opusTCv20230926max50+bt+jhubc_transformer-big_2024-08-17 download
ron Tatoeba-MT-models/deu+eng+fra+por+spa-itc/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
slk Tatoeba-MT-models/eng-ces+slk/opusTCv20210807+bt_transformer-big_2022-03-13 download
slv Tatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
spa Tatoeba-MT-models/eng-spa/opusTCv20210807+bt_transformer-big_2022-03-13 download
sqi HPLT-MT-models/en-sq/translate-en-sq-v1.0-hplt_opus download
srp_Cyrl Tatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30 download
swe Tatoeba-MT-models/eng-gmq/opusTCv20210807+bt_transformer-big_2022-03-17 download
tur Tatoeba-MT-models/eng-tur/opusTCv20210807+bt_transformer-big_2022-02-25 download
ukr Tatoeba-MT-models/eng-zle/opusTCv20210807+bt_transformer-big_2022-03-13 download

Acknowledgements

None of this would be possible without the enormous work done by Common Crawl providing the essential data that most open datasets for language modeling are based on. Furthermore, we are also grateful for the data preparation work done by Hugging Face and the community on top of the crawled data from Common Crawl published under the label of fineweb-edu. Important for this work is also the availability of parallel data through OPUS and the public translation models based on that data. Furthermore, this project was supported by the European Union's Horizon Europe research and innovation programme through the HPLT project under grant agreement No 101070350. Finally, we also want to acknowledge the computational resource made available from the Finnish national allocation for the LUMI supercomputer (https://www.lumi-supercomputer.eu) through the extreme scale project MaMuLaM: Massively Multilingual Language Models. None of the translated data would exist without this infrastructure and the compute.