source languages: nl; target languages: en; OPUS readme: nl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
Search public pages, research tools, and SAVRN solutions.
At the University of Helsinki, we focus on: - NLP for morphologically-rich languages - Cross-lingual NLP - NLP in the humanities
source languages: nl; target languages: en; OPUS readme: nl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: fr; target languages: en; OPUS readme: fr-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: en; target languages: ru; OPUS readme: en-ru; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: en-de
hfname: kor-eng - sourcelanguages: kor - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/kor-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'korHani', 'korHang', 'korLatn', 'kor'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.test.txt - srcalpha3: kor - tgtalpha3: eng - shortpair: ko-en - chrF2score: 0.588 - brevitypenalty: 0.9590000000000001 - reflen: 17711.0 - srcname: Korean - tgtname: English - traindate…
source languages: de; target languages: en; OPUS readme: de-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: ru-en
hfname: spa-eng - sourcelanguages: spa - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/spa-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'spa'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/spa-eng/opus-2020-08-18.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/spa-eng/opus-2020-08-18.test.txt - srcalpha3: spa - tgtalpha3: eng - shortpair: es-en - chrF2score: 0.7390000000000001 - brevitypenalty: 0.9740000000000001 - reflen: 79376.0 - srcname: Spanish - tgtname: English - traindate: 2020-08-18 00:00:00…
source languages: en; target languages: fr; OPUS readme: en-fr; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: zho-eng helsinkigitsha: 480fcbe0ee1bf4774bcbe6226ad9f58e63f6c535 transformersgitsha: 2207e5d8cb224e954a7cba69fa4ac2309e9ff30b portmachine: brutasse porttime: 2020-08-21-14:41 srcmultilingual: False tgtmultilingual: False reflen: 82826.0 brevitypenalty…
hfname: eng-spa - sourcelanguages: eng - targetlanguages: spa - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eng-spa/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eng'} - tgtconstituents: {'spa'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-spa/opus-2020-08-18.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-spa/opus-2020-08-18.test.txt - srcalpha3: eng - tgtalpha3: spa - shortpair: en-es - chrF2score: 0.721 - brevitypenalty: 0.978 - reflen: 77311.0 - srcname: English - tgtname: Spanish - traindate: 2020-08-18 00:00:00 - srcalpha2: en - tgtalpha2…
Neural machine translation model for translating from English (en) to Turkish (tr). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…
Neural machine translation model for translating from Turkish (tr) to English (en). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…
source languages: ar; target languages: en; OPUS readme: ar-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: it; target languages: en; OPUS readme: it-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
Neural machine translation model for translating from Korean (ko) to English (en). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. - More information about released models for this language pair: OPUS-MT kor-eng README - Tatoeba Translation…
a sentence initial language token is required in the form of >>id<< (id = valid target language ID) - hfname: eng-zho - sourcelanguages: eng - targetlanguages: zho - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eng-zho/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eng'} - tgtconstituents: {'cmnHans', 'nan', 'nanHani', 'gan', 'yue', 'cmnKana', 'yueHani', 'wuuBopo', 'cmnLatn', 'yueHira', 'cmnHani', 'cjyHans', 'cmn', 'lzhHang', 'lzhHira', 'cmnHant', 'lzhBopo', 'zho', 'zhoHans', 'zhoHant', 'lzhHani', 'yueHang', 'wuu', 'yueKana', 'wuuLatn', 'yueBopo', 'cjyHant', 'yueHans', 'lzh', 'cmnHira', 'lzhYiii', 'lzhHans', 'cmnBopo', 'cmnHang'…
source languages: fr,frBE,frCA,frFR,wa,frp,oc,ca,rm,lld,fur,lij,lmo,es,esAR,esCL,esCO,esCR,esDO,esEC,esES,esGT,esHN,esMX,esNI,esPA,esPE,esPR,esSV,esUY,esVE,pt,ptbr,ptBR,ptPT,gl,lad,an,mwl,it,itIT,co,nap,scn,vec,sc,ro,la; target languages: en; OPUS readme: fr+frBE+frCA+frFR+wa+frp+oc+ca+rm+lld+fur+lij+lmo+es+esAR+esCL+esCO+esCR+esDO+esEC+esES+esGT+esHN+esMX+esNI+esPA+esPE+esPR+esSV+esUY+esVE+pt+ptbr+ptBR+ptPT+gl+lad+an+mwl+it+itIT+co+nap+scn+vec+sc+ro+la-en; dataset: opus; model: transformer; pre-processing: normalization + SentencePiece.
hfname: mul-eng - sourcelanguages: mul - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/mul-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'sjnLatn', 'cat', 'nan', 'spa', 'ileLatn', 'pap', 'mwl', 'uzbLatn', 'mww', 'hil', 'lij', 'avkLatn', 'ladLatn', 'latLatn', 'bosLatn', 'oss', 'epo', 'ron', 'fry', 'cym', 'toiLatn', 'awa', 'swg', 'zsmLatn', 'zhoHant', 'gcfLatn', 'uzbCyrl', 'isl', 'lfnLatn', 'shsLatn', 'novLatn', 'bho', 'ltz', 'lzh', 'kurLatn', 'sun', 'arg', 'pesThaa', 'sqi', 'uigArab', 'csbLatn', 'fra', 'hat', 'livLatn', 'nonLatn', 'sco', 'cmnHans', 'pnb', 'roh', 'chv', 'ibo', 'bulLatn', 'amh', 'lfnCyrl'…
source languages: fr; target languages: es; OPUS readme: fr-es; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: it; target languages: es; OPUS readme: it-es; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: tr; target languages: en; OPUS readme: tr-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
hfname: eus-spa - sourcelanguages: eus - targetlanguages: spa - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eus-spa/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eus'} - tgtconstituents: {'spa'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/eus-spa/opus-2020-06-17.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/eus-spa/opus-2020-06-17.test.txt - srcalpha3: eus - tgtalpha3: spa - shortpair: eu-es - chrF2score: 0.6729999999999999 - brevitypenalty: 0.9640000000000001 - reflen: 12469.0 - srcname: Basque - tgtname: Spanish - traindate: 2020-06-17 - srcalpha2…
Neural machine translation model for translating from Arabic (ar) to English (en). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…
source languages: da; target languages: en; OPUS readme: da-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: pl; target languages: en; OPUS readme: pl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
hfname: fin-eng - sourcelanguages: fin - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/fin-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'fin'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/fin-eng/opus-2020-08-05.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/fin-eng/opus-2020-08-05.test.txt - srcalpha3: fin - tgtalpha3: eng - shortpair: fi-en - chrF2score: 0.6970000000000001 - brevitypenalty: 0.99 - reflen: 74651.0 - srcname: Finnish - tgtname: English - traindate: 2020-08-05 - srcalpha2: fi…
source languages: ja; target languages: en; OPUS readme: ja-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: bg; target languages: en; OPUS readme: bg-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: nl; target languages: fr; OPUS readme: nl-fr; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: fr; target languages: de; OPUS readme: fr-de; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: sv; target languages: en; OPUS readme: sv-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
source languages: de; target languages: fr; OPUS readme: de-fr; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.
a sentence initial language token is required in the form of >>id<< (id = valid target language ID) - hfname: eng-ara - sourcelanguages: eng - targetlanguages: ara - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eng-ara/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eng'} - tgtconstituents: {'apc', 'ara', 'arqLatn', 'arq', 'afb', 'araLatn', 'apcLatn', 'arz'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-ara/opus-2020-07-03.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-ara/opus-2020-07-03.test.txt - srcalpha3: eng - tgtalpha3: ara - shortpair: en-ar…
Neural machine translation model for translating from English (en) to Bulgarian (bg). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…
fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations are added for Czech (ces), Ukrainian (ukr) and Finnish (fin). For Czech and Ukrainian, this release doubles the data and for Finnish, we include translations for the entire fineweb-edu data set with its 350B token release. More information about how…