SAVRN
Search Contact SAVRN

Organization

Helsinki-NLP Research Group

Helsinki-NLP · github.com

At the University of Helsinki, we focus on: - NLP for morphologically-rich languages - Cross-lingual NLP - NLP in the humanities

Models in Library35
Datasets in Library1
Models on Hugging Face1,563
Followers1.1k

Models

source languages: nl; target languages: en; OPUS readme: nl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: fr; target languages: en; OPUS readme: fr-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 75M parameters 512 tokens transformers

source languages: en; target languages: ru; OPUS readme: en-ru; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: en-de

Open weights cc-by-4.0 512 tokens transformers

hfname: kor-eng - sourcelanguages: kor - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/kor-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'korHani', 'korHang', 'korLatn', 'kor'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.test.txt - srcalpha3: kor - tgtalpha3: eng - shortpair: ko-en - chrF2score: 0.588 - brevitypenalty: 0.9590000000000001 - reflen: 17711.0 - srcname: Korean - tgtname: English - traindate…

Open weights apache-2.0 512 tokens transformers

source languages: de; target languages: en; OPUS readme: de-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: ru-en

Open weights cc-by-4.0 512 tokens transformers

hfname: spa-eng - sourcelanguages: spa - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/spa-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'spa'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/spa-eng/opus-2020-08-18.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/spa-eng/opus-2020-08-18.test.txt - srcalpha3: spa - tgtalpha3: eng - shortpair: es-en - chrF2score: 0.7390000000000001 - brevitypenalty: 0.9740000000000001 - reflen: 79376.0 - srcname: Spanish - tgtname: English - traindate: 2020-08-18 00:00:00…

Open weights apache-2.0 512 tokens transformers

source languages: en; target languages: fr; OPUS readme: en-fr; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: zho-eng helsinkigitsha: 480fcbe0ee1bf4774bcbe6226ad9f58e63f6c535 transformersgitsha: 2207e5d8cb224e954a7cba69fa4ac2309e9ff30b portmachine: brutasse porttime: 2020-08-21-14:41 srcmultilingual: False tgtmultilingual: False reflen: 82826.0 brevitypenalty…

Open weights cc-by-4.0 512 tokens transformers

hfname: eng-spa - sourcelanguages: eng - targetlanguages: spa - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eng-spa/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eng'} - tgtconstituents: {'spa'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-spa/opus-2020-08-18.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-spa/opus-2020-08-18.test.txt - srcalpha3: eng - tgtalpha3: spa - shortpair: en-es - chrF2score: 0.721 - brevitypenalty: 0.978 - reflen: 77311.0 - srcname: English - tgtname: Spanish - traindate: 2020-08-18 00:00:00 - srcalpha2: en - tgtalpha2…

Open weights apache-2.0 512 tokens transformers

Neural machine translation model for translating from English (en) to Turkish (tr). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…

Open weights cc-by-4.0 1,024 tokens transformers

Neural machine translation model for translating from Turkish (tr) to English (en). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…

Open weights cc-by-4.0 235M parameters 1,024 tokens transformers

source languages: ar; target languages: en; OPUS readme: ar-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: it; target languages: en; OPUS readme: it-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

Neural machine translation model for translating from Korean (ko) to English (en). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. - More information about released models for this language pair: OPUS-MT kor-eng README - Tatoeba Translation…

Open weights cc-by-4.0 209M parameters 1,024 tokens transformers

a sentence initial language token is required in the form of >>id<< (id = valid target language ID) - hfname: eng-zho - sourcelanguages: eng - targetlanguages: zho - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eng-zho/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eng'} - tgtconstituents: {'cmnHans', 'nan', 'nanHani', 'gan', 'yue', 'cmnKana', 'yueHani', 'wuuBopo', 'cmnLatn', 'yueHira', 'cmnHani', 'cjyHans', 'cmn', 'lzhHang', 'lzhHira', 'cmnHant', 'lzhBopo', 'zho', 'zhoHans', 'zhoHant', 'lzhHani', 'yueHang', 'wuu', 'yueKana', 'wuuLatn', 'yueBopo', 'cjyHant', 'yueHans', 'lzh', 'cmnHira', 'lzhYiii', 'lzhHans', 'cmnBopo', 'cmnHang'…

Open weights apache-2.0 512 tokens transformers

source languages: fr,frBE,frCA,frFR,wa,frp,oc,ca,rm,lld,fur,lij,lmo,es,esAR,esCL,esCO,esCR,esDO,esEC,esES,esGT,esHN,esMX,esNI,esPA,esPE,esPR,esSV,esUY,esVE,pt,ptbr,ptBR,ptPT,gl,lad,an,mwl,it,itIT,co,nap,scn,vec,sc,ro,la; target languages: en; OPUS readme: fr+frBE+frCA+frFR+wa+frp+oc+ca+rm+lld+fur+lij+lmo+es+esAR+esCL+esCO+esCR+esDO+esEC+esES+esGT+esHN+esMX+esNI+esPA+esPE+esPR+esSV+esUY+esVE+pt+ptbr+ptBR+ptPT+gl+lad+an+mwl+it+itIT+co+nap+scn+vec+sc+ro+la-en; dataset: opus; model: transformer; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

hfname: mul-eng - sourcelanguages: mul - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/mul-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'sjnLatn', 'cat', 'nan', 'spa', 'ileLatn', 'pap', 'mwl', 'uzbLatn', 'mww', 'hil', 'lij', 'avkLatn', 'ladLatn', 'latLatn', 'bosLatn', 'oss', 'epo', 'ron', 'fry', 'cym', 'toiLatn', 'awa', 'swg', 'zsmLatn', 'zhoHant', 'gcfLatn', 'uzbCyrl', 'isl', 'lfnLatn', 'shsLatn', 'novLatn', 'bho', 'ltz', 'lzh', 'kurLatn', 'sun', 'arg', 'pesThaa', 'sqi', 'uigArab', 'csbLatn', 'fra', 'hat', 'livLatn', 'nonLatn', 'sco', 'cmnHans', 'pnb', 'roh', 'chv', 'ibo', 'bulLatn', 'amh', 'lfnCyrl'…

Open weights apache-2.0 512 tokens transformers

source languages: fr; target languages: es; OPUS readme: fr-es; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: it; target languages: es; OPUS readme: it-es; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: tr; target languages: en; OPUS readme: tr-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

hfname: eus-spa - sourcelanguages: eus - targetlanguages: spa - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eus-spa/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eus'} - tgtconstituents: {'spa'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/eus-spa/opus-2020-06-17.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/eus-spa/opus-2020-06-17.test.txt - srcalpha3: eus - tgtalpha3: spa - shortpair: eu-es - chrF2score: 0.6729999999999999 - brevitypenalty: 0.9640000000000001 - reflen: 12469.0 - srcname: Basque - tgtname: Spanish - traindate: 2020-06-17 - srcalpha2…

Open weights apache-2.0 512 tokens transformers

Neural machine translation model for translating from Arabic (ar) to English (en). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…

Open weights cc-by-4.0 1,024 tokens transformers

source languages: da; target languages: en; OPUS readme: da-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: pl; target languages: en; OPUS readme: pl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

hfname: fin-eng - sourcelanguages: fin - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/fin-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'fin'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/fin-eng/opus-2020-08-05.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/fin-eng/opus-2020-08-05.test.txt - srcalpha3: fin - tgtalpha3: eng - shortpair: fi-en - chrF2score: 0.6970000000000001 - brevitypenalty: 0.99 - reflen: 74651.0 - srcname: Finnish - tgtname: English - traindate: 2020-08-05 - srcalpha2: fi…

Open weights apache-2.0 512 tokens transformers

source languages: ja; target languages: en; OPUS readme: ja-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: bg; target languages: en; OPUS readme: bg-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: nl; target languages: fr; OPUS readme: nl-fr; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: fr; target languages: de; OPUS readme: fr-de; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: sv; target languages: en; OPUS readme: sv-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: de; target languages: fr; OPUS readme: de-fr; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

a sentence initial language token is required in the form of >>id<< (id = valid target language ID) - hfname: eng-ara - sourcelanguages: eng - targetlanguages: ara - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/eng-ara/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'eng'} - tgtconstituents: {'apc', 'ara', 'arqLatn', 'arq', 'afb', 'araLatn', 'apcLatn', 'arz'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-ara/opus-2020-07-03.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/eng-ara/opus-2020-07-03.test.txt - srcalpha3: eng - tgtalpha3: ara - shortpair: en-ar…

Open weights apache-2.0 512 tokens transformers

Neural machine translation model for translating from English (en) to Bulgarian (bg). This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. You can also use OPUS-MT models with the transformers pipelines, for example: The work is supported by the European Language Grid as pilot…

Open weights cc-by-4.0 238M parameters 1,024 tokens transformers

Datasets

fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations are added for Czech (ces), Ukrainian (ukr) and Finnish (fin). For Czech and Ukrainian, this release doubles the data and for Finnish, we include translations for the entire fineweb-edu data set with its 350B token release. More information about how…

Publicly accessible odc-by