SAVRN
Search Contact SAVRN

Open-weight model · Translation

nllb-200-3.3B

by AI at Meta facebook/nllb-200-3.3B

This is the model card of NLLB-200's 3.3B variant. Here are the metrics for that particular checkpoint. - Information about training algorithms, parameters, fairness constraints or other applied approaches, and features.

Parameters
Context1,024
Weights17.6 GB
Licensecc-by-nc-4.0
AccessOpen weights
Monthly Downloads186.7k

Model Card

This is the model card of NLLB-200's 3.3B variant. Here are the metrics for that particular checkpoint. - Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. - Paper or other resource for more information NLLB Team et al, No Language Left Behind: Scaling Human-Centered Machine Translation, Arxiv, 2022 - Where to send questions or comments about the model: https://github.com/facebookresearch/fairseq/issues • Model performance measures: NLLB-200 model was…

Excerpt from the card by AI at Meta, licensed cc-by-nc-4.0.

Configuration

Architecture
M2M100ForConditionalGeneration
Context length (tokens)
1,024
Layers
24
Vocabulary size
256,206
Stored precision
float32
Model type
m2m_100

Identity and Version

Repository
facebook/nllb-200-3.3B
Publisher
AI at Meta
Task
Translation
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
ace, acm, acq, aeb, af, ajp, ak, als
Revision
1a07f7d195896b2114afcb79b7b57ab512e7b43e
First published
2022-07-08
Last updated
2023-02-11

Files and Weights

12 files, 17.6 GB in total. The weights are 3 files totalling 17.6 GB in bin.

Weights3 files · 17.6 GB
Configuration4 files · 94.6 KB
Tokenizer2 files · 17.3 MB
Documentation1 file · 7.6 KB
Other1 file · 4.9 MB
Repository1 file · 1.2 KB
Every file
FileTypeSizeSHA-256
pytorch_model-00001-of-00003.binWeights6.9 GB 06aec913c891
pytorch_model-00002-of-00003.binWeights8.5 GB 1c5a44acb028
pytorch_model-00003-of-00003.binWeights2.1 GB 2cf744009996
config.jsonConfiguration808 B
generation_config.jsonConfiguration189 B
pytorch_model.bin.index.jsonConfiguration90.0 KB
special_tokens_map.jsonConfiguration3.5 KB
README.mdDocumentation7.6 KB
sentencepiece.bpe.modelOther4.9 MB 14bb8dfb35c0
.gitattributesRepository1.2 KB
tokenizer.jsonTokenizer17.3 MB e316b82de11d
tokenizer_config.jsonTokenizer564 B

License and Download

License
cc-by-nc-4.0
Access
Open weights, no gate
Download size
17.6 GB
Download from AI at Meta

Released by AI at Meta through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) flores-200

Memory Requirements

PrecisionWeights in memory
As published17.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About nllb-200-3.3B

Can I use nllb-200-3.3B commercially?

Not without separate permission. nllb-200-3.3B is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

What is nllb-200-3.3B's context length?

1,024 tokens, from the maximum position embeddings in its published configuration.

Similar Models

source languages: nl; target languages: en; OPUS readme: nl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

Model · Translation

nllb-200-distilled-600M

AI at Meta

This is the model card of NLLB-200's distilled 600M variant. Here are the metrics for that particular checkpoint. - Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. - Paper or other resource for more information NLLB Team et al, No Language Left Behind: Scaling Human-Centered Machine Translation, Arxiv, 2022 - Where to send questions or comments about the model: https://github.com/facebookresearch/fairseq/issues • Model performance measures: NLLB-200…

Open weights cc-by-nc-4.0 1,024 tokens transformers

source languages: en; target languages: ru; OPUS readme: en-ru; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: en-de

Open weights cc-by-4.0 512 tokens transformers

hfname: kor-eng - sourcelanguages: kor - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/kor-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'korHani', 'korHang', 'korLatn', 'kor'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.test.txt - srcalpha3: kor - tgtalpha3: eng - shortpair: ko-en - chrF2score: 0.588 - brevitypenalty: 0.9590000000000001 - reflen: 17711.0 - srcname: Korean - tgtname: English - traindate…

Open weights apache-2.0 512 tokens transformers

source languages: de; target languages: en; OPUS readme: de-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers