SAVRN
Search Contact SAVRN

Open-weight model · Translation

nllb-200-distilled-600M

by AI at Meta facebook/nllb-200-distilled-600M

This is the model card of NLLB-200's distilled 600M variant. Here are the metrics for that particular checkpoint. - Information about training algorithms, parameters, fairness constraints or other applied approaches, and features.

Parameters
Context1,024
Weights2.5 GB
Licensecc-by-nc-4.0
AccessOpen weights
Monthly Downloads1.1M

Model Card

This is the model card of NLLB-200's distilled 600M variant. Here are the metrics for that particular checkpoint. - Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. - Paper or other resource for more information NLLB Team et al, No Language Left Behind: Scaling Human-Centered Machine Translation, Arxiv, 2022 - Where to send questions or comments about the model: https://github.com/facebookresearch/fairseq/issues • Model performance measures: NLLB-200…

Excerpt from the card by AI at Meta, licensed cc-by-nc-4.0.

Configuration

Architecture
M2M100ForConditionalGeneration
Context length (tokens)
1,024
Layers
12
Vocabulary size
256,206
Stored precision
float32
Model type
m2m_100

Identity and Version

Repository
facebook/nllb-200-distilled-600M
Publisher
AI at Meta
Task
Translation
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
ace, acm, acq, aeb, af, ajp, ak, als
Revision
f8d333a098d19b4fd9a8b18f94170487ad3f821d
First published
2022-07-08
Last updated
2024-02-14

Files and Weights

9 files, 2.5 GB in total. The weights are 1 file totalling 2.5 GB in bin.

Weights1 file · 2.5 GB
Configuration3 files · 4.6 KB
Tokenizer2 files · 17.3 MB
Documentation1 file · 7.7 KB
Other1 file · 4.9 MB
Repository1 file · 1.3 KB
Every file
FileTypeSizeSHA-256
pytorch_model.binWeights2.5 GB c266c2cfd197
config.jsonConfiguration846 B
generation_config.jsonConfiguration189 B
special_tokens_map.jsonConfiguration3.5 KB
README.mdDocumentation7.7 KB
sentencepiece.bpe.modelOther4.9 MB 14bb8dfb35c0
.gitattributesRepository1.3 KB
tokenizer.jsonTokenizer17.3 MB e316b82de11d
tokenizer_config.jsonTokenizer564 B

License and Download

License
cc-by-nc-4.0
Access
Open weights, no gate
Download size
2.5 GB
Download from AI at Meta

Released by AI at Meta through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) flores-200

Memory Requirements

PrecisionWeights in memory
As published2.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About nllb-200-distilled-600M

Can I use nllb-200-distilled-600M commercially?

Not without separate permission. nllb-200-distilled-600M is released under Creative Commons Attribution-NonCommercial 4.0. CC BY-NC 4.0 permits sharing and adapting with credit for non-commercial purposes only. Commercial use needs separate permission from the rights holder.

What is nllb-200-distilled-600M's context length?

1,024 tokens, from the maximum position embeddings in its published configuration.

Similar Models

source languages: nl; target languages: en; OPUS readme: nl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

source languages: en; target languages: ru; OPUS readme: en-ru; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: en-de

Open weights cc-by-4.0 512 tokens transformers

hfname: kor-eng - sourcelanguages: kor - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/kor-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'korHani', 'korHang', 'korLatn', 'kor'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.test.txt - srcalpha3: kor - tgtalpha3: eng - shortpair: ko-en - chrF2score: 0.588 - brevitypenalty: 0.9590000000000001 - reflen: 17711.0 - srcname: Korean - tgtname: English - traindate…

Open weights apache-2.0 512 tokens transformers

source languages: de; target languages: en; OPUS readme: de-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: ru-en

Open weights cc-by-4.0 512 tokens transformers