SAVRN
Search Contact SAVRN

Open-weight model · Translation

Mythos2.0-2B

by Adithyan bm AdithyanAI/Mythos2.0-2B

A Frontier Sparse Mixture-of-Experts (SMoE) Foundation Translation Model for 500+ Global Languages Mythos is an open-source foundation model family built by Adithyan AI.

Parameters
Context
Weights27.6 GB
Licenseapache-2.0
AccessAccess requested at publisher
Monthly Downloads

Model Card

By Adithyan bm, published under apache-2.0, revision 688abee06a4a.

A Frontier Sparse Mixture-of-Experts (SMoE) Foundation Translation Model for 500+ Global Languages Mythos is an open-source foundation model family built by Adithyan AI. In this organization, we develop and open-source state-of-the-art Sparse Mixture-of-Experts (SMoE) language models, universal translation engines, parallel multilingual datasets, and ultra-efficient inference runtimes targeting 500+ languages. 100% Free and Open-Source: Released under the permissive Apache 2.0 license with zero paywalls, metered tokens, or subscription fees. Mythos2.0-2B was trained on a massive 16.5B sentence-pair parallel corpus covering 500+ languages and regional dialects across Africa, the Americas…

Read Adithyan bm's full model card
# Mythos2.0-2B **A Frontier Sparse Mixture-of-Experts (SMoE) Foundation Translation Model for 500+ Global Languages**

Organization  ·   Supported Languages (500+)  ·   Collections  ·   Architecture  ·   Quickstart  ·   Collaboration


Welcome to Mythos AI

Mythos is an open-source foundation model family built by Adithyan AI. In this organization, we develop and open-source state-of-the-art Sparse Mixture-of-Experts (SMoE) language models, universal translation engines, parallel multilingual datasets, and ultra-efficient inference runtimes targeting 500+ languages.

  • Mission: Bridge the digital divide for underserved languages worldwide through efficient, open-weights AI architectures.
  • 100% Free and Open-Source: Released under the permissive Apache 2.0 license with zero paywalls, metered tokens, or subscription fees.
  • Community: Connect with our core team and contributors on Discord or explore our models on Hugging Face.

Supported Languages Directory (500+ Languages)

Mythos2.0-2B was trained on a massive 16.5B sentence-pair parallel corpus covering 500+ languages and regional dialects across Africa, the Americas, Asia, Europe, and Oceania.

To translate any text into a desired language, simply prepend the target language tag <2code> (e.g. <2es> for Spanish, <2ml> for Malayalam, <2hi> for Hindi, <2fr> for French, <2de> for German, <2ta> for Tamil).

Major Language Hubs Supported:

  • Global Commercial Languages: English (eng), Spanish (spa), French (fra), German (deu), Italian (ita), Portuguese (por), Russian (rus), Mandarin Chinese (cmn), Japanese (jpn), Korean (kor), Arabic (ara), Turkish (tur), Vietnamese (vie), Indonesian (ind), Dutch (nld), Polish (pol).
  • South Asian & Indian Languages: Hindi (hin), Malayalam (mal), Tamil (tam), Telugu (tel), Bengali (ben), Marathi (mar), Gujarati (guj), Kannada (kan), Punjabi (pan), Urdu (urd), Odia (ori), Assamese (asm), Sanskrit (san), Nepali (nep), Sinhala (sin), Maithili (mai), Bhojpuri (bho), Sindhi (snd), Kashmiri (kas), Konkani (kok).
  • African Languages: Swahili (swa), Amharic (amh), Yoruba (yor), Igbo (ibo), Hausa (hau), Somali (som), Oromo (orm), Zulu (zul), Xhosa (xho), Shona (sna), Tigrinya (tir), Malagasy (mlg), Kinyarwanda (kin), Lingala (lin), Bambara (bam), Wolof (wol).
  • European & Slavic Languages: Ukrainian (ukr), Czech (ces), Romanian (ron), Greek (ell), Hungarian (hun), Danish (dan), Finnish (fin), Norwegian (nob), Swedish (swe), Bulgarian (bul), Croatian (hrv), Serbian (srp), Slovak (slk), Catalan (cat), Basque (eus), Galician (glg), Irish (gle), Welsh (cym), Scottish Gaelic (gla).
  • Southeast Asian & Middle Eastern Languages: Thai (tha), Burmese (mya), Khmer (khm), Lao (lao), Tagalog / Filipino (fil), Cebuano (ceb), Persian / Farsi (pes), Hebrew (heb), Pashto (pus), Kurdish (kmr/ckb), Uyghur (uig), Kazakh (kaz), Uzbek (uzb), Azerbaijani (aze).

  • Americas & Indigenous Languages: Quechua (que), Guarani (grn), Aymara (aym), Nahuatl (nah), Navajo (nav), Mayan languages (myn), Inuktitut (iku), Cherokee (chr).

[!TIP] Universal Language Prompting:
To translate into any supported language, prepend <2{iso_code}> to your source text (e.g., <2es> for Spanish, <2hi> for Hindi, <2fr> for French, <2de> for German, <2ml> for Malayalam). All 500+ ISO-639 language codes are mapped directly into the model vocabulary and registered in the metadata above for automatic Hugging Face search filtering.


Mythos2.0 Model Collections

Following the modular design of frontier foundation families like Qwen, the Mythos2.0 Series spans foundation models, specialized context engines, and quantized edge runtimes:

Collection / Model Architecture Parameters (Total / Active) Context Window Target Capability Status
Mythos2.0-2B Sparse MoE (8E, Top-2) 2.04B / 678.7M 8,192 Flagship Universal 500+ Language Translation Active Run
Mythos2.0-4B Sparse MoE (16E, Top-2) 4.10B / 1.10B 16,384 Long Document and Legal/Technical Translation Pipeline
Mythos2.0-Edge-2B 2-bit / 4-bit SMoE 2.04B (~1.2 GB RAM) 4,096 Sub-2-bit Edge and Mobile Phone Deployment In Dev
Mythos-Tokenizer Byte-Level BPE 128,000 Vocab - Balanced Compression for 552 Global Languages Available
Mythos-16B-Corpus Parallel Bilingual Corpus 16 Billion Pairs - Bicleaner & LASER Curated Parallel Training Data Open Data

Explore all models in the official Hugging Face Collection:
https://huggingface.co/collections/AdithyanAI


Key Features of Mythos2.0-2B

  1. 500+ Global Languages Supported: Native, high-fidelity translation across major world languages plus 250+ underserved African, Indigenous American, and Regional South/Central Asian languages with zero coverage in commercial translation APIs.
  2. Sparse Mixture-of-Experts Efficiency: Employs 8 SwiGLU experts with Top-2 routing. With 2.04B total parameters, only 678.7M parameters are activated per token, delivering the translation capacity of a 7B-class model with the inference speed and memory footprint of a sub-1B model.
  3. 8k Native Context with Document Packing: Features a native 8,192-token context window (4,096 encoder + 4,096 decoder) with Block-Diagonal Attention Packing. Translates whole articles, SRT/VTT subtitles, and markdown documents without chunking or losing discourse context.
  4. FP32 Master Precision Embeddings: Maintains a 131M-parameter shared 3-way tied embedding table (src_embed, tgt_embed, proj.weight) in full 32-bit FP32 master weights, ensuring stable representation across rare scripts.

Comparison with Frontier Models

Specification Mythos2.0-2B TranslateGemma-7B NLLB-200 (3.3B) Google Cloud API
Architecture Sparse MoE (8E, Top-2) Dense Transformer Dense Enc-Dec Proprietary LLM
Total Parameters 2.04B 7.0B 3.3B Closed
Active Parameters / Token 678.7M 7.0B 3.3B Closed
Context Window 8,192 tokens 2,048 tokens 1,024 tokens Dynamic
Supported Languages 500+ 55 200 189
Min Inference VRAM ~4 GB 16 GB 8 GB Cloud API
License Apache 2.0 (100% Free) Community License CC-BY-NC 4.0 Paid Metered API

Model Architecture Overview

  • Model Family: Mythos2.0
  • Model ID: AdithyanAI/Mythos2.0-2B
  • Architecture: Encoder-Decoder Sparse Mixture-of-Experts (SMoE)
  • Total Parameters: 2.04B (2,037,643,264)
  • Active Parameters / Token: 678.7M (678,688,768)
  • Layers: 24 Transformer Blocks (12 Encoder Layers + 12 Decoder Layers)
  • Hidden Dimension ($d_{\text{model}}$): 1,024
  • Attention Mechanism: Grouped Query Attention (GQA)
  • 16 Query Heads, 4 Key-Value Head Groups (4x KV compression)
  • Head Dimension: 64
  • RoPE Base Frequency: $\theta = 100,000.0$
  • Feed-Forward Network (Sparse MoE):
  • 8 SwiGLU Experts per layer (Intermediate Dim: 3,072)
  • Top-2 Routing with Switch-Transformer Capacity Factor (1.35)
  • Calibrated Load Balancing Loss (aux_loss_weight = 0.01)
  • Context Capacity: 8,192 tokens (4k Source + 4k Target Document-Packed)
  • Vocabulary: 128,000 Byte-Level BPE Tokens (552 languages)

Quickstart: Free Offline Inference

1. Installation

pip install torch transformers tokenizers sacrebleu

2. Python Inference Code

import torch
from tokenizers import Tokenizer

# Load Tokenizer
tokenizer = Tokenizer.from_file("multilingual_tokenizer.json")
sos_id = tokenizer.token_to_id("[SOS]")
eos_id = tokenizer.token_to_id("[EOS]")
pad_id = tokenizer.token_to_id("[PAD]")

def translate(model, text: str, tgt_lang: str = "fra", max_len: int = 128, device: str = "cuda:0"):
    clean_tgt = tgt_lang.split("_")[0]
    prompt = f"<2{clean_tgt}> {text}"
    tokens = tokenizer.encode(prompt).ids
    src_tensor = torch.tensor([tokens], dtype=torch.long, device=device)
    src_mask = (src_tensor != pad_id).unsqueeze(1).unsqueeze(2)

    with torch.no_grad():
        enc_out = model.encode(src_tensor, src_mask)
        gen_tokens = torch.tensor([[sos_id, tokenizer.token_to_id(f"<2{clean_tgt}>")]], device=device)

        for _ in range(max_len):
            cur_len = gen_tokens.size(1)
            causal_mask = torch.tril(torch.ones((cur_len, cur_len), dtype=torch.bool, device=device)).unsqueeze(0).unsqueeze(0)
            dec_out = model.decode(gen_tokens, enc_out, src_mask, causal_mask)
            logits = model.project(dec_out[:, -1:])
            next_token = logits.argmax(dim=-1).item()
            if next_token == eos_id:
                break
            gen_tokens = torch.cat([gen_tokens, torch.tensor([[next_token]], device=device)], dim=1)

    return tokenizer.decode(gen_tokens[0].tolist()[2:])

Training Infrastructure and Engineering

  • Distributed Engine: PyTorch Fully Sharded Data Parallel (FSDP) and DDP.
  • Precision Policy: Native 32-bit FP32 Master Weights with FP16 compute and unscaled FP32 logits projection.
  • Zero-Host RAM Footprint: Streaming disk-spooler architecture keeping host CPU memory strictly below < 1.0 GB throughout training.
  • Router Stabilization: Switch-Transformer dynamic capacity factor capping (capacity_factor = 1.35) with calibrated auxiliary loss (0.01) preventing expert collapse.

Community and Collaboration

We welcome researchers, linguists, and compute sponsors to join the Mythos AI initiative:

  • Compute Sponsors: Pooling idle GPU hours (RTX 3090/4090, A100, H100) to scale Mythos2.0-4B and Mythos2.0-7B pre-training.
  • Researchers and Engineers: Optimizing routing loss, sparse kernels, and sub-2-bit quantization.
  • Native Linguists: Auditing translation quality and expanding low-resource parallel corpora.
  • Official Discord: https://discord.gg/KKVN5BShGj

Citation

If you use Mythos2.0-2B in your research or applications, please cite:

@misc{mythos2026multilingual,
  author       = {Adithyan AI and Community Contributors},
  title        = {Mythos2.0-2B: A Free and Open-Source Sparse Mixture-of-Experts Translation Foundation Model for 500+ Languages},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/AdithyanAI/Mythos2.0-2B}}
}

Identity and Version

Repository
AdithyanAI/Mythos2.0-2B
Publisher
Adithyan bm
Task
Translation
Modality
Text
Library
transformers
Parameters
Not stated by the source
Languages
en, es, fr, de, it, pt, hi, zh
Revision
688abee06a4ac3ff4180ccf45a2047abfd5af191
First published
2026-08-23
Last updated
2026-09-18

Files and Weights

12 files, 27.6 GB in total. The weights are 4 files totalling 27.6 GB in pt.

Weights4 files · 27.6 GB
Tokenizer1 file · 10.0 MB
Documentation1 file · 18.0 KB
Other5 files · 903.8 KB
Repository1 file · 2.4 KB
Every file
FileTypeSizeSHA-256
Mythos_Translator.ptWeights9.2 GB
best_checkpoint.ptWeights9.2 GB
fixed_val_benchmark.ptWeights1.9 MB
last_checkpoint.ptWeights9.2 GB
README.mdDocumentation18.0 KB
figures/mythos_icon.pngOther482.8 KB
figures/mythos_logo.svgOther2.5 KB
runs/events.out.tfevents.1789653379.276b212c64a7.74.0Other238.1 KB
runs/events.out.tfevents.1789699125.7b69f9e914cb.74.0Other179.5 KB
validation_metrics.csvOther905 B
.gitattributesRepository2.4 KB
multilingual_tokenizer.jsonTokenizer10.0 MB

License and Download

License
apache-2.0
Access
Access requested at publisher
Download size
27.6 GB
Request access from Adithyan bm

Adithyan bm grants access through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published27.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Mythos2.0-2B

Can I use Mythos2.0-2B commercially?

Yes. Mythos2.0-2B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

source languages: nl; target languages: en; OPUS readme: nl-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

Model · Translation

nllb-200-distilled-600M

AI at Meta

This is the model card of NLLB-200's distilled 600M variant. Here are the metrics for that particular checkpoint. - Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. - Paper or other resource for more information NLLB Team et al, No Language Left Behind: Scaling Human-Centered Machine Translation, Arxiv, 2022 - Where to send questions or comments about the model: https://github.com/facebookresearch/fairseq/issues • Model performance measures: NLLB-200…

Open weights cc-by-nc-4.0 1,024 tokens transformers

source languages: en; target languages: ru; OPUS readme: en-ru; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers

This model can be used for translation and text-to-text generation. CONTENT WARNING: Readers should be aware this section contains content that is disturbing, offensive, and can propagate historical and current stereotypes. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Further details about the dataset for this model can be found in the OPUS readme: en-de

Open weights cc-by-4.0 512 tokens transformers

hfname: kor-eng - sourcelanguages: kor - targetlanguages: eng - opusreadmeurl: https://github.com/Helsinki-NLP/Tatoeba-Challenge/tree/master/models/kor-eng/README.md - originalrepo: Tatoeba-Challenge - srcconstituents: {'korHani', 'korHang', 'korLatn', 'kor'} - tgtconstituents: {'eng'} - srcmultilingual: False - tgtmultilingual: False - urlmodel: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.zip - urltestset: https://object.pouta.csc.fi/Tatoeba-MT-models/kor-eng/opus-2020-06-17.test.txt - srcalpha3: kor - tgtalpha3: eng - shortpair: ko-en - chrF2score: 0.588 - brevitypenalty: 0.9590000000000001 - reflen: 17711.0 - srcname: Korean - tgtname: English - traindate…

Open weights apache-2.0 512 tokens transformers

source languages: de; target languages: en; OPUS readme: de-en; dataset: opus; model: transformer-align; pre-processing: normalization + SentencePiece.

Open weights apache-2.0 512 tokens transformers