Bayon is a 100M-parameter decoder-only generative language model pretrained from scratch for Khmer (km). The model was designed specifically for Khmer rather than being adapted from an existing multilingual or English-focused pretrained model. Its architecture and tokenizer were developed with Khmer text generation as the primary target. Bayon serves as the base pretrained model for Bayon Instruct. Research into language-specific model and tokenizer design Studying efficient language modeling for low-resource languages Bayon is a base pretrained model, not an instruction-tuned assistant. It may therefore produce continuations rather than direct answers when given natural-language questions…
Organization
Bayon: AACL-IJCNLP 2026
attentionlab
Models
Bayon Instruct is an instruction-tuned version of Bayon, a 100M-parameter decoder-only language model designed specifically for Khmer. Bayon was pretrained from scratch using a custom 5,000-token Khmer BPE tokenizer and subsequently adapted for instruction following with LoRA. The instruction-tuning data consists of 18,000 Gemini-distilled, Khmer-focused SFT examples. The model is intended primarily for Khmer text generation and instruction-following tasks, particularly where maintaining Khmer-language output is important. The tokenizer retains byte fallback, although byte fallback was not observed in the evaluation described in the associated research. Research on language-specific and…