# Bayon: AACL-IJCNLP 2026: Open-Weight Models and Datasets
Source: https://savrn.com/model-publishers/attentionlab
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Models

Model · Text generation

### [bayon](https://savrn.com/models/bayon)

[Bayon: AACL-IJCNLP 2026](https://savrn.com/model-publishers/attentionlab)

Bayon is a 100M-parameter decoder-only generative language model pretrained from scratch for Khmer (km). The model was designed specifically for Khmer rather than being adapted from an existing multilingual or English-focused pretrained model. Its architecture and tokenizer were developed with Khmer text generation as the primary target. Bayon serves as the base pretrained model for Bayon Instruct. Research into language-specific model and tokenizer design Studying efficient language modeling for low-resource languages Bayon is a base pretrained model, not an instruction-tuned assistant. It may therefore produce continuations rather than direct answers when given natural-language questions…

Open weights apache-2.0 109M parameters 8,192 tokens

[View model](https://savrn.com/models/bayon)

Model · Text generation

### [bayon-it](https://savrn.com/models/bayon-it)

[Bayon: AACL-IJCNLP 2026](https://savrn.com/model-publishers/attentionlab)

Bayon Instruct is an instruction-tuned version of Bayon, a 100M-parameter decoder-only language model designed specifically for Khmer. Bayon was pretrained from scratch using a custom 5,000-token Khmer BPE tokenizer and subsequently adapted for instruction following with LoRA. The instruction-tuning data consists of 18,000 Gemini-distilled, Khmer-focused SFT examples. The model is intended primarily for Khmer text generation and instruction-following tasks, particularly where maintaining Khmer-language output is important. The tokenizer retains byte fallback, although byte fallback was not observed in the evaluation described in the associated research. Research on language-specific and…

Open weights apache-2.0 109M parameters 8,192 tokens

[View model](https://savrn.com/models/bayon-it)

## Explore More

- [All model publishers](https://savrn.com/model-publishers)
- [The model directory](https://savrn.com/models)
- [The dataset directory](https://savrn.com/datasets)

## Source

- Listed from their public repositories, read 2026-10-04.
- [Hugging Face profile](https://huggingface.co/attentionlab)
- [How the hub is built](https://savrn.com/model-hub/methodology)
