# bayon-it by Bayon: AACL-IJCNLP 2026: Open-Weight Model
Source: https://savrn.com/models/bayon-it
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve bayon-it (109M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 0.2 GB | 0.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[bayon-it on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/bayon-it/gpus)

## Model Card

By Bayon: AACL-IJCNLP 2026, published under apache-2.0, revision 6c78fcf7aaf6.

### Bayon Instruct

Bayon Instruct is an instruction-tuned version of [Bayon](https://savrn.com/models/bayon), a 100M-parameter decoder-only language model designed specifically for Khmer.

Bayon was pretrained from scratch using a custom 5,000-token Khmer BPE tokenizer and subsequently adapted for instruction following with LoRA. The instruction-tuning data consists of 18,000 Gemini-distilled, Khmer-focused SFT examples.

The model is intended primarily for Khmer text generation and instruction-following tasks, particularly where maintaining Khmer-language output is important.

### Model Details

| Property | Value |
| --- | --- |
| Base model | [attentionlab/bayon](https://savrn.com/models/bayon) |
| Parameters | ~100M |
| Architecture | Decoder-only Transformer |
| Language | Khmer (km) |
| Tokenizer | Custom 5,000-token Khmer BPE |
| Pretraining data | 1.361B Khmer tokens from FineWeb-2 |
| Instruction-tuning data | 18,000 SFT examples |
| Fine-tuning method | LoRA |
| LoRA rank | 16 |
| LoRA alpha | 32 |
| LoRA dropout | 0.1 |
| Context/training sequence length | 1,024 tokens |
| License | Apache-2.0 |

The tokenizer retains byte fallback, although byte fallback was not observed in the evaluation described in the associated research.

### Intended Use

Bayon Instruct is intended for:

[Read the full model card (863 words)](https://savrn.com/models/bayon-it/card)

## Configuration

Architecture

GemmaForCausalLM

Context length (tokens)

8,192

Layers

28

Hidden size

512

Feed-forward size

2,048

Attention heads

8

Key/value heads

2

Head dimension

64

Vocabulary size

5,000

RoPE base

10000

Stored precision

float32

Model type

gemma

## Identity and Version

Repository

attentionlab/bayon-it

Publisher

Bayon: AACL-IJCNLP 2026

Task

Text generation

Modality

Text

Library

Not stated by the source

Parameters

109M parameters

Languages

km

Revision

6c78fcf7aaf6b8169545498584725e8085f75731

First published

2026-07-24

Last updated

2026-10-04

## Files and Weights

7 files, 436.6 MB in total. The weights are 1 file totalling 436.1 MB in safetensors.

Weights1 file · 436.1 MB

Configuration2 files · 800 B

Tokenizer2 files · 458.1 KB

Documentation1 file · 8.2 KB

Repository1 file · 1.5 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 436.1 MB | 400b5c67564c |
| config.json | Configuration | 661 B | — |
| generation_config.json | Configuration | 139 B | — |
| README.md | Documentation | 8.2 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.model | Tokenizer | 352.9 KB | 60cb773534c3 |
| tokenizer.vocab | Tokenizer | 105.1 KB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

436.1 MB

[Download from Bayon: AACL-IJCNLP 2026](https://huggingface.co/attentionlab/bayon-it)

Released by Bayon: AACL-IJCNLP 2026 through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from [attentionlab/bayon](https://savrn.com/models/bayon)
- Trained on (disclosed) attentionlab/bayon-sft

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 436.1 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About bayon-it

### How much GPU memory does bayon-it need?

About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (109M parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run bayon-it on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use bayon-it commercially?

Yes. bayon-it is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

### What is bayon-it's context length?

8,192 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [bayon](https://savrn.com/models/bayon)

[Bayon: AACL-IJCNLP 2026](https://savrn.com/model-publishers/attentionlab)

Bayon is a 100M-parameter decoder-only generative language model pretrained from scratch for Khmer (km). The model was designed specifically for Khmer rather than being adapted from an existing multilingual or English-focused pretrained model. Its architecture and tokenizer were developed with Khmer text generation as the primary target. Bayon serves as the base pretrained model for Bayon Instruct. Research into language-specific model and tokenizer design Studying efficient language modeling for low-resource languages Bayon is a base pretrained model, not an instruction-tuned assistant. It may therefore produce continuations rather than direct answers when given natural-language questions…

Open weights apache-2.0 109M parameters 8,192 tokens

[View model](https://savrn.com/models/bayon)

Model · Text generation

### [lightning-105m](https://savrn.com/models/lightning-105m)

[AobanZ](https://savrn.com/model-publishers/aobangaming)

Lightning is a small, autoregressive transformer which utilizes FlashAttention and MHA. This model is trained on a variety of books from a dataset(300 MB). This is the expanded version of Lightning-60m. Lightning utilizes FlashAttention and AdamW for performance and capability. Lightning is designed to provide quick, coherent outputs, improved with a larger size and weight. Lightning is intended to be used for research, analysis and fine-tuning, stories and other. It is not intended to be used for professional advice, real writing or any kind of heavy work as generated outputs may be incorrect. Lightning can be used directly for text generation, experimentation, and conversational…

Open weights mit 105M parameters transformers

[View model](https://savrn.com/models/lightning-105m)

Model · Text generation

### [quipu-114m](https://savrn.com/models/quipu-114m)

[Aneek Chattopadhyay](https://savrn.com/model-publishers/aneekc)

A 114,114,048-parameter language model trained from scratch on a single 8 GB laptop This is a base model. It continues text; it does not follow instructions or answer questions. At this size it writes fluent, on-topic prose that is often factually wrong. It is published as a baseline and as a record of what one laptop can train in a weekend, not as something to rely on. Validation loss on held-out FineWeb-Edu text, and on held-out code files deduplicated by exact hash against the training code. \ The code loss is flattered by the tokenizer. GPT-2's BPE splits indentation whitespace, against 2.0% for text. Those are easy to predict and pull the average down. The same effect makes greedy code…

Open weights apache-2.0 114M parameters pytorch

[View model](https://savrn.com/models/quipu-114m)

Model · Text generation

### [SharperSwarm](https://savrn.com/models/sharperswarm)

[Convergent Intelligence](https://savrn.com/model-publishers/reaperdoesntknow)

SAGI (Swarm AGI) is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. V3.2 introduces a revolutionary Self-Assessment Layer, allowing the system to predict its own performance, identify skill gaps, and autonomously design its own learning curriculum. 1. Pre-Assessment: Predict success, identify risks, recommend strategy. 2. Execution: Generate with selected strategy. 3. Real-Time Monitoring: Catch and correct errors during generation. 4. Post-Assessment: Update skill…

Open weights apache-2.0 103M parameters 1,024 tokens transformers

[View model](https://savrn.com/models/sharperswarm)

Model · Text generation

### [macbert4csc-base-chinese](https://savrn.com/models/macbert4csc-base-chinese)

[Ming Xu (徐明)](https://savrn.com/model-publishers/shibing624)

macbert4csc-base-chinese evaluate SIGHAN2015 test data： 由于训练使用的数据使用了SIGHAN2015的训练集（复现paper），在SIGHAN2015的测试集上达到SOTA水平。 模型结构，魔改于softmaskedbert： 本项目开源在中文文本纠错项目：pycorrector，可支持macbert4csc模型，通过如下命令调用： 当然，你也可使用transformers调用： SIGHAN+Wang271K中文纠错数据集，数据格式： 如果需要训练macbert4csc，请参考https://github.com/shibing624/pycorrector/tree/master/pycorrector/macbert MacBERT is an improved BERT with novel MLM as correction pre-training task, which mitigates the discrepancy of pre-training and fine-tuning. Here is an example of our pre-training task. Except for the new pre-training task, we also incorporate the following techniques. Note that our MacBERT can be directly replaced with the original BERT as there is no…

Open weights apache-2.0 102M parameters 512 tokens transformers

[View model](https://savrn.com/models/macbert4csc-base-chinese)

Model · Text generation

### [uuu_fine_tune_gpt2](https://savrn.com/models/uuu-fine-tune-gpt2)

[David Lanz](https://savrn.com/model-publishers/davidlanz)

Fine tuning pre-trained language models for text generation. Pretrained model on Chinese language using a GPT2 for Large Language Head Model objective. transferlearning from DavidLanz/uuufinetunetaipower and fine-tuning with medical dataset for the GPT-2 architecture. You can use this model directly with a pipeline for text generation. Since the generation relies on some randomness, we

Open weights gpl 102M parameters transformers

[View model](https://savrn.com/models/uuu-fine-tune-gpt2)

## Bayon: AACL-IJCNLP 2026

[All models and datasets](https://savrn.com/model-publishers/attentionlab)

## Versions

- [6c78fcf7aaf6](https://savrn.com/models/bayon-it/versions/6c78fcf7aaf6) · current 2026-10-04

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-04.
- [Hugging Face record](https://huggingface.co/attentionlab/bayon-it)
- [How the hub is built](https://savrn.com/model-hub/methodology)
