Bayon is a 100M-parameter decoder-only generative language model pretrained from scratch for Khmer (km). The model was designed specifically for Khmer rather than being adapted from an existing multilingual or English-focused pretrained model. Its architecture and tokenizer were developed with Khmer text generation as the primary target. Bayon serves as the base pretrained model for Bayon Instruct. Research into language-specific model and tokenizer design Studying efficient language modeling for low-resource languages Bayon is a base pretrained model, not an instruction-tuned assistant. It may therefore produce continuations rather than direct answers when given natural-language questions…
quipu-114m is an open-weight model for text generation from Aneek Chattopadhyay, released under Apache License 2.0. It has 114M parameters. At 16-bit it needs about 0.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
A 114,114,048-parameter language model trained from scratch on a single 8 GB laptop This is a base model. It continues text; it does not follow instructions or answer questions. At this size it writes fluent, on-topic prose that is often factually wrong.
Runs On
What it takes to serve quipu-114m (114M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
quipu-114m on every accelerator the SAVRN Index prices, at every precision
Model Card
By Aneek Chattopadhyay, published under apache-2.0, revision 4666bea7604e.
A 114,114,048-parameter language model trained from scratch on a single 8 GB laptop This is a base model. It continues text; it does not follow instructions or answer questions. At this size it writes fluent, on-topic prose that is often factually wrong. It is published as a baseline and as a record of what one laptop can train in a weekend, not as something to rely on. Validation loss on held-out FineWeb-Edu text, and on held-out code files deduplicated by exact hash against the training code. \ The code loss is flattered by the tokenizer. GPT-2's BPE splits indentation whitespace, against 2.0% for text. Those are easy to predict and pull the average down. The same effect makes greedy code…
Read Aneek Chattopadhyay's full model card
A 114,114,048-parameter language model trained from scratch on a single 8 GB laptop GPU (RTX 5060 Laptop): own weights, own data pipeline, no fine-tuned base.
This is a base model. It continues text; it does not follow instructions or answer questions. At this size it writes fluent, on-topic prose that is often factually wrong. It is published as a baseline and as a record of what one laptop can train in a weekend, not as something to rely on.
Project page: https://quipu-lm.vercel.app
Architecture
| Type | Decoder-only transformer, pre-norm |
| Parameters | 114,114,048 (input and output embeddings tied) |
| Layers / width | 12 layers, d_model 768, SwiGLU feed-forward 2,048 |
| Attention | Grouped-query: 12 query heads over 4 key/value heads, head dim 64 |
| Positions | Rotary embeddings (GPT-NeoX half-split), base 10,000 |
| Norm | RMSNorm, eps 1e-6 |
| Context | 1,024 tokens |
| Tokenizer | GPT-2 BPE via tiktoken (vocab 50,257) |
Training
| Tokens | 2,999,975,936 (one pass, no data repeated) |
| Mix | 80% FineWeb-Edu (sample-10BT), 20% permissively licensed code from github-code-clean |
| Code filter | MIT, Apache-2.0, BSD, ISC, CC0, Unlicense only; minified, vendored and >16k-token files removed; HTML capped at 10% |
| Steps | 5,722 × 524,288 tokens |
| Optimiser | AdamW, cosine schedule, 200 warmup steps, bf16 autocast |
| Hardware | 1 × RTX 5060 Laptop GPU (8 GB), Windows |
| Wall clock | 45 h 22 min, ~18,400 tokens/s |
| Incidents | 0 resumes, 0 skipped (non-finite) steps |
Results
Validation loss on held-out FineWeb-Edu text, and on held-out code files deduplicated by exact hash against the training code.
| Step | Tokens | Text val loss | Code val loss* |
|---|---|---|---|
| 100 | 52M | 6.40 | 4.36 |
| 500 | 262M | 4.40 | 2.06 |
| 1,000 | 524M | 3.85 | 1.42 |
| 2,000 | 1.05B | 3.58 | 1.20 |
| 4,000 | 2.10B | 3.38 | 1.02 |
| 5,722 | 3.00B | 3.31 | 0.97 |
* The code loss is flattered by the tokenizer. GPT-2's BPE splits indentation into many whitespace tokens: 36.5% of the code validation tokens are pure whitespace, against 2.0% for text. Those are easy to predict and pull the average down. The same effect makes greedy code generation collapse into runs of spaces. Do not compare this number with code models that use a code-aware tokenizer. The next Quipu model uses one.
The text loss is not directly comparable with other small models either: the data mix (20% code) and evaluation set differ.
What it writes
Final model, greedy decoding:
Photosynthesis is the process by which plants convert carbon dioxide into oxygen and release carbon dioxide into the atmosphere.
The history of the printing press is a fascinating one. It was invented in 1848 by the French printer Pierre-Louis Leclerc.
Both are fluent and both are wrong, which is the honest summary of a 114M base model. Greedy decoding also repeats itself. Sampling (temperature 0.8, top-k 50) reads better. Samples from every milestone, for every prompt, are in the project repository.
Files
| File | |
|---|---|
model.safetensors |
Final weights, fp32 (step 5,722) |
milestones/step_*.safetensors |
bf16 snapshots at steps 100, 250, 500, 1,000, 2,000, 4,000 and 5,722, for studying how the model learns |
config.json |
Architecture |
modeling_quipu.py |
Standalone PyTorch implementation and loader |
lm_head.weight is not stored; it is the embedding matrix.
Usage
pip install torch safetensors tiktoken huggingface_hub
from huggingface_hub import hf_hub_download
import importlib.util, sys
path = hf_hub_download("quipu-lm/quipu-114m", "modeling_quipu.py")
spec = importlib.util.spec_from_file_location("modeling_quipu", path)
mq = importlib.util.module_from_spec(spec)
sys.modules["modeling_quipu"] = mq
spec.loader.exec_module(mq)
model = mq.load("quipu-lm/quipu-114m") # final model
early = mq.load("quipu-lm/quipu-114m", weights="milestones/step_001000.safetensors")
print(mq.generate(model, "Photosynthesis is the process by which",
max_new_tokens=60, temperature=0.8, seed=1337))
This is not a transformers model; there is no AutoModel class for it.
Limitations
- Base model only: no instruction tuning, no safety tuning, no chat format.
- Frequently states false facts with confidence. Do not use its output as information.
- Code output is syntax-shaped but rarely correct, and greedy decoding degenerates into whitespace (see the tokenizer note above).
- English only. 1,024-token context.
- Trained on web text, which carries the biases of web text.
Data and licence
Weights and code: Apache-2.0.
Training data: FineWeb-Edu (ODC-By 1.0, © Hugging Face) and the permissively licensed subset of codeparrot/github-code-clean. Code files keep their original licences; only MIT, Apache-2.0, BSD, ISC, CC0 and Unlicense files were used.
Configuration
- Vocabulary size
- 50,257
Identity and Version
- Repository
- AneekC/quipu-114m
- Publisher
- Aneek Chattopadhyay
- Task
- Text generation
- Modality
- Text
- Library
- pytorch
- Parameters
- 114M parameters
- Languages
- en
- Revision
- 4666bea7604e0f47a4f9c7bb53fe7172c8bd3d97
- First published
- 2026-09-28
- Last updated
- 2026-09-28
Files and Weights
12 files, 2.1 GB in total. The weights are 8 files totalling 2.1 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| milestones/step_000100.safetensors | Weights | 228.2 MB | fd818bfbc711 |
| milestones/step_000250.safetensors | Weights | 228.2 MB | 23de897d5ff2 |
| milestones/step_000500.safetensors | Weights | 228.2 MB | 65a8d2fff13a |
| milestones/step_001000.safetensors | Weights | 228.2 MB | a9d363833fb2 |
| milestones/step_002000.safetensors | Weights | 228.2 MB | 5ed74120cbb9 |
| milestones/step_004000.safetensors | Weights | 228.2 MB | 4f739e5d1c20 |
| milestones/step_005722.safetensors | Weights | 228.2 MB | 3497f20858f7 |
| model.safetensors | Weights | 456.5 MB | 3580f319f97a |
| config.json | Configuration | 311 B | — |
| modeling_quipu.py | Configuration | 7.6 KB | — |
| README.md | Documentation | 5.2 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 2.1 GB
Released by Aneek Chattopadhyay through its official repository on Hugging Face. Read the license.
Built From
- Trained on (disclosed) HuggingFaceFW/fineweb-edu
- Trained on (disclosed) codeparrot/github-code-clean
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 2.1 GB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About quipu-114m
How much GPU memory does quipu-114m need?
About 0.3 GB at 16-bit and 0.1 GB at 4-bit: the weights (114M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run quipu-114m on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use quipu-114m commercially?
Yes. quipu-114m is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Similar Models
Bayon Instruct is an instruction-tuned version of Bayon, a 100M-parameter decoder-only language model designed specifically for Khmer. Bayon was pretrained from scratch using a custom 5,000-token Khmer BPE tokenizer and subsequently adapted for instruction following with LoRA. The instruction-tuning data consists of 18,000 Gemini-distilled, Khmer-focused SFT examples. The model is intended primarily for Khmer text generation and instruction-following tasks, particularly where maintaining Khmer-language output is important. The tokenizer retains byte fallback, although byte fallback was not observed in the evaluation described in the associated research. Research on language-specific and…
Lightning is a small, autoregressive transformer which utilizes FlashAttention and MHA. This model is trained on a variety of books from a dataset(300 MB). This is the expanded version of Lightning-60m. Lightning utilizes FlashAttention and AdamW for performance and capability. Lightning is designed to provide quick, coherent outputs, improved with a larger size and weight. Lightning is intended to be used for research, analysis and fine-tuning, stories and other. It is not intended to be used for professional advice, real writing or any kind of heavy work as generated outputs may be incorrect. Lightning can be used directly for text generation, experimentation, and conversational…
This model is a fine-tuned version of GeorgeUwaifo/iviegpt2new01cresults on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 4 - evalbatchsize: 8 - gradientaccumulationsteps: 4 - totaltrainbatchsize: 16 - lrschedulertype: linear - lrschedulerwarmupsteps: 387 - numepochs: 5 - mixedprecisiontraining: Native AMP - Transformers 5.16.1 - Pytorch 2.11.0+cu128 - Datasets 4.8.5 - Tokenizers 0.23.1
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
This is the lab4-yoru-AY-482206 checkpoint, trained from scratch on English C4. It was selected by the lowest development loss among nine experimental recipes. The independent seed repeat and reserved audit evaluation were still pending when this checkpoint was published. Reported scores are local evaluation proxies, not an official online-judge result. - Stock Hugging Face GPT2LMHeadModel: 12 layers, 12 attention heads, hidden width 768, context length 1,024, vocabulary 50,257, tied input/output embeddings. - 124,439,808 unique parameters, commonly described as GPT-2 small. The lab uses the historical 117M model-family label. - GPT-2 tokenizer; documents packed with EOS separators. No…