LibraMind (565M parameters)
LibraMind is a 565M-parameter chat model pretrained from scratch by alby13 on a single consumer GPU (NVIDIA RTX 4090). It is a hybrid architecture that interleaves 18 Gated DeltaNet linear-recurrent layers with 6 Gated Attention layers (a repeating DDDA × 6 pattern), with Block Attention Residuals and logit soft-capping, trained with the Muon optimizer. It went through a full five-stage pipeline: pretraining, midtraining, supervised fine-tuning, preference tuning, and RL with verifiable rewards.
Research model: no safety training.
LibraMind has not undergone safety training or red-teaming. It can produce inaccurate, biased, offensive, or harmful text, and it will not reliably refuse requests. It also makes factual and arithmetic mistakes (see the examples below). It is released as a research artifact, provided "as is" without warranty of any kind, and the authors accept no liability for its outputs or for any use of the model. You are solely responsible for evaluating it and for any safeguards, filtering, and human oversight your use case needs. Do not deploy it in user-facing, high-stakes, or safety-critical settings without your own safety work.
This notice is guidance. It does not add to or modify the terms of the Apache-2.0 license.
Model Summary
|
|
| Developer |
alby13 |
| Parameters |
565M (565.2M) |
| Model type |
Chat model |
| Architecture |
Hybrid Gated DeltaNet + Gated Attention, decoder-only |
| Layers / width |
24 layers, hidden size 1,024 (DDDA × 6: 3 DeltaNet layers then 1 Attention layer, repeated 6 times) |
| Vocabulary |
32,768 (custom byte-level BPE) |
| Context length |
4,096 tokens (pretrained at 2,048, extended during midtraining) |
| Weights precision |
bf16 (about 1.1 GB) |
| Language |
English |
| License |
Apache-2.0 |
| Release |
10/6/2026 |
Intended use
- Light assistant tasks: casual conversation, rewriting and summarizing short texts, formatted answers such as lists, sections and word limits, and simple function calling.
- Hobby and research: a fully documented small-model training run (architecture, data, every stage, evaluations).
- Game NPCs and interactive characters: dialogue driven by a character card in the system prompt, with optional schema-forced JSON for emotions and actions.
Intended Use
Intended for:
- Research on small hybrid recurrent/attention language models
- Studying low-resource, single-GPU training pipelines (pretraining through RL) and the Muon optimizer at small scale
- Experiments with on-device chat, role-play, and structured output (JSON / tool calls), with your own validation of the results (for example, validate JSON and compute numbers in code)
Not intended for:
- High-stakes or safety-critical decisions (medical, legal, financial, etc.)
- Unsupervised, user-facing deployment: the model has no safety training
- Use as a source of factual information without independent verification
How to Use
LibraMind uses a custom architecture, so it doesn't load in transformers, llama.cpp or other GGUF runtimes. It needs this repository's model.py, inference.py and chat_format.py, plus:
PyTorch with CUDA (trained with 2.14)
flash-linear-attention 0.5.2, with Triton
tokenizers
llguidance (optional, for grammar-forced JSON)
In this repository, the weights are runs/rl1/model_final.pt (float32, 2.26 GB) and the tokenizer is data/chat/sft/tokenizer.json.
import torch
from tokenizers import Tokenizer
from chat_format import ChatEncoder
from inference import GrammarFactory, generate
from model import LM, ModelConfig
ck = torch.load("model_final.pt", map_location="cpu", weights_only=False)
model = LM(ModelConfig(**{**ck["model_config"], "grad_ckpt": False}))
model.load_state_dict(ck["model"])
model = model.cuda().to(torch.bfloat16).eval()
enc = ChatEncoder(Tokenizer.from_file("tokenizer.json"))
def chat(messages, schema=None, max_new=300):
matcher = [GrammarFactory("tokenizer.json", enc.im_end, model.config.vocab_size).json(schema)] if schema else None
with torch.autocast("cuda", dtype=torch.bfloat16):
out = generate(model, [enc.encode_prompt(messages)], max_new, temperature=0.7, top_p=0.9, rep_penalty=1.1,
stop_ids=[enc.im_end], matchers=matcher, vocab=enc.tok.get_vocab_size())[0]
return enc.tok.decode(out)
print(chat([{"role": "system", "content": "You are Brom, a gruff dwarven blacksmith. Stay in character."},
{"role": "user", "content": "Can you fix my sword?"}]))
generate uses cached decoding: Gated DeltaNet recurrent state plus an attention KV cache, replayed as a CUDA graph. It accepts many prompts at once, at about 1,050 tokens/s across 64 parallel conversations. The repository also has a browser chat window (chatui.cmd) and a console chat (chat.cmd).
Recommended settings:
- temperature 0.4, top-p 0.9, repetition penalty 1.1
- System prompt:
"You are LibraMind, a friendly and thoughtful conversation partner made by alby13. Keep replies natural and to the point: a few sentences for casual chat, more detail only when it's needed. Ask a follow-up question when it helps the conversation.
- temperature 0 for JSON and tool calls
- always pass stop_ids=[<|im_end|>]
Requirements:
Memory: about 1.1 GB for the weights in bf16, plus activations. The recurrent layers use constant-size state, and only the 6 attention layers grow a KV cache with context length.
Chat format
ChatML with a beginning-of-text token. Roles are system, user, assistant and tool:
<|bos|><|im_start|>system
{system prompt}<|im_end|>
<|im_start|>user
{message}<|im_end|>
<|im_start|>assistant
{reply}<|im_end|>
The system prompt is optional. Without one, the model behaves as the LibraMind assistant.
Characters and Personas
Put the character card in the system prompt: name, setting, personality, how they speak, and any facts they must get right. That means prices, quest details, names and places. Roleplay training used this pattern:
You are {name}. Stay in character: speak as {name} would, in the first person, and never say you are an AI.
Character profile:
{profile}
For game actions, force the reply into a JSON schema with GrammarFactory.json(schema). The output is then guaranteed to parse and match the schema. Example schema: examples/npc_schema.json, with fields line, emotion and action. Describe the fields in the system prompt too, so the content makes sense as well as the format.
Function calling
List the tools in the system prompt in the format the model was trained on:
You can call the following functions:
<tools>
[{"type": "function", "function": {"name": "get_weather", "description": "...", "parameters": {...}}}]
</tools>
To call a function, reply with a JSON object inside <tool_call></tool_call> tags, for example:
<tool_call>
{"name": "function_name", "arguments": {"arg": "value"}}
</tool_call>
You may make several calls. Function results come back in <tool_response> tags.
GrammarFactory.tool_calls(tools) can force calls to name a listed function, with arguments that match its schema.
Example Outputs
sampling settings used (temperature 0.4, top-p 0.9, repetition penalty 1.1).
General chat
User: Hi! Who are you?
LibraMind: I'm LibraMind, a small AI language model trained from scratch by alby13. How can I help?
Role-play
System Prompt: You are Grom, a gruff dwarven blacksmith in the mountain town of Bringham. Stay in character and keep replies short.
User: Where can I find some armor? I need some that also protects against magic.
LibraMind: Well, I ain't no scholar, but I've heard tales from travelers. You might find some armor in the blacksmith's forge. Ask the old sage, Thorne. He knows all about such things.
System Prompt: You are Jin, the stern innkeeper of the Gilded Redbird inn in the village of Ashfield. A room costs 15 silver a night and a bowl of stew costs 3 silver. (...)
User: How much for a room tonight and some supper?
LibraMind: A room for 15 silver is a good price. A bowl of stew for 3 silver will be 9 silver.
Structured output (schema-forced decoding; the system prompt asks for JSON with line, emotion, and action)
System: You are Grom (...). Reply only with JSON: "line" is what you say, "emotion" is how you feel, and "action" is what you do (open_shop if the player wants to buy or get something repaired).
User: I'd like to buy a new axe.
Sample 1: {"line": "axe", "emotion": "happy", "action": "open_shop"}
Sample 2: {"line": "You've come to the right place, friend. I've got just the thing for you. What's your name, and what kind of axe are you looking for?", "emotion": "neutral", "action": "open_shop"}
Instruction following
User: Give me three tips for memorizing for studying as a numbered list.
LibraMind:
1. Create a dedicated study space: Having a specific area for studying can help you focus and stay motivated.
2. Use mnemonic devices: These are memory techniques that can help you remember information by associating it with something you already know or have in mind.
3. Break your study sessions into smaller intervals: This can help maintain your concentration and prevent burnout.
Architecture
The layout follows Qwen3.5-style hybrid models: three linear-attention layers for every full-attention layer. The plain residual stream is replaced by Block Attention Residuals.
| Component |
Details |
| Gated DeltaNet (18 layers) |
Constant-memory linear recurrent state using the delta rule, with a short 1D convolution (kernel 4) and output gating. About 10.5M parameters per layer. |
| Gated Attention (6 layers, every 4th) |
Grouped-query attention: 16 query heads × 128 dim, 4 shared KV heads, QK-Norm, RoPE (base 10,000), and a sigmoid output gate. About 7.3M parameters per layer. Only these 6 layers keep a KV cache. |
| SwiGLU MLP (all 24 layers) |
Hidden width 3,584, about 11.0M parameters per layer. |
| Block Attention Residuals |
8 blocks of 6 sub-layers; each sub-layer attends over summaries of earlier blocks with a learned query. About 50K parameters in total. |
| Normalization / bias |
Pre-RMSNorm without learnable scale; no bias terms anywhere. |
| Embeddings / head |
Untied: 33.6M input embedding + 33.6M output head. |
| Logit soft-cap |
logits = 15 · tanh(logits / 15) |
| Tokenizer |
Byte-level BPE, 32,768 vocab, GPT-4-style regex splitting, about 4.7 bytes per token. Trained on 2B characters of web text. |
Parameter split: about 498.1M in the transformer body (24 layers) and 67.1M in embeddings and output head.
Training
| Stage |
Data |
| 1. Pretraining |
Enhanced FineWeb text |
| 2. Midtraining |
50% unseen FineWeb, 50% conversations |
| 3. Supervised fine-tuning (SFT) |
| 4. Preference tuning |
| 5. RL with verifiable rewards (GRPO) |
Pretraining details:
- Optimizers: Muon for the ~497M 2D matrix weights (LR 0.02, cautious weight decay 0.07 on a cosine schedule); AdamW (LR 0.001) for embeddings, output head, gates, and Block Attention Residual queries
- Schedule: 40-step warmup, constant, then linear decay to 5% over the last 65% of training
- Precision / compute: bf16, torch.compile, gradient clipping at 1.0, about 17.1 GB VRAM on a single RTX 4090
Evaluation
Pretrained base model
Zero-shot evaluation with lighteval in cloze form (accuracy normalized by length; Winogrande and LAMBADA use plain accuracy). The SmolLM2-360M column is from its model card, which reports the same style of evaluation. SmolLM2-360M saw about 4 trillion training tokens, roughly 400× more than LibraMind.
| Benchmark |
LibraMind base (10.2B tokens) |
SmolLM2-360M (4T tokens) |
| HellaSwag |
55.3 |
54.5 |
| ARC (Easy / Challenge) |
65.6 / 35.4 (average 50.5) |
average 53.0 |
| PIQA |
73.9 |
71.7 |
| OpenBookQA |
31.2 |
37.4 |
| CommonsenseQA |
40.6 |
38.0 |
| Social IQa |
44.6 |
— |
| Winogrande |
56.6 |
52.5 |
| MMLU (cloze) |
30.3 |
35.8 |
| LAMBADA (accuracy / perplexity) |
46.3 / 13.3 |
— |
Chat model, by training stage
- IFEval: 541 prompts, using the lm-eval implementation.
- Tool calling: 200 held-out function-calling conversations from the SFT sources. "Call decision" means it correctly chose whether to call a tool. Numbers in parentheses use grammar-forced (constrained) decoding.
- Validation loss: 1,800 held-out conversations.
|
mid |
SFT |
preference |
RL (released) |
| IFEval prompt-level strict / loose |
40.7 / 42.3 |
45.8 / 49.7 |
50.5 / 53.8 |
51.4 / 54.5 |
| IFEval instruction-level strict / loose |
54.1 / 56.2 |
57.6 / 61.0 |
63.0 / 66.0 |
64.0 / 66.7 |
| Tools: correct call decision |
99.5 |
99.0 |
99.5 |
99.5 |
| Tools: valid JSON |
99.0 |
98.5 |
99.5 |
99.5 |
| Tools: right function (grammar-forced) |
87.8 (88.8) |
91.3 (92.9) |
92.4 (92.9) |
92.4 (92.4) |
| Tools: exact arguments |
71.4 |
80.1 |
81.1 |
80.6 |
| Validation loss, conversations |
1.035 |
0.950 |
0.966 |
0.968 |
On 40 held-out role-play prompts, the released model never described itself as an AI. Its average reply on general prompts is 156 words (175 after SFT).
IFEval caveat: these scores flatter the model's general instruction-following. Its training data includes IFEval-style constraint exercises, and the RL stage rewarded the same 19 instruction checkers that IFEval scores (on different prompts). For reference, SmolLM2-360M-Instruct reports 41.0 on IFEval (the average of its prompt- and instruction-level scores). LibraMind's equivalent strict average is 57.7.
Limitations and Bias
- Facts: it often states wrong or invented facts confidently. For example, it described the sitcom Last of the Summer Wine as a science-fiction series. Don't use it as a source of information.
- Math: it fails at arithmetic beyond single digits (see above). Part of the cause is the tokenizer, which splits numbers into irregular 1–3-digit chunks (6340 → 6|34|0, 12452 → 124|52). Do calculations in code, or give the model a calculator tool.
- Code: generated code is rarely correct.
- Using provided information: it usually uses facts from the system prompt, but can garble them, especially numbers. In game use, validate anything that matters.
- Roleplay: it stays in character but can drift from the character's details, hedge ("my expertise lies elsewhere, but…") or invent backstory. Long conversations lose coherence, and the context limit is 4,096 tokens.
- Language: English only.
- Safety: there was no dedicated safety training or red-teaming. Any refusal behavior comes only from the fine-tuning and preference data. Like any web-trained model, it can produce biased, offensive or harmful text. Filter outputs before showing them to players or the public.
Training Data and Licenses
- FineWeb: released under the permissive ODC-By 1.0 (Open Data Commons Attribution) license.
- SmolTalk, smol-smoltalk, and SmolTalk2: Apache 2.0
- GSM8K and MMLU: MIT.
License
The model weights and code are released under the Apache License 2.0; see the LICENSE file. Apache-2.0 includes a disclaimer of warranty and a limitation of liability (Sections 7 and 8). Training data carries its own terms; see above.
Acknowledgements
LibraMind builds on public research and open tools:
- Gated DeltaNet (Yang, Kautz and Hatamizadeh, 2024) and the flash-linear-attention library
- The Qwen3.5 hybrid layout and gated attention
- Block Attention Residuals from the Kimi team at Moonshot AI
- The Muon optimizer (Keller Jordan, 2024)
- Anchored Preference Optimization (D'Oosterlinck et al., 2024)
- GRPO (DeepSeekMath, 2024)
- IFEval (Zhou et al., 2023)
- lighteval
- The SmolLM2/SmolLM3 data and recipes from Hugging Face, Tulu 3 from Ai2, GSM8K, and MMLU
Citation
@misc{alby13_libramind_2026,
author = {alby13},
title = {LibraMind: A 565M Parameter Hybrid Gated DeltaNet-Attention Chat Model},
year = {2026},
publisher = {Local AI Research},
url = {https://github.com/alby13/LibraMind-AI}
}
Contact
GitHub issues discussions or Hugging Face Discussions tab.