SAVRN
Search Contact SAVRN

Open-weight model · Text generation

LibraMindMini

by Alby13 alby13/LibraMindMini

LibraMindMini is an open-weight model for text generation from Alby13, released under Apache License 2.0. It has 565M parameters. At 16-bit it needs about 1.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

LibraMind is a 565M-parameter chat model pretrained from scratch by alby13 on a single consumer GPU (NVIDIA RTX 4090).

Parameters565M
Context—
Weights1.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve LibraMindMini (565M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1.1 GB 1.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.6 GB 0.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.3 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

LibraMindMini on every accelerator the SAVRN Index prices, at every precision

Model Card

By Alby13, published under apache-2.0, revision bcda93862829.

LibraMind is a 565M-parameter chat model pretrained from scratch by alby13 on a single consumer GPU (NVIDIA RTX 4090). It is a hybrid architecture that interleaves 18 Gated DeltaNet linear-recurrent layers with 6 Gated Attention layers (a repeating DDDA × 6 pattern), with Block Attention Residuals and logit soft-capping, trained with the Muon optimizer. It went through a full five-stage pipeline: pretraining, midtraining, supervised fine-tuning, preference tuning, and RL with verifiable rewards. - Research on small hybrid recurrent/attention language models - Studying low-resource, single-GPU training pipelines (pretraining through RL) and the Muon optimizer at small scale - Experiments…

Read Alby13's full model card

LibraMind (565M parameters)

LibraMind is a 565M-parameter chat model pretrained from scratch by alby13 on a single consumer GPU (NVIDIA RTX 4090). It is a hybrid architecture that interleaves 18 Gated DeltaNet linear-recurrent layers with 6 Gated Attention layers (a repeating DDDA × 6 pattern), with Block Attention Residuals and logit soft-capping, trained with the Muon optimizer. It went through a full five-stage pipeline: pretraining, midtraining, supervised fine-tuning, preference tuning, and RL with verifiable rewards.

Research model: no safety training. LibraMind has not undergone safety training or red-teaming. It can produce inaccurate, biased, offensive, or harmful text, and it will not reliably refuse requests. It also makes factual and arithmetic mistakes (see the examples below). It is released as a research artifact, provided "as is" without warranty of any kind, and the authors accept no liability for its outputs or for any use of the model. You are solely responsible for evaluating it and for any safeguards, filtering, and human oversight your use case needs. Do not deploy it in user-facing, high-stakes, or safety-critical settings without your own safety work. This notice is guidance. It does not add to or modify the terms of the Apache-2.0 license.

Model Summary

Developer alby13
Parameters 565M (565.2M)
Model type Chat model
Architecture Hybrid Gated DeltaNet + Gated Attention, decoder-only
Layers / width 24 layers, hidden size 1,024 (DDDA × 6: 3 DeltaNet layers then 1 Attention layer, repeated 6 times)
Vocabulary 32,768 (custom byte-level BPE)
Context length 4,096 tokens (pretrained at 2,048, extended during midtraining)
Weights precision bf16 (about 1.1 GB)
Language English
License Apache-2.0
Release 10/6/2026

Intended use

  • Light assistant tasks: casual conversation, rewriting and summarizing short texts, formatted answers such as lists, sections and word limits, and simple function calling.
  • Hobby and research: a fully documented small-model training run (architecture, data, every stage, evaluations).
  • Game NPCs and interactive characters: dialogue driven by a character card in the system prompt, with optional schema-forced JSON for emotions and actions.

Intended Use

Intended for: - Research on small hybrid recurrent/attention language models - Studying low-resource, single-GPU training pipelines (pretraining through RL) and the Muon optimizer at small scale - Experiments with on-device chat, role-play, and structured output (JSON / tool calls), with your own validation of the results (for example, validate JSON and compute numbers in code)

Not intended for: - High-stakes or safety-critical decisions (medical, legal, financial, etc.) - Unsupervised, user-facing deployment: the model has no safety training - Use as a source of factual information without independent verification

How to Use

LibraMind uses a custom architecture, so it doesn't load in transformers, llama.cpp or other GGUF runtimes. It needs this repository's model.py, inference.py and chat_format.py, plus:

PyTorch with CUDA (trained with 2.14) flash-linear-attention 0.5.2, with Triton tokenizers llguidance (optional, for grammar-forced JSON) In this repository, the weights are runs/rl1/model_final.pt (float32, 2.26 GB) and the tokenizer is data/chat/sft/tokenizer.json.

import torch
from tokenizers import Tokenizer

from chat_format import ChatEncoder
from inference import GrammarFactory, generate
from model import LM, ModelConfig

ck = torch.load("model_final.pt", map_location="cpu", weights_only=False)
model = LM(ModelConfig(**{**ck["model_config"], "grad_ckpt": False}))
model.load_state_dict(ck["model"])
model = model.cuda().to(torch.bfloat16).eval()
enc = ChatEncoder(Tokenizer.from_file("tokenizer.json"))


def chat(messages, schema=None, max_new=300):
    matcher = [GrammarFactory("tokenizer.json", enc.im_end, model.config.vocab_size).json(schema)] if schema else None
    with torch.autocast("cuda", dtype=torch.bfloat16):
        out = generate(model, [enc.encode_prompt(messages)], max_new, temperature=0.7, top_p=0.9, rep_penalty=1.1,
                       stop_ids=[enc.im_end], matchers=matcher, vocab=enc.tok.get_vocab_size())[0]
    return enc.tok.decode(out)


print(chat([{"role": "system", "content": "You are Brom, a gruff dwarven blacksmith. Stay in character."},
            {"role": "user", "content": "Can you fix my sword?"}]))

generate uses cached decoding: Gated DeltaNet recurrent state plus an attention KV cache, replayed as a CUDA graph. It accepts many prompts at once, at about 1,050 tokens/s across 64 parallel conversations. The repository also has a browser chat window (chatui.cmd) and a console chat (chat.cmd).

Recommended settings:

  • temperature 0.4, top-p 0.9, repetition penalty 1.1
  • System prompt:
"You are LibraMind, a friendly and thoughtful conversation partner made by alby13. Keep replies natural and to the point: a few sentences for casual chat, more detail only when it's needed. Ask a follow-up question when it helps the conversation.
  • temperature 0 for JSON and tool calls
  • always pass stop_ids=[<|im_end|>]

Requirements:

Memory: about 1.1 GB for the weights in bf16, plus activations. The recurrent layers use constant-size state, and only the 6 attention layers grow a KV cache with context length.

Chat format

ChatML with a beginning-of-text token. Roles are system, user, assistant and tool:

<|bos|><|im_start|>system
{system prompt}<|im_end|>
<|im_start|>user
{message}<|im_end|>
<|im_start|>assistant
{reply}<|im_end|>

The system prompt is optional. Without one, the model behaves as the LibraMind assistant.

Characters and Personas

Put the character card in the system prompt: name, setting, personality, how they speak, and any facts they must get right. That means prices, quest details, names and places. Roleplay training used this pattern:

You are {name}. Stay in character: speak as {name} would, in the first person, and never say you are an AI.

Character profile:
{profile}

For game actions, force the reply into a JSON schema with GrammarFactory.json(schema). The output is then guaranteed to parse and match the schema. Example schema: examples/npc_schema.json, with fields line, emotion and action. Describe the fields in the system prompt too, so the content makes sense as well as the format.

Function calling

List the tools in the system prompt in the format the model was trained on:

You can call the following functions:
<tools>
[{"type": "function", "function": {"name": "get_weather", "description": "...", "parameters": {...}}}]
</tools>

To call a function, reply with a JSON object inside <tool_call></tool_call> tags, for example:
<tool_call>
{"name": "function_name", "arguments": {"arg": "value"}}
</tool_call>
You may make several calls. Function results come back in <tool_response> tags.

GrammarFactory.tool_calls(tools) can force calls to name a listed function, with arguments that match its schema.

Example Outputs

sampling settings used (temperature 0.4, top-p 0.9, repetition penalty 1.1).

General chat

User: Hi! Who are you? LibraMind: I'm LibraMind, a small AI language model trained from scratch by alby13. How can I help?

Role-play

System Prompt: You are Grom, a gruff dwarven blacksmith in the mountain town of Bringham. Stay in character and keep replies short. User: Where can I find some armor? I need some that also protects against magic. LibraMind: Well, I ain't no scholar, but I've heard tales from travelers. You might find some armor in the blacksmith's forge. Ask the old sage, Thorne. He knows all about such things.

System Prompt: You are Jin, the stern innkeeper of the Gilded Redbird inn in the village of Ashfield. A room costs 15 silver a night and a bowl of stew costs 3 silver. (...) User: How much for a room tonight and some supper? LibraMind: A room for 15 silver is a good price. A bowl of stew for 3 silver will be 9 silver.

Structured output (schema-forced decoding; the system prompt asks for JSON with line, emotion, and action)

System: You are Grom (...). Reply only with JSON: "line" is what you say, "emotion" is how you feel, and "action" is what you do (open_shop if the player wants to buy or get something repaired). User: I'd like to buy a new axe.

Sample 1: {"line": "axe", "emotion": "happy", "action": "open_shop"}
Sample 2: {"line": "You've come to the right place, friend. I've got just the thing for you. What's your name, and what kind of axe are you looking for?", "emotion": "neutral", "action": "open_shop"}

Instruction following

User: Give me three tips for memorizing for studying as a numbered list. LibraMind: 1. Create a dedicated study space: Having a specific area for studying can help you focus and stay motivated. 2. Use mnemonic devices: These are memory techniques that can help you remember information by associating it with something you already know or have in mind. 3. Break your study sessions into smaller intervals: This can help maintain your concentration and prevent burnout.

Architecture

The layout follows Qwen3.5-style hybrid models: three linear-attention layers for every full-attention layer. The plain residual stream is replaced by Block Attention Residuals.

Component Details
Gated DeltaNet (18 layers) Constant-memory linear recurrent state using the delta rule, with a short 1D convolution (kernel 4) and output gating. About 10.5M parameters per layer.
Gated Attention (6 layers, every 4th) Grouped-query attention: 16 query heads × 128 dim, 4 shared KV heads, QK-Norm, RoPE (base 10,000), and a sigmoid output gate. About 7.3M parameters per layer. Only these 6 layers keep a KV cache.
SwiGLU MLP (all 24 layers) Hidden width 3,584, about 11.0M parameters per layer.
Block Attention Residuals 8 blocks of 6 sub-layers; each sub-layer attends over summaries of earlier blocks with a learned query. About 50K parameters in total.
Normalization / bias Pre-RMSNorm without learnable scale; no bias terms anywhere.
Embeddings / head Untied: 33.6M input embedding + 33.6M output head.
Logit soft-cap logits = 15 · tanh(logits / 15)
Tokenizer Byte-level BPE, 32,768 vocab, GPT-4-style regex splitting, about 4.7 bytes per token. Trained on 2B characters of web text.

Parameter split: about 498.1M in the transformer body (24 layers) and 67.1M in embeddings and output head.

Training

Stage Data
1. Pretraining Enhanced FineWeb text
2. Midtraining 50% unseen FineWeb, 50% conversations
3. Supervised fine-tuning (SFT)
4. Preference tuning
5. RL with verifiable rewards (GRPO)

Pretraining details: - Optimizers: Muon for the ~497M 2D matrix weights (LR 0.02, cautious weight decay 0.07 on a cosine schedule); AdamW (LR 0.001) for embeddings, output head, gates, and Block Attention Residual queries - Schedule: 40-step warmup, constant, then linear decay to 5% over the last 65% of training - Precision / compute: bf16, torch.compile, gradient clipping at 1.0, about 17.1 GB VRAM on a single RTX 4090

Evaluation

Pretrained base model

Zero-shot evaluation with lighteval in cloze form (accuracy normalized by length; Winogrande and LAMBADA use plain accuracy). The SmolLM2-360M column is from its model card, which reports the same style of evaluation. SmolLM2-360M saw about 4 trillion training tokens, roughly 400× more than LibraMind.

Benchmark LibraMind base (10.2B tokens) SmolLM2-360M (4T tokens)
HellaSwag 55.3 54.5
ARC (Easy / Challenge) 65.6 / 35.4 (average 50.5) average 53.0
PIQA 73.9 71.7
OpenBookQA 31.2 37.4
CommonsenseQA 40.6 38.0
Social IQa 44.6 —
Winogrande 56.6 52.5
MMLU (cloze) 30.3 35.8
LAMBADA (accuracy / perplexity) 46.3 / 13.3 —

Chat model, by training stage

  • IFEval: 541 prompts, using the lm-eval implementation.
  • Tool calling: 200 held-out function-calling conversations from the SFT sources. "Call decision" means it correctly chose whether to call a tool. Numbers in parentheses use grammar-forced (constrained) decoding.
  • Validation loss: 1,800 held-out conversations.
mid SFT preference RL (released)
IFEval prompt-level strict / loose 40.7 / 42.3 45.8 / 49.7 50.5 / 53.8 51.4 / 54.5
IFEval instruction-level strict / loose 54.1 / 56.2 57.6 / 61.0 63.0 / 66.0 64.0 / 66.7
Tools: correct call decision 99.5 99.0 99.5 99.5
Tools: valid JSON 99.0 98.5 99.5 99.5
Tools: right function (grammar-forced) 87.8 (88.8) 91.3 (92.9) 92.4 (92.9) 92.4 (92.4)
Tools: exact arguments 71.4 80.1 81.1 80.6
Validation loss, conversations 1.035 0.950 0.966 0.968

On 40 held-out role-play prompts, the released model never described itself as an AI. Its average reply on general prompts is 156 words (175 after SFT).

IFEval caveat: these scores flatter the model's general instruction-following. Its training data includes IFEval-style constraint exercises, and the RL stage rewarded the same 19 instruction checkers that IFEval scores (on different prompts). For reference, SmolLM2-360M-Instruct reports 41.0 on IFEval (the average of its prompt- and instruction-level scores). LibraMind's equivalent strict average is 57.7.

Limitations and Bias

  • Facts: it often states wrong or invented facts confidently. For example, it described the sitcom Last of the Summer Wine as a science-fiction series. Don't use it as a source of information.
  • Math: it fails at arithmetic beyond single digits (see above). Part of the cause is the tokenizer, which splits numbers into irregular 1–3-digit chunks (6340 → 6|34|0, 12452 → 124|52). Do calculations in code, or give the model a calculator tool.
  • Code: generated code is rarely correct.
  • Using provided information: it usually uses facts from the system prompt, but can garble them, especially numbers. In game use, validate anything that matters.
  • Roleplay: it stays in character but can drift from the character's details, hedge ("my expertise lies elsewhere, but…") or invent backstory. Long conversations lose coherence, and the context limit is 4,096 tokens.
  • Language: English only.
  • Safety: there was no dedicated safety training or red-teaming. Any refusal behavior comes only from the fine-tuning and preference data. Like any web-trained model, it can produce biased, offensive or harmful text. Filter outputs before showing them to players or the public.

Training Data and Licenses

  • FineWeb: released under the permissive ODC-By 1.0 (Open Data Commons Attribution) license.
  • SmolTalk, smol-smoltalk, and SmolTalk2: Apache 2.0
  • GSM8K and MMLU: MIT.

License

The model weights and code are released under the Apache License 2.0; see the LICENSE file. Apache-2.0 includes a disclaimer of warranty and a limitation of liability (Sections 7 and 8). Training data carries its own terms; see above.

Acknowledgements

LibraMind builds on public research and open tools:

  • Gated DeltaNet (Yang, Kautz and Hatamizadeh, 2024) and the flash-linear-attention library
  • The Qwen3.5 hybrid layout and gated attention
  • Block Attention Residuals from the Kimi team at Moonshot AI
  • The Muon optimizer (Keller Jordan, 2024)
  • Anchored Preference Optimization (D'Oosterlinck et al., 2024)
  • GRPO (DeepSeekMath, 2024)
  • IFEval (Zhou et al., 2023)
  • lighteval
  • The SmolLM2/SmolLM3 data and recipes from Hugging Face, Tulu 3 from Ai2, GSM8K, and MMLU

Citation

@misc{alby13_libramind_2026,
  author    = {alby13},
  title     = {LibraMind: A 565M Parameter Hybrid Gated DeltaNet-Attention Chat Model},
  year      = {2026},
  publisher = {Local AI Research},
  url       = {https://github.com/alby13/LibraMind-AI}
}

Contact

GitHub issues discussions or Hugging Face Discussions tab.

Configuration

Head dimension
128
Vocabulary size
32,832
RoPE base
10000

Identity and Version

Repository
alby13/LibraMindMini
Publisher
Alby13
Task
Text generation
Modality
Text
Library
pytorch
Parameters
565M parameters
Languages
en
Revision
bcda93862829cbc71fab7c035c8f9166af5272aa
First published
2026-10-07
Last updated
2026-10-07

Files and Weights

8 files, 1.1 GB in total. The weights are 1 file totalling 1.1 GB in safetensors.

Weights1 file · 1.1 GB
Configuration4 files · 32.4 KB
Tokenizer1 file · 2.3 MB
Documentation1 file · 16.8 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.1 GB 2e432bee4331
chat_format.pyConfiguration2.7 KB —
config.jsonConfiguration343 B —
inference.pyConfiguration17.3 KB —
model.pyConfiguration12.0 KB —
README.mdDocumentation16.8 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer2.3 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.1 GB
Download from Alby13

Released by Alby13 through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published1.1 GB
16-bit1.1 GB
8-bit0.6 GB
4-bit0.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About LibraMindMini

How much GPU memory does LibraMindMini need?

About 1.4 GB at 16-bit and 0.3 GB at 4-bit: the weights (565M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run LibraMindMini on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use LibraMindMini commercially?

Yes. LibraMindMini is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text generation

bloomz-560m

BigScience Workshop

We recommend using the model to perform tasks expressed in natural language. For example, given the prompt "Translate to English: Je t’aime.", the model will most likely answer "I love you.". Some prompt ideas from our paper: - Suggest at least five related search terms to "Mạng neural nhân tạo". - Write a fairy tale about a troll saving a princess from a dangerous dragon. The fairy tale is a masterpiece that has achieved praise worldwide and its moral is "Heroes Come in All Shapes and Sizes". Story (in Spanish): - Explain in a sentence in Telugu what is backpropagation in neural networks. Feel free to share your generations in the Community tab! Prompt Engineering: The performance may vary…

Open weights bigscience-bloom-rail-1.0 559M parameters transformers

Model · Text generation

Qwen3-0.6B-Base

Qwen

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Building upon extensive advancements in training data, model architecture, and optimization techniques, Qwen3 delivers the following key improvements over the previously released Qwen2.5: Qwen3-0.6B-Base has the following features: For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our blog, GitHub, and Documentation. The code of Qwen3 has been in the latest Hugging Face transformers and we advise you to use the latest version of transformers. With transformers<4.51.0, you will…

Open weights apache-2.0 596M parameters 32,768 tokens transformers

Model · Text generation

Symbiotic-1B

Convergent Intelligence

Purpose: Lightweight, memory-augmented reasoning model for CPU and embedded inference SymbioticLM-1B is the compact version of the SymbioticAI architecture. It fuses Qwen’s rotary transformer design with a symbolic processing pipeline and a persistent episodic memory. Though smaller in parameter count, it retains the full cognitive engine: symbolic memory, dynamic thought evolution, and entropy-gated control. This model is ideal for symbolic reasoning in constrained environments — like research agents, lightweight assistants, and memory-efficient logical processing. - Procedural planning, math modeling, small-code generation - Less fluent in free-form language than larger variants…

Open weights afl-3.0 596M parameters 40,960 tokens transformers

This model is a fine-tuned version of Qwen/Qwen3-0.6B on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 2e-05 - trainbatchsize: 4 - evalbatchsize: 8 - gradientaccumulationsteps: 16 - totaltrainbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 100 - numepochs: 3 - Transformers 5.17.0 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.23.2

Open weights apache-2.0 596M parameters 40,960 tokens transformers

Model · Text generation

S1-mini-MLX-4bit

David Larrea

Converted with mlx-lm 0.31.3. A six-case deterministic sanity check matched three upstream reference strings exactly. The remaining differences included a retained leading “So,” punctuation/ordinal variation, and omission of “tomorrow” in one correction case. This is a small functional check, not the upstream 7,519-case evaluation; assess the 4-bit build on your own transcripts. A 0.6B-parameter text normalizer for speech-to-text output. It takes a raw ASR transcript and rewrites it as clean written text: fillers removed, false starts and self-corrections resolved to the value the speaker landed on, punctuation and capitalization applied, and spoken numbers, dates, times, currency and email…

Open weights other 596M parameters 40,960 tokens mlx

Model · Text generation

S1-mini-MLX-8bit

David Larrea

Converted with mlx-lm 0.31.3. A six-case deterministic sanity check matched four upstream reference strings exactly. The two differences were a retained leading “So,” and one comma variation. This is a small functional check, not the upstream 7,519-case evaluation. A 0.6B-parameter text normalizer for speech-to-text output. It takes a raw ASR transcript and rewrites it as clean written text: fillers removed, false starts and self-corrections resolved to the value the speaker landed on, punctuation and capitalization applied, and spoken numbers, dates, times, currency and email addresses rendered in written form. On a held-out set of 7,519 English cases it reaches 94.8% token accuracy, and…

Open weights other 596M parameters 40,960 tokens mlx