SAVRN
Search Contact SAVRN

Open-weight model · Text generation

slm-125m-dpo

by Tijani Ohiokpehai tohio/slm-125m-dpo

slm-125m-dpo is an open-weight model for text generation from Tijani Ohiokpehai, released under MIT License. It has 165M parameters and a 1,024-token context. At 16-bit it needs about 0.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

tohio/slm-125m-dpo is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine.

Parameters165M
Context1,024
Weights330.8 MB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve slm-125m-dpo (165M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

slm-125m-dpo on every accelerator the SAVRN Index prices, at every precision

Model Card

By Tijani Ohiokpehai, published under mit, revision 287c9df2e806.

tohio/slm-125m-dpo is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Read Tijani Ohiokpehai's full model card

tohio/slm-125m-dpo is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine.

Architecture Highlights

  • Decoder-Only Transformer with Rotary Position Embeddings (RoPE)
  • Attention: Grouped-Query Attention (GQA, 12:4 ratio)
  • Activation: SwiGLU Feed-Forward Network ($d_{ffn} = 2048$)
  • Normalization: Bias-free Pre-LayerNorm / RMSNorm
  • Weight Tying: Tied input embedding (embed_tokens) and output head (lm_head)
  • Hugging Face Native: Directly compatible with LlamaForCausalLM

Pre-training Curriculum

The base model was pre-trained across an interleaved multi-source domain mixture composed of:

  • FineWeb-Edu
  • DCLM-Edu
  • The Stack-Edu
  • NuminaMath-CoT
  • OpenMathReasoning
  • SLM-Synthetic-Pretrain

Alignment Lineage

  1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards.
  2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignore_index=-100).
  3. Direct Preference Optimization (DPO): Single-stage pairwise preference alignment optimizing chosen vs. rejected generations.

Prompt Format (ChatML)

This model adheres strictly to the ChatML template:

<|im_start|>system
You are a helpful AI assistant.<|im_end|>
<|im_start|>user
Write a Python script to compute the Fibonacci sequence efficiently.<|im_end|>
<|im_start|>assistant

Quick Start via transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "tohio/slm-125m-dpo"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {"role": "system", "content": "You are a concise AI assistant."},
    {"role": "user", "content": "Explain quantum computing in one sentence."},
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=100, temperature=0.7, top_p=0.9)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Trained and published autonomously via slm-gpt.

Configuration

Architecture
LlamaForCausalLM
Context length (tokens)
1,024
Layers
14
Hidden size
768
Feed-forward size
2,048
Attention heads
12
Key/value heads
4
Vocabulary size
50,304
RoPE base
10000
Stored precision
bfloat16
Model type
llama

Identity and Version

Repository
tohio/slm-125m-dpo
Publisher
Tijani Ohiokpehai
Task
Text generation
Modality
Text
Library
transformers
Parameters
165M parameters
Languages
en
Revision
287c9df2e806a28ed7c24a568966b344ce80ccd3
First published
2026-10-01
Last updated
2026-10-01

Files and Weights

6 files, 330.8 MB in total. The weights are 1 file totalling 330.8 MB in safetensors.

Weights1 file · 330.8 MB
Configuration2 files · 882 B
Tokenizer1 file · 1.1 KB
Documentation1 file · 2.7 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights330.8 MB babc47f7efbb
config.jsonConfiguration641 B —
generation_config.jsonConfiguration241 B —
README.mdDocumentation2.7 KB —
.gitattributesRepository1.5 KB —
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
330.8 MB
Download from Tijani Ohiokpehai

Released by Tijani Ohiokpehai through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published330.8 MB
16-bit0.3 GB
8-bit0.2 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About slm-125m-dpo

How much GPU memory does slm-125m-dpo need?

About 0.4 GB at 16-bit and 0.1 GB at 4-bit: the weights (165M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run slm-125m-dpo on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use slm-125m-dpo commercially?

Yes. slm-125m-dpo is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is slm-125m-dpo's context length?

1,024 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

slm-125m-sft

Tijani Ohiokpehai

tohio/slm-125m-sft is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 1,024 tokens transformers

Model · Text generation

slm-125m-base

Tijani Ohiokpehai

tohio/slm-125m-base is a small language model (125M parameters) built from scratch and aligned end-to-end using the slm-gpt engine. - Decoder-Only Transformer with Rotary Position Embeddings (RoPE) The base model was pre-trained across an interleaved multi-source domain mixture composed of: - FineWeb-Edu - DCLM-Edu - The Stack-Edu - NuminaMath-CoT - OpenMathReasoning - SLM-Synthetic-Pretrain 1. Pre-training: Multi-GPU Distributed Data Parallel (DDP) over rank-disjoint memory-mapped token shards. 2. Supervised Fine-Tuning (SFT): Full parameter instruction-tuning on ChatML formatted dialogues with prompt loss masking (ignoreindex=-100). 3. Direct Preference Optimization (DPO): Single-stage…

Open weights mit 165M parameters 2,048 tokens transformers

Model · Text generation

Ru-Small-Instruct

LongTime

Ru-Small-Instruct — экспериментальная компактная русскоязычная языковая модель класса SLM (Small Language Model) с объемом параметров ~0.2B (~165M). Разработана с упором на суверенность весов (Zero-Fingerprint): модель обучена с нуля без заимствования базовых чекпоинтов у сторонних корпоративных сетей (Llama 3 от Meta, Qwen от Alibaba, Mistral). Модель предназначена для исследований локального инференса, работы на маломощном оборудовании, CPU и мобильных чипах, где критичны нулевая задержка (Time-To-First-Token) и полная независимость весов. Для компактной модели в 165M параметров, обученной на одном домашнем GPU за 48 часов, способность держать роль, грамотно формулировать сложные термины…

Open weights mit 165M parameters 512 tokens transformers

This model is a fine-tuned version of qing-yao/ppt-pythia-160m-uniform250-previousmse-seed324-stage1 on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 500 - trainingsteps: 10000 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1

Open weights apache-2.0 162M parameters 2,048 tokens transformers

This model is a fine-tuned version of qing-yao/ppt-pythia-160m-uniform250-previousmse-seed324-stage1 on the None dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 0.001 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosinewithminlr - lrschedulerwarmupsteps: 500 - trainingsteps: 10000 - Transformers 5.4.0 - Pytorch 2.8.0+cu128 - Datasets 3.2.0 - Tokenizers 0.22.1

Open weights apache-2.0 162M parameters 2,048 tokens transformers

Model · Text generation

CasualSwarms

Convergent Intelligence

SAGI is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. The enhancements were integrated with the existing AGI system through: 1. Compatibility Layer: Ensuring new components work with existing AGI Core 2. Unified State Representation: Combining enhanced capabilities with existing state 3. Enhanced Continuous Learning: Upgrading the learning system with new capabilities 4. Performance Monitoring: Tracking improvements through validation systems - Successfully…

Open weights apache-2.0 170M parameters 1,024 tokens transformers