SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Ru-Small-Instruct

by LongTime longtimedevs/Ru-Small-Instruct

Ru-Small-Instruct — экспериментальная компактная русскоязычная языковая модель класса SLM (Small Language Model) с объемом параметров ~0.2B (~165M).

Parameters165M
Context512
Weights330.4 MB
Licensemit
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve Ru-Small-Instruct (165M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By LongTime, published under mit, revision 9f2174d24448.

Ru-Small-Instruct — экспериментальная компактная русскоязычная языковая модель класса SLM (Small Language Model) с объемом параметров ~0.2B (~165M). Разработана с упором на суверенность весов (Zero-Fingerprint): модель обучена с нуля без заимствования базовых чекпоинтов у сторонних корпоративных сетей (Llama 3 от Meta, Qwen от Alibaba, Mistral). Модель предназначена для исследований локального инференса, работы на маломощном оборудовании, CPU и мобильных чипах, где критичны нулевая задержка (Time-To-First-Token) и полная независимость весов. Для компактной модели в 165M параметров, обученной на одном домашнем GPU за 48 часов, способность держать роль, грамотно формулировать сложные термины…

Read LongTime's full model card

Ru-Small-Instruct (0.2B): Sovereign Lightweight Language Model

Project Homepage & Dev Blog


Introduction

Ru-Small-Instruct — экспериментальная компактная русскоязычная языковая модель класса SLM (Small Language Model) с объемом параметров ~0.2B (~165M). Разработана с упором на суверенность весов (Zero-Fingerprint): модель обучена с нуля без заимствования базовых чекпоинтов у сторонних корпоративных сетей (Llama 3 от Meta, Qwen от Alibaba, Mistral).

Модель предназначена для исследований локального инференса, работы на маломощном оборудовании, CPU и мобильных чипах, где критичны нулевая задержка (Time-To-First-Token) и полная независимость весов.


Текущий статус модели и честный взгляд (Limitations)

Модель находится на ранней стадии и пока далека от совершенства. При объеме всего ~0.2B параметров она может иногда выдавать неожиданные ассоциации, путать факты или уходить в галлюцинации на коротких или неоднозначных запросах.

Этому не стоит удивляться:
Модель Ru-Small-Instruct была создана с нуля всего за 2 дня на одной домашней видеокарте NVIDIA GeForce RTX 5060 Ti (16GB) — от обучения собственного BPE-токенизатора и фундаментального чтения статей Википедии до финальной диалоговой огранки (SFT).

Пример реального диалога с текущей версией:

Вы: Привет!
Бот: Здравствуйте! Чем я могу вам помочь?

Вы: Ты кто?
Бот: Я искусственный интеллект — русскоязычная языковая модель, которая позволяет компьютерам 
     понимать и интерпретировать человеческий язык. Он включает в себя использование алгоритмов 
     для обучения и прогнозирования будущих событий на основе данных.

Для компактной модели в 165M параметров, обученной на одном домашнем GPU за 48 часов, способность держать роль, грамотно формулировать сложные термины и помнить синтаксис русского языка — это отличный практический результат.


Architecture & Specifications

| Спецификация | Значение | Описание | | :--- | :---: | :--- | | **Backbone Parameters** | ~165M – 200M | ~0.2B по классификации Hugging Face | | **Training Hardware** | 1x NVIDIA RTX 5060 Ti 16GB | Обучение на домашнем железе | | **Training Time** | ~2 дня (48 часов) | Полный сквозной цикл с нуля | | **Hidden Dimension ($d_{model}$)** | 896 | Размерность латентного пространства | | **Intermediate Dimension ($d_{ff}$)** | 2560 | Размерность слоев SwiGLU MLP | | **Layers** | 16 | Глубина трансформерных блоков | | **Attention Heads (Q / KV)** | 14 / 2 | Grouped-Query Attention (GQA) | | **Context Length** | 512 | Окно внимания с поддержкой RoPE | | **Vocab Size** | 28 672 | Чистый русский ByteLevel BPE | | **Native Precision** | `bfloat16` / `float16` | Базовый формат хранения | | **Special Chat Tags** | Custom Cyrillic | Нативная разметка без утечек английских токенов |

Training Recipe

1. Pre-training (Фундаментальное чтение)

Модель прошла фундаментальное обучение с нуля на русскоязычном корпусе: * Энциклопедический блок: срез статей русской Википедии (естественные науки, география, история, технологии). * Связная речь и аналитика: очищенные новостные тексты открытого корпуса «Газета». * Упаковка данных: потоковая конкатенация токенов (Data Packing) с удалением паразитных паддингов.

2. SFT & Alignment (Инструкционная доводка)

Для стабилизации ответов проведена диалоговая огранка (Instruction Tuning) с маскированием пользовательского ввода: * Инвариантность к регистру: устойчивость к строчным и заглавным запросам (привет, Привет, ПРИВЕТ). * Фактологическое якорение: очистка от ложных ассоциаций и деловых писем. * Нативный диалоговый протокол: управляющие маркеры <|юзер|>, <|юзерзакончил|>, <|я_начал|>, <|я_закончил|>.


Benchmark & Performance

Оценка производительности на локальном оборудовании (NVIDIA GeForce RTX, FP16/BF16):

| Метрика | Значение | | :--- | :---: | | **Generation Speed (RTX Desktop)** | **180 – 210 t/s** | | **Memory Footprint (Inference)** | **~380 – 490 MB** | | **Cold Start Time** | **< 0.15 s** | | **Supported Hardware** | CUDA, Apple Silicon (MPS), CPU, Edge |

Prompt Encoding

Для получения точных и связных ответов структурируйте промпт с переносами строк:

<|юзер|>
Кто такой слон?
<|юзерзакончил|>
<|я_начал|>

Генерация завершается моделью на служебном маркере <|я_закончил|>.


Minimal Inference

Запуск инференса на чистом Python через библиотеку transformers:

import torch
from transformers import PreTrainedTokenizerFast, LlamaForCausalLM

MODEL_ID = "longtimedevs/Ru-Small-Instruct"
DEVICE = "cuda:0" if torch.cuda.is_available() else "cpu"
DTYPE = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16

# 1. Загрузка модели и токенизатора
tokenizer = PreTrainedTokenizerFast.from_pretrained(MODEL_ID)
model = LlamaForCausalLM.from_pretrained(MODEL_ID, torch_dtype=DTYPE).to(DEVICE)
model.eval()

# 2. Подготовка запроса
user_query = "Привет! Кто ты?"
prompt = f"<|юзер|>\n{user_query}\n<|юзерзакончил|>\n<|я_начал|>\n"
inputs = tokenizer(prompt, return_tensors="pt").to(DEVICE)

# 3. Генерация с остановкой по токену
ai_end_id = tokenizer.convert_tokens_to_ids("<|я_закончил|>")

with torch.no_grad():
    output_tokens = model.generate(
        **inputs,
        max_new_tokens=120,
        temperature=0.25,
        top_p=0.85,
        repetition_penalty=1.15,
        eos_token_id=[ai_end_id, tokenizer.eos_token_id],
        pad_token_id=tokenizer.pad_token_id,
        do_sample=True
    )

# 4. Декодирование ответа
response = tokenizer.decode(output_tokens[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(f"Ru-Small-Instruct: {response.strip()}")

Recommended Generation Parameters

Параметр Рекомендуемое значение Описание
temperature 0.25 Минимизирует дрейф внимания на компактных весах
top_p 0.85 Отрезает маловероятные токены
repetition_penalty 1.15 Предотвращает зацикливание списков
max_new_tokens 120 – 150 Оптимальная длина реплики

License

Модель и связанные материалы распространяются под лицензией MIT License. Веса свободны для коммерческого и исследовательского применения.


Author & Project Info

Разработка, поддержка и сопутствующие проекты:
Официальный сайт автора: https://long-time.ru

Configuration

Architecture
LlamaForCausalLM
Context length (tokens)
512
Layers
16
Hidden size
896
Feed-forward size
2,560
Attention heads
14
Key/value heads
2
Head dimension
64
Vocabulary size
28,672
Model type
llama

Identity and Version

Repository
longtimedevs/Ru-Small-Instruct
Publisher
LongTime
Task
Text generation
Modality
Text
Library
transformers
Parameters
165M parameters
Languages
ru
Revision
9f2174d2444890064a3d31425ad557f3e3073f56
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

7 files, 333.4 MB in total. The weights are 1 file totalling 330.4 MB in safetensors.

Weights1 file · 330.4 MB
Configuration2 files · 979 B
Tokenizer2 files · 3.1 MB
Documentation1 file · 11.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights330.4 MB b45db05f8703
config.jsonConfiguration753 B
generation_config.jsonConfiguration226 B
README.mdDocumentation11.3 KB
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer3.1 MB
tokenizer_config.jsonTokenizer315 B

License and Download

License
mit
Access
Open weights, no gate
Download size
330.4 MB
Download from LongTime

Released by LongTime through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published330.4 MB
16-bit0.3 GB
8-bit0.2 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Ru-Small-Instruct

How much GPU memory does Ru-Small-Instruct need?

About 0.4 GB at 16-bit and 0.1 GB at 4-bit: the weights (165M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Ru-Small-Instruct on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Ru-Small-Instruct commercially?

Yes. Ru-Small-Instruct is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is Ru-Small-Instruct's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

CasualSwarms

Convergent Intelligence

SAGI is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. The enhancements were integrated with the existing AGI system through: 1. Compatibility Layer: Ensuring new components work with existing AGI Core 2. Unified State Representation: Combining enhanced capabilities with existing state 3. Enhanced Continuous Learning: Upgrading the learning system with new capabilities 4. Performance Monitoring: Tracking improvements through validation systems - Successfully…

Open weights apache-2.0 170M parameters 1,024 tokens transformers

Model · Text generation

Haidass-Translate-143M

DALab

English | 中文 A 143M-parameter bidirectional Chinese↔English translation model, instruction-tuned on the Haidass1.5-143M base — the strongest zh⇄en translator at this scale among general chat-architecture models. Drafter-143M: a control model with identical configuration, data and training recipe, except that it starts from random initialization instead of the pretrained base — used to quantify the contribution of base-model pretraining. OPUS-MT models are single-directional — one independent 78M model per direction; "-" marks directions a model does not serve. The same models re-evaluated on FLORES+ devtest (released 2026; zero overlap with dev): devtest sentences do not overlap with dev.…

Open weights apache-2.0 143M parameters 4,096 tokens

English | 中文 The instruction-tolerant sibling of DALabCommunity/Haidass-Translate-143M: same 143M zh⇄en translation training, plus 9.1% cleaned general-domain data (STEPFUN ShareGPT) mixed in. Translation scores are within 0.3 BLEU of the pure-translation version, and the model retains limited general instruction-following ability that the pure-translation version does not have. Drafter-143M: a control model with identical configuration, data and training recipe, except that it starts from random initialization instead of the pretrained base — used to quantify the contribution of base-model pretraining. OPUS-MT models are single-directional — one independent 78M model per direction; "-"…

Open weights apache-2.0 143M parameters 4,096 tokens

Model · Text generation

gpt2

OpenAI community

Test the whole generation capabilities here: https://transformer.huggingface.co/doc/gpt2-large Pretrained model on English language using a causal language modeling (CLM) objective. It was introduced in and first released at this page. model. Content from this model card has been written by the Hugging Face team to complete the information they provided and give specific examples of bias. GPT-2 is a transformers model pretrained on a very large corpus of English data in a self-supervised fashion. This means it was pretrained on the raw texts only, with no humans labelling them in any way (which is why it can use lots of publicly available data) with an automatic process to generate inputs…

Open weights mit 137M parameters transformers

SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters. They are capable of solving a wide range of tasks while being lightweight enough to run on-device. More details in our paper: https://arxiv.org/abs/2502.02737 SmolLM2 demonstrates significant advances over its predecessor SmolLM1, particularly in instruction following, knowledge, reasoning. The 135M model was trained on 2 trillion tokens using a diverse dataset combination: FineWeb-Edu, DCLM, The Stack, along with new filtered datasets we curated and will release soon. We developed the instruct version through supervised fine-tuning (SFT) using a combination of public datasets and our…

Open weights apache-2.0 135M parameters 8,192 tokens transformers

SmolLM2 is a family of compact language models available in three size: 135M, 360M, and 1.7B parameters. They are capable of solving a wide range of tasks while being lightweight enough to run on-device. More details in our paper https://arxiv.org/abs/2502.02737 SmolLM2 demonstrates significant advances over its predecessor SmolLM1, particularly in instruction following, knowledge, reasoning. The 135M model was trained on 2 trillion tokens using a diverse dataset combination: FineWeb-Edu, DCLM, The Stack, along with new filtered datasets we curated and will release soon. We developed the instruct version through supervised fine-tuning (SFT) using a combination of public datasets and our own…

Open weights apache-2.0 135M parameters 8,192 tokens transformers