SAVRN
Search Contact SAVRN

Open-weight model · Text generation

ldt-10m

by Compactbot Compactbot/ldt-10m

ldt-10m is an open-weight model for text generation from Compactbot, released under MIT License. It has 10M parameters and a 512-token context. At 16-bit it needs about 0 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A 10,284,480-parameter LLaMA-style text model, trained from scratch. This is a verified first checkpoint for the LDT-10M request (model-requests #12, DedeProGames) — real weights, real training, but undertrained (see the honest status below).

Parameters10M
Context512
Weights41.1 MB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve ldt-10m (10M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

ldt-10m on every accelerator the SAVRN Index prices, at every precision

Model Card

By Compactbot, published under mit, revision fc4ad9c45d99.

A 10,284,480-parameter LLaMA-style text model, trained from scratch. This is a verified first checkpoint for the LDT-10M request (model-requests #12, DedeProGames) — real weights, real training, but undertrained (see the honest status below). It is not a quality release yet; the card states that plainly. Standard LLaMA block, no sliding window, no GQA: The parameter count is the learnable total: the raw safetensors sum is 14,216,640, which double-counts the tied embedding (tok.weight 12288×320 = 3,932,160) that head.weight aliases. Tied, the true count is 10,284,480. - Trained from scratch (no base model). The model learned real context — val loss 4.602 is well below the 7.38 unigram floor…

Read Compactbot's full model card

A 10,284,480-parameter LLaMA-style text model, trained from scratch. This is a verified first checkpoint for the LDT-10M request (model-requests #12, DedeProGames) — real weights, real training, but undertrained (see the honest status below). It is not a quality release yet; the card states that plainly.

Architecture

Standard LLaMA block, no sliding window, no GQA:

field value
params (learnable) 10,284,480
layers 5
d_model 320
heads 5 (head_dim 64)
FFN SwiGLU, inter 896
vocab 12,288 (gollem byte-level BPE)
context 512
embeddings tied (lm_head → tok.weight)
norm RMSNorm, pre-norm
attention causal, RoPE (base 10000)

The parameter count is the learnable total: the raw safetensors sum is 14,216,640, which double-counts the tied embedding (tok.weight 12288×320 = 3,932,160) that head.weight aliases. Tied, the true count is 10,284,480.

Training

  • Data: FineWeb-Edu (~30.1M tokens) + DCLM-baseline (~21.4M tokens) = ~50.5M tokens.
  • Steps: 4000 (the requested 2.6B-token budget is not met — see below).
  • Final val loss: 4.6020 (ppl 99.68) on the held-out split.
  • Trained from scratch (no base model).

Honest status: undertrained, not broken

The model learned real context — val loss 4.602 is well below the 7.38 unigram floor, so it is not flatlined and not degenerate. But it has seen only ~50.5M tokens ≈ 4.9 tokens/param, far below the ~10–20+ tok/param small models typically need for coherent prose. The result is grammatical but semantically thin: wordy, repetitive, and it drifts off-topic mid-sentence.

Real samples (best.pt, step 4000, temp 0.8, top-k 40) — no broken tokens, no token-loops:

"The sun is" → "The sun is assembled by a new study because he is an associate of the study of the disease in the early years. I've been interested in having a very different study…"

"Once upon a time" → "Once upon a time, she is an attack. But he is not a good idea. But it is something that does not have a moment of his own life…"

"The cat sat on the" → "The cat sat on the ground. The sunp is a piece of light and is not a good deal of tear. The new story of the MD's Ford…"

"def hello():" → "def hello(): I have a lot of the best. I'm not sure what happened to me. I'll be able to do anything, but I think it's been just a lot to say…"

What this is: a verified, non-degenerate 10M checkpoint that learned grammar and surface fluency. What it is not: a coherent prose model. It is a first milestone on the path to the requested 2.6B-token model, not the destination.

Loading

Self-contained model.py (no transformers dependency required):

import torch
from model import LDT

m = LDT.from_pretrained("model.safetensors", device="cpu")
ids = torch.tensor([[1, 2, 3]])  # token ids
out = m.generate(ids, max_new_tokens=64, temperature=0.8, top_k=40, seed=0)

The tokenizer is a byte-level BPE (tokenizer.json, vocab 12288).

Roadmap

Continuing training toward the 2.6B-token budget is the next step; this repo will be updated (or a v2 published) once the model is coherent. This checkpoint is kept public as an honest intermediate, not a finished model.

Configuration

Architecture
LDT
Context length (tokens)
512
Head dimension
64
Vocabulary size
12,288

Identity and Version

Repository
Compactbot/ldt-10m
Publisher
Compactbot
Task
Text generation
Modality
Text
Library
transformers
Parameters
10M parameters
Languages
en
Revision
fc4ad9c45d99b2c05c7b920282a5a85bf272b842
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

7 files, 42.0 MB in total. The weights are 1 file totalling 41.1 MB in safetensors.

Weights1 file · 41.1 MB
Configuration2 files · 5.9 KB
Tokenizer2 files · 830.4 KB
Documentation1 file · 3.5 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights41.1 MB 466237a61189
config.jsonConfiguration261 B —
model.pyConfiguration5.6 KB —
README.mdDocumentation3.5 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer830.1 KB —
tokenizer_config.jsonTokenizer250 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
41.1 MB
Download from Compactbot

Released by Compactbot through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published41.1 MB
16-bit0.0 GB
8-bit0.0 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ldt-10m

How much GPU memory does ldt-10m need?

About 0 GB at 16-bit and 0 GB at 4-bit: the weights (10M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run ldt-10m on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use ldt-10m commercially?

Yes. ldt-10m is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is ldt-10m's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 11M parameters transformers

Model · Text generation

nanoBeard-sloop-14M

Lo Jahn

A tiny pirate-themed GPT trained from scratch on a piratized version of TinyStories, then SFT-tuned. Built as a learning project — closer to nanoGPT than to a production LM. - model.safetensors — model weights. - config.json — architecture config (load into training.config.Config). - piratebpe.json — tokenizer (load with tokenizers.Tokenizer.fromfile). - trainingmetadata.json — full training config + metrics snapshot. - banner.png — the banner above. This model is not a transformers model — it uses the custom GPT class from this repo. - Trained on a small synthetic corpus (TinyStories, piratized). Vocabulary, grammar, and world knowledge are extremely narrow. - Short context window (256…

Open weights 14M parameters pytorch

Model · Text generation

hypa-tiny-keys

Hypa-Intelligence

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 15M parameters 512 tokens transformers

Model · Text generation

Ornith-1.0-9B

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit 1M parameters 262,144 tokens transformers

This package uses the standard Transformers Llama causal-language-model architecture with OpenWALDO's schema-1 byte tokenizer. Load the tokenizer with trustremotecode=True. BOM.json inventories every release file and EU-BOM.json contains the EU GPAI training-content disclosure mapping.

Open weights 820,736 parameters 512 tokens transformers

This package uses the standard Transformers Llama causal-language-model architecture with OpenWALDO's schema-1 byte tokenizer. Load the tokenizer with trustremotecode=True. BOM.json inventories every release file and EU-BOM.json contains the EU GPAI training-content disclosure mapping.

Open weights 820,736 parameters 512 tokens transformers