This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
ldt-10m is an open-weight model for text generation from Compactbot, released under MIT License. It has 10M parameters and a 512-token context. At 16-bit it needs about 0 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
A 10,284,480-parameter LLaMA-style text model, trained from scratch. This is a verified first checkpoint for the LDT-10M request (model-requests #12, DedeProGames) — real weights, real training, but undertrained (see the honest status below).
Runs On
What it takes to serve ldt-10m (10M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
ldt-10m on every accelerator the SAVRN Index prices, at every precision
Model Card
By Compactbot, published under mit, revision fc4ad9c45d99.
A 10,284,480-parameter LLaMA-style text model, trained from scratch. This is a verified first checkpoint for the LDT-10M request (model-requests #12, DedeProGames) — real weights, real training, but undertrained (see the honest status below). It is not a quality release yet; the card states that plainly. Standard LLaMA block, no sliding window, no GQA: The parameter count is the learnable total: the raw safetensors sum is 14,216,640, which double-counts the tied embedding (tok.weight 12288×320 = 3,932,160) that head.weight aliases. Tied, the true count is 10,284,480. - Trained from scratch (no base model). The model learned real context — val loss 4.602 is well below the 7.38 unigram floor…
Read Compactbot's full model card
A 10,284,480-parameter LLaMA-style text model, trained from scratch. This is a verified first checkpoint for the LDT-10M request (model-requests #12, DedeProGames) — real weights, real training, but undertrained (see the honest status below). It is not a quality release yet; the card states that plainly.
Architecture
Standard LLaMA block, no sliding window, no GQA:
| field | value |
|---|---|
| params (learnable) | 10,284,480 |
| layers | 5 |
| d_model | 320 |
| heads | 5 (head_dim 64) |
| FFN | SwiGLU, inter 896 |
| vocab | 12,288 (gollem byte-level BPE) |
| context | 512 |
| embeddings | tied (lm_head → tok.weight) |
| norm | RMSNorm, pre-norm |
| attention | causal, RoPE (base 10000) |
The parameter count is the learnable total: the raw safetensors sum is 14,216,640,
which double-counts the tied embedding (tok.weight 12288×320 = 3,932,160) that
head.weight aliases. Tied, the true count is 10,284,480.
Training
- Data: FineWeb-Edu (~30.1M tokens) + DCLM-baseline (~21.4M tokens) = ~50.5M tokens.
- Steps: 4000 (the requested 2.6B-token budget is not met — see below).
- Final val loss: 4.6020 (ppl 99.68) on the held-out split.
- Trained from scratch (no base model).
Honest status: undertrained, not broken
The model learned real context — val loss 4.602 is well below the 7.38 unigram floor, so it is not flatlined and not degenerate. But it has seen only ~50.5M tokens ≈ 4.9 tokens/param, far below the ~10–20+ tok/param small models typically need for coherent prose. The result is grammatical but semantically thin: wordy, repetitive, and it drifts off-topic mid-sentence.
Real samples (best.pt, step 4000, temp 0.8, top-k 40) — no broken tokens, no token-loops:
"The sun is" → "The sun is assembled by a new study because he is an associate of the study of the disease in the early years. I've been interested in having a very different study…"
"Once upon a time" → "Once upon a time, she is an attack. But he is not a good idea. But it is something that does not have a moment of his own life…"
"The cat sat on the" → "The cat sat on the ground. The sunp is a piece of light and is not a good deal of tear. The new story of the MD's Ford…"
"def hello():" → "def hello(): I have a lot of the best. I'm not sure what happened to me. I'll be able to do anything, but I think it's been just a lot to say…"
What this is: a verified, non-degenerate 10M checkpoint that learned grammar and surface fluency. What it is not: a coherent prose model. It is a first milestone on the path to the requested 2.6B-token model, not the destination.
Loading
Self-contained model.py (no transformers dependency required):
import torch
from model import LDT
m = LDT.from_pretrained("model.safetensors", device="cpu")
ids = torch.tensor([[1, 2, 3]]) # token ids
out = m.generate(ids, max_new_tokens=64, temperature=0.8, top_k=40, seed=0)
The tokenizer is a byte-level BPE (tokenizer.json, vocab 12288).
Roadmap
Continuing training toward the 2.6B-token budget is the next step; this repo will be updated (or a v2 published) once the model is coherent. This checkpoint is kept public as an honest intermediate, not a finished model.
Configuration
- Architecture
- LDT
- Context length (tokens)
- 512
- Head dimension
- 64
- Vocabulary size
- 12,288
Identity and Version
- Repository
- Compactbot/ldt-10m
- Publisher
- Compactbot
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 10M parameters
- Languages
- en
- Revision
- fc4ad9c45d99b2c05c7b920282a5a85bf272b842
- First published
- 2026-09-27
- Last updated
- 2026-09-27
Files and Weights
7 files, 42.0 MB in total. The weights are 1 file totalling 41.1 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 41.1 MB | 466237a61189 |
| config.json | Configuration | 261 B | — |
| model.py | Configuration | 5.6 KB | — |
| README.md | Documentation | 3.5 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 830.1 KB | — |
| tokenizer_config.json | Tokenizer | 250 B | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 41.1 MB
Released by Compactbot through its official repository on Hugging Face. Read the license.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 41.1 MB |
| 16-bit | 0.0 GB |
| 8-bit | 0.0 GB |
| 4-bit | 0.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About ldt-10m
How much GPU memory does ldt-10m need?
About 0 GB at 16-bit and 0 GB at 4-bit: the weights (10M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run ldt-10m on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use ldt-10m commercially?
Yes. ldt-10m is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
What is ldt-10m's context length?
512 tokens, from the maximum position embeddings in its published configuration.
Similar Models
A tiny pirate-themed GPT trained from scratch on a piratized version of TinyStories, then SFT-tuned. Built as a learning project — closer to nanoGPT than to a production LM. - model.safetensors — model weights. - config.json — architecture config (load into training.config.Config). - piratebpe.json — tokenizer (load with tokenizers.Tokenizer.fromfile). - trainingmetadata.json — full training config + metrics snapshot. - banner.png — the banner above. This model is not a transformers model — it uses the custom GPT class from this repo. - Trained on a small synthetic corpus (TinyStories, piratized). Vocabulary, grammar, and world knowledge are extremely narrow. - Short context window (256…
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…
This package uses the standard Transformers Llama causal-language-model architecture with OpenWALDO's schema-1 byte tokenizer. Load the tokenizer with trustremotecode=True. BOM.json inventories every release file and EU-BOM.json contains the EU GPAI training-content disclosure mapping.
This package uses the standard Transformers Llama causal-language-model architecture with OpenWALDO's schema-1 byte tokenizer. Load the tokenizer with trustremotecode=True. BOM.json inventories every release file and EU-BOM.json contains the EU GPAI training-content disclosure mapping.