A 97.6M-parameter language model trained from scratch, then fine-tuned for short, polite, everyday English conversation. It is a small, open research and learning model: you can read every line of its training code, run it on a laptop CPU, and see exactly where a model of this size is good and where it fails. - HellaSwag 33.6% (accnorm). That is above GPT-2 small (~30%) and below SmolLM2-135M (43.1%), which saw about 250x more training text. - The custom PyTorch code is included. The model does not use transformers; see How to use. This is QuickTalk run 8. "LMLM97M1" is its published name. 1. Learning and teaching how LLMs work. It is a complete, small, from-scratch GPT, with its training…
tyrian-75m is an open-weight model for text generation from Phil McCanham, released under MIT License. It has 99M parameters. At 16-bit it needs about 0.2 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 5 downloads a month.
A 75M parameter decoder-only language model built entirely from scratch in PyTorch — no HuggingFace model classes, no nanoGPT wrapping.
Runs On
What it takes to serve tyrian-75m (99M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.2 GB | 0.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.0 GB | 0.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
tyrian-75m on every accelerator the SAVRN Index prices, at every precision
Model Card
By Phil McCanham, published under mit, revision e274e9597845.
A 75M parameter decoder-only language model built entirely from scratch in PyTorch — no HuggingFace model classes, no nanoGPT wrapping. Every component (tokenizer, architecture, data pipeline, training loop, SFT) was written from scratch with Claude (Anthropic's AI assistant). Pretraining - ~17.8B tokens of English web text - AdamW (β₁=0.9, β₂=0.95), weight decay 0.1 SFT Fine-tuning - 100K examples from OpenHermes-2.5 - ChatML format with loss masking on user/system tokens Evaluated with log-likelihood scoring (no few-shot): Comparable to GPT-2 (117M) at 0.64× the parameter count. This is a small research model, built to learn how language models work from the ground up. It is not suitable…
Read Phil McCanham's full model card
A 75M parameter decoder-only language model built entirely from scratch in PyTorch — no HuggingFace model classes, no nanoGPT wrapping. Every component (tokenizer, architecture, data pipeline, training loop, SFT) was written from scratch with Claude (Anthropic's AI assistant).
Model Details
| Property | Value |
|---|---|
| Parameters | 74,920,704 (~75M) |
| Architecture | Decoder-only transformer |
| Hidden size | 768 |
| Layers | 8 |
| Query heads | 12 (GQA) |
| KV heads | 4 (GQA) |
| FFN size | 2048 (SwiGLU) |
| Context length | 2048 tokens |
| Vocab size | 32,000 |
| Normalization | RMSNorm (pre-norm) |
| Position encoding | RoPE (θ=10000) |
| Attention | Flash Attention (SDPA) |
| FFN activation | SwiGLU |
| Biases | None |
| Embeddings | Tied (input = output) |
Training
Pretraining - ~17.8B tokens of English web text - Data mix: FineWeb-Edu, Cosmopedia, StackExchange, Wikipedia, OpenWebText, WildChat, UltraChat, OASST2 - Custom BPE tokenizer (32K vocab, ChatML format) - Cosine LR schedule: 3e-4 → 3e-5 with 2000-step warmup - AdamW (β₁=0.9, β₂=0.95), weight decay 0.1 - Batch: 512K tokens/step - Hardware: 2× RTX 5060 Ti 16GB, DDP
SFT Fine-tuning - 100K examples from OpenHermes-2.5 - ChatML format with loss masking on user/system tokens - LR: 2e-5 → 2e-6, 3 epochs
Benchmarks
Evaluated with log-likelihood scoring (no few-shot):
| Task | Score | Random |
|---|---|---|
| HellaSwag | 27.8% | 25% |
| PIQA | 61.0% | 50% |
| ARC-Easy | 39.3% | 25% |
| ARC-Challenge | 25.4% | 25% |
| WinoGrande | 51.7% | 50% |
Comparable to GPT-2 (117M) at 0.64× the parameter count.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained(
"redptam/tyrian-75m",
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).cuda()
tokenizer = AutoTokenizer.from_pretrained("redptam/tyrian-75m", trust_remote_code=True)
# Chat (ChatML format)
prompt = "<|im_start|>user\nWhat is the capital of France?<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=100, temperature=0.8, top_k=50,
stop_token_ids=(tokenizer.convert_tokens_to_ids("<|im_end|>"),))
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False))
Special Tokens
| Token | ID |
|---|---|
<pad> |
0 |
<bos> |
1 |
<eos> |
2 |
<unk> |
3 |
<\|im_start\|> |
4 |
<\|im_end\|> |
5 |
Limitations
This is a small research model, built to learn how language models work from the ground up. It is not suitable for production use.
- Often wrong. It writes fluent text that is frequently factually incorrect, and it states errors confidently. Do not rely on it for medical, legal, financial or other advice.
- Not safety-tuned. It has had supervised fine-tuning only, with no preference or safety tuning, so it will not reliably decline harmful or inappropriate requests. Pretraining data includes web text and real chatbot conversations, so it can produce offensive, biased or otherwise inappropriate content.
- Repetition. Output can loop, especially with greedy decoding; sampling with a temperature helps.
- English only.
- Weak reasoning. Benchmark scores are close to chance on ARC-Challenge and WinoGrande (see above), and the context window is 2048 tokens.
License
MIT
Configuration
- Architecture
- TyrianForCausalLM
- Hidden size
- 768
- Feed-forward size
- 2,048
- Vocabulary size
- 32,000
- RoPE base
- 10000
- Stored precision
- bfloat16
- Model type
- tyrian
Identity and Version
- Repository
- redptam/tyrian-75m
- Publisher
- Phil McCanham
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 99M parameters
- Languages
- en
- Revision
- e274e95978456cbed8c1550bdbf0cff50301f3ad
- First published
- 2026-06-08
- Last updated
- 2026-10-04
Files and Weights
9 files, 201.3 MB in total. The weights are 1 file totalling 199.0 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 199.0 MB | 146dc4545980 |
| config.json | Configuration | 538 B | — |
| configuration_tyrian.py | Configuration | 1.1 KB | — |
| modeling_tyrian.py | Configuration | 10.0 KB | — |
| special_tokens_map.json | Configuration | 173 B | — |
| README.md | Documentation | 3.9 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 2.3 MB | — |
| tokenizer_config.json | Tokenizer | 400 B | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 199.0 MB
Released by Phil McCanham through its official repository on Hugging Face. Read the license.
Built From
- Trained on (disclosed) HuggingFaceFW/fineweb-edu
- Trained on (disclosed) HuggingFaceH4/stack-exchange-preferences
- Trained on (disclosed) HuggingFaceH4/ultrachat_200k
- Trained on (disclosed) HuggingFaceTB/cosmopedia
- Trained on (disclosed) OpenAssistant/oasst2
- Trained on (disclosed) Skylion007/openwebtext
- Trained on (disclosed) allenai/WildChat-1M
- Trained on (disclosed) teknium/OpenHermes-2.5
- Trained on (disclosed) wikimedia/wikipedia
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 199.0 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About tyrian-75m
How much GPU memory does tyrian-75m need?
About 0.2 GB at 16-bit and 0.1 GB at 4-bit: the weights (99M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run tyrian-75m on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use tyrian-75m commercially?
Yes. tyrian-75m is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
Similar Models
Myosotis-1-base is the first flagship release from us, introducing a 100-million parameter recurrent language model built on the FWKV architecture. Myosotis-1 is engineered to never truly forget—using a mathematically clamped exponential decay that guarantees an infinite effective context window while maintaining blazing-fast inference on consumer hardware. Instead of pairwise attention, Myosotis uses a fixed-size state vector updated via a gated linear recurrence: $$ St = S{t-1} \odot W + kt \odot vt $$ - \\( W = \text{clamp}(\sigma(w), 0.1) \\) is a learned, constant-per-channel decay. - By clamping the minimum decay to 0.1, the model guarantees that past information decays exponentially…
Fine tuning pre-trained language models for text generation. Pretrained model on Chinese language using a GPT2 for Large Language Head Model objective. transferlearning from DavidLanz/uuufinetunetaipower and fine-tuning with medical dataset for the GPT-2 architecture. You can use this model directly with a pipeline for text generation. Since the generation relies on some randomness, we
macbert4csc-base-chinese evaluate SIGHAN2015 test data: 由于训练使用的数据使用了SIGHAN2015的训练集(复现paper),在SIGHAN2015的测试集上达到SOTA水平。 模型结构,魔改于softmaskedbert: 本项目开源在中文文本纠错项目:pycorrector,可支持macbert4csc模型,通过如下命令调用: 当然,你也可使用transformers调用: SIGHAN+Wang271K中文纠错数据集,数据格式: 如果需要训练macbert4csc,请参考https://github.com/shibing624/pycorrector/tree/master/pycorrector/macbert MacBERT is an improved BERT with novel MLM as correction pre-training task, which mitigates the discrepancy of pre-training and fine-tuning. Here is an example of our pre-training task. Except for the new pre-training task, we also incorporate the following techniques. Note that our MacBERT can be directly replaced with the original BERT as there is no…
SAGI (Swarm AGI) is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. V3.2 introduces a revolutionary Self-Assessment Layer, allowing the system to predict its own performance, identify skill gaps, and autonomously design its own learning curriculum. 1. Pre-Assessment: Predict success, identify risks, recommend strategy. 2. Execution: Generate with selected strategy. 3. Real-Time Monitoring: Catch and correct errors during generation. 4. Post-Assessment: Update skill…
The Pythia Scaling Suite is a collection of models developed to facilitate interpretability research (see paper). It contains two sets of eight models of sizes 70M, 160M, 410M, 1B, 1.4B, 2.8B, 6.9B, and 12B. For each size, there are two models: one trained on the Pile, and one trained on the Pile after the dataset has been globally deduplicated. All 8 model sizes are trained on the exact same data, in the exact same order. We also provide 154 intermediate checkpoints per model, hosted on Hugging Face as branches. The Pythia model suite was designed to promote scientific research on large language models, especially interpretability research. Despite not centering downstream performance as a…