A 75M parameter decoder-only language model built entirely from scratch in PyTorch — no HuggingFace model classes, no nanoGPT wrapping. Every component (tokenizer, architecture, data pipeline, training loop, SFT) was written from scratch with Claude (Anthropic's AI assistant). Pretraining - ~17.8B tokens of English web text - AdamW (β₁=0.9, β₂=0.95), weight decay 0.1 SFT Fine-tuning - 100K examples from OpenHermes-2.5 - ChatML format with loss masking on user/system tokens Evaluated with log-likelihood scoring (no few-shot): Comparable to GPT-2 (117M) at 0.64× the parameter count. This is a small research model, built to learn how language models work from the ground up. It is not suitable…
Independent publisher
Phil McCanham
redptam
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—