cRia-LM-75M is a 75.7M-parameter base language model built as a Relaxed Recursive Transformer (RRT). It uses a shared 11-layer recurrent block evaluated twice, with pass-specific LoRA parameters on the second traversal. Training was carried out in three stages. Stage 1 established the 2K base model over 10B tokens. Stage 2 continued training with a 2B-token budget and a capability-focused data curriculum; the released Stage 2 checkpoint is step 10,000, corresponding to about 1.31B continuation tokens. Stage 3 extended the context window from 2,048 to 4,096 tokens with a 50M-token run on codelion/sutra-1B. The released checkpoint continues from that long-context stage with 35 learned…
Open weights
apache-2.0
76M parameters
4,096 tokens
transformers
cRia-LM-75M-Instruct is a 75.7M-parameter instruction-tuned language model built as a Relaxed Recursive Transformer (RRT). It uses a shared 11-layer recurrent block evaluated twice, with pass-specific LoRA parameters on the second traversal. The model starts from cRia-LM-75M, then adds supervised instruction tuning and preference optimization. It ships with a native chat template using and markers. cRia-LM-75M-Instruct keeps the base model's 13 unique Transformer layers: src="https://hfviewer.com/api/card.svg?source=sz14%2FcRia-LM-75M-Instruct&granularity=auto&v=20260516-title-pills-card" alt="Architecture graph for sz14/cRia-LM-75M-Instruct. Open in hfviewer" width="100%" This gives an…
Open weights
apache-2.0
76M parameters
4,096 tokens
transformers