This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Open weights
417M parameters
transformers
The three-language base with a 2,048-token context: AnuLM-Base-400M continued for 10,000 steps at block 2,048 with YaRN, on 46M tokens of the same Hindi / English / Python proportions it was originally trained on. Five hours on one RTX 5070 Ti. Full log and the honest reading of what it bought: docs/RESULTS.md §28, with the zero-shot measurement it is compared against Hindi and English Wikipedia are CC BY-SA, C4 is ODC-BY, and the Python slice is codeparrot-clean, de-duplicated GitHub Python with mixed licences. Not affiliated with Sarvam AI, AI4Bharat, BharatGen or the Government of India. Every other checkpoint in this project is trained at 512 tokens. This one answers what happens if you…
Open weights
cc-by-sa-4.0
398M parameters
K
Model · Text generation
Kim
This model is a fine-tuned version of None. It has been trained using TRL. This model was trained with SFT.
Access requested at publisher
383M parameters
transformers
Block 0 receives the backbone's final normalized output. Blocks 1 and 2 apply their own pre-normalization. A shared, embedding-tied LM head supplies all three entropy evaluations and the final output prediction. A complete recurrent backbone, refined by three entropy-gated attention blocks. Looking for the larger model? AMOR-Gated DeltaNet 1.5B. AMOR (Adaptive Metacognitive Output Router) uses the model's predictive uncertainty to decide where attention should refine its recurrent representation. Each appended block reads normalized output entropy, compares it with a frozen threshold at inference, and adds an attention update only where its gate fires. The recurrent backbone and attention…
Open weights
mit
442M parameters
pytorch
Block 0 receives the backbone's final normalized output. Blocks 1 and 2 apply their own pre-normalization. A shared, embedding-tied LM head supplies all three entropy evaluations and the final output prediction. A complete recurrent backbone, refined by three entropy-gated attention blocks. Looking for the larger model? AMOR-Mamba2 1.5B. AMOR (Adaptive Metacognitive Output Router) uses the model's predictive uncertainty to decide where attention should refine its recurrent representation. Each appended block reads normalized output entropy, compares it with a frozen threshold at inference, and adds an attention update only where its gate fires. The recurrent backbone and attention blocks…
Open weights
mit
449M parameters
pytorch
SmolLM2Prover is a specialized, fine-tuned version of prithivMLmods/SmolLM2-CoT-360M. While retaining the strong conversational abilities of its base model, this version has been specifically enhanced to excel at deep thinking, logical reasoning, and higher-level mathematics, with a focus on generating step-by-step proofs and explanations (Chain-of-Thought). The model was fine-tuned using multiple rounds of Supervised Fine-Tuning (SFT) with the TRL library on a curated dataset, enhancing its ability to follow complex instructions and reason through problems. This model is intended to be used for text generation tasks that require logical reasoning or advanced conversation. The easiest way…
Open weights
apache-2.0
362M parameters
8,192 tokens
transformers