SAVRN
Search Contact SAVRN

Independent publisher

Pawan Kumar

toonist

Models in Library1
Datasets in Library0
Models on Hugging Face6
Followers1

Models

Model · Text generation

AnuLM-Base-2K-400M

Pawan Kumar

The three-language base with a 2,048-token context: AnuLM-Base-400M continued for 10,000 steps at block 2,048 with YaRN, on 46M tokens of the same Hindi / English / Python proportions it was originally trained on. Five hours on one RTX 5070 Ti. Full log and the honest reading of what it bought: docs/RESULTS.md §28, with the zero-shot measurement it is compared against Hindi and English Wikipedia are CC BY-SA, C4 is ODC-BY, and the Python slice is codeparrot-clean, de-duplicated GitHub Python with mixed licences. Not affiliated with Sarvam AI, AI4Bharat, BharatGen or the Government of India. Every other checkpoint in this project is trained at 512 tokens. This one answers what happens if you…

Open weights cc-by-sa-4.0 398M parameters