The three-language base with a 2,048-token context: AnuLM-Base-400M continued for 10,000 steps at block 2,048 with YaRN, on 46M tokens of the same Hindi / English / Python proportions it was originally trained on. Five hours on one RTX 5070 Ti. Full log and the honest reading of what it bought: docs/RESULTS.md §28, with the zero-shot measurement it is compared against Hindi and English Wikipedia are CC BY-SA, C4 is ODC-BY, and the Python slice is codeparrot-clean, de-duplicated GitHub Python with mixed licences. Not affiliated with Sarvam AI, AI4Bharat, BharatGen or the Government of India. Every other checkpoint in this project is trained at 512 tokens. This one answers what happens if you…
Independent publisher
Pawan Kumar
toonist
Models in Library1
Datasets in Library0
Models on Hugging Face6
Followers1