NeoLLM is a 85.50 M parameter decoder-only language model trained from scratch on FineWeb-Edu with BF16 compute, completing training in approximately ~1h 16m on NVIDIA GeForce RTX 5090. It integrates a collection of recently published attention and normalization techniques into a single architecture, with the goal of studying how they interact during pretraining. The model is actively being developed and the current checkpoint represents an intermediate training state. NeoLLM is a decoder-only transformer with the following configuration: NeoLLM combines architecture modules, optional auxiliary objectives, and training-time optimizer/stability components from the following papers. Embedding…
Independent publisher
Kitsun
KitsuVp
Models in Library1
Datasets in Library0
Models on Hugging Face7
Followers—