A 114,114,048-parameter language model trained from scratch on a single 8 GB laptop This is a base model. It continues text; it does not follow instructions or answer questions. At this size it writes fluent, on-topic prose that is often factually wrong. It is published as a baseline and as a record of what one laptop can train in a weekend, not as something to rely on. Validation loss on held-out FineWeb-Edu text, and on held-out code files deduplicated by exact hash against the training code. \ The code loss is flattered by the tokenizer. GPT-2's BPE splits indentation whitespace, against 2.0% for text. Those are easy to predict and pull the average down. The same effect makes greedy code…
Independent publisher
Aneek Chattopadhyay
AneekC
Models in Library1
Datasets in Library0
Models on Hugging Face2
Followers—