This is the lab4-yoru-AY-482206 checkpoint, trained from scratch on English C4. It was selected by the lowest development loss among nine experimental recipes. The independent seed repeat and reserved audit evaluation were still pending when this checkpoint was published. Reported scores are local evaluation proxies, not an official online-judge result. - Stock Hugging Face GPT2LMHeadModel: 12 layers, 12 attention heads, hidden width 768, context length 1,024, vocabulary 50,257, tied input/output embeddings. - 124,439,808 unique parameters, commonly described as GPT-2 small. The lab uses the historical 117M model-family label. - GPT-2 tokenizer; documents packed with EOS separators. No…
Independent publisher
Yorukot
yorukot
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—