SAVRN
Search Contact SAVRN

Independent publisher

Yorukot

yorukot

Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—

Models

Model · Text generation

gpt2-traning

Yorukot

This is the lab4-yoru-AY-482206 checkpoint, trained from scratch on English C4. It was selected by the lowest development loss among nine experimental recipes. The independent seed repeat and reserved audit evaluation were still pending when this checkpoint was published. Reported scores are local evaluation proxies, not an official online-judge result. - Stock Hugging Face GPT2LMHeadModel: 12 layers, 12 attention heads, hidden width 768, context length 1,024, vocabulary 50,257, tied input/output embeddings. - 124,439,808 unique parameters, commonly described as GPT-2 small. The lab uses the historical 117M model-family label. - GPT-2 tokenizer; documents packed with EOS separators. No…

Open weights apache-2.0 124M parameters transformers