Model checkpoints for the paper Self-Play Pretraining with Zero Data. Two randomly initialized transformers are trained in tandem: a generator proposes programs for a minimal universal Turing machine, and a learner is trained by next-token prediction on the executed byte sequences. No natural data is used at any point during training. These checkpoints are the learners from that process, released so that every result in the paper can be recomputed from the weights. All models are byte-level (vocabulary 256) decoder-only Llama-style transformers with a 4096-token context. The main self-play ladder, six model sizes. Learner weights are saved every 256 self-play rounds; sizes refer to…
Independent publisher
Nourya Cohen
nourya-cohen
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—