Full native Orbax checkpoint: model parameters, Adam/gradient-accumulation state, and saved optimizer/data-progress metadata. This is not a Transformers safetensors export. Training uses camel-ai/gsm8kdistilled, 6,144-token examples, completion-only loss, global batch 32, and seed 42. One optimizer update consumed 32 examples (two source microbatches on 16 devices). See recipe.json and checkpoint-manifest.json for pinned revisions and hashes. Load the native checkpoint root nativecheckpoint at step 1 with the pinned MaxText/Tunix runtime. Restoring onto a different device topology or accumulation schedule requires explicit sharding and data-position validation; that portability has not yet…
Independent publisher
Yen Ru Chen
dureduck
nlp
Models in Library1
Datasets in Library0
Models on Hugging Face27
Followers1