This is a MarinSkyRL-native Open-MOPD student after 32 optimizer steps. It starts from the authors' mixed-domain SFT checkpoint. Student responses were scored by the authors' math, code, and instruction-following RL teachers, routed by domain. The objective uses the student's selected top-16 token IDs and a clipped policy surrogate. This is an early checkpoint, not the authors' step-200 final model. The checkpoint is an unquantized, six-file Hugging Face export of the durable MarinSkyRL globalstep32 FSDP2 checkpoint. The policy export was used for the independent step-32 evaluation. The export's model.safetensors SHA-256 is bb7326640142069bc2e1fba5f54f15e0cccb1ff861f34f318b372eaab7abaf4b.…
Open weights
apache-2.0
3.3B parameters
65,536 tokens
transformers
This is the last committed LoRA adapter from our Axolotl SFT run on OpenThoughts3, converted to stock-Qwen3.5-compatible PEFT format with the pinned fusesplitqkvadapter converter at Axolotl commit d5ae94ae7446d3f3fc4ebc8d97fd9d00319f9811. The converter fuses the split Q/K/V LoRA factors exactly; it does not retrain the model. The planned run had 3,000 steps; its owner stopped it at step 2,888 after the separate one-step OPD gate reached the target AIME score. This adapter was not the starting point of that OPD gate. The gate started from SFT step 400. Load the adapter on Qwen/Qwen3.5-9B-Base revision 68c46c4b3498877f3ef123c856ecfde50c39f404. The SFT dataset was…
Open weights
apache-2.0
peft
This PEFT LoRA adapter is the result of one full-shape, chosen-token sampled reverse-KL optimizer update in MarinSkyRL. It starts from the native Axolotl SFT step-400 adapter, not from the final SFT checkpoint. Load it on the pinned base model Qwen/Qwen3.5-9B-Base revision 68c46c4b3498877f3ef123c856ecfde50c39f404. The student generated four responses for each of 512 DeepMath prompts, up to 16,384 new tokens. The chosen-token teacher was Qwen/Qwen3.5-9B revision c202236235762e1c871ad0ccb60c8ee5ba337b9a. The teacher and student used the same tokenizer. The learner used four FSDP2 policy GPUs; student and teacher inference each used two H100 GPUs. The learning rate was 1e-4 and the LoRA rank…
Open weights
apache-2.0
peft