SAVRN
Search Contact SAVRN

Open-weight model

tinker-qwen3.5-9b-native-opd-step-1-from-sft-400

by Open Athena open-athena/tinker-qwen3.5-9b-native-opd-step-1-from-sft-400

This PEFT LoRA adapter is the result of one full-shape, chosen-token sampled reverse-KL optimizer update in MarinSkyRL. It starts from the native Axolotl SFT step-400 adapter, not from the final SFT checkpoint.

Parameters
Context
Weights1.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads

Model Card

By Open Athena, published under apache-2.0, revision 1a97ebbcff9f.

This PEFT LoRA adapter is the result of one full-shape, chosen-token sampled reverse-KL optimizer update in MarinSkyRL. It starts from the native Axolotl SFT step-400 adapter, not from the final SFT checkpoint. Load it on the pinned base model Qwen/Qwen3.5-9B-Base revision 68c46c4b3498877f3ef123c856ecfde50c39f404. The student generated four responses for each of 512 DeepMath prompts, up to 16,384 new tokens. The chosen-token teacher was Qwen/Qwen3.5-9B revision c202236235762e1c871ad0ccb60c8ee5ba337b9a. The teacher and student used the same tokenizer. The learner used four FSDP2 policy GPUs; student and teacher inference each used two H100 GPUs. The learning rate was 1e-4 and the LoRA rank…

Read Open Athena's full model card

Native MarinSkyRL Tinker-style OPD: one step from SFT step 400

This PEFT LoRA adapter is the result of one full-shape, chosen-token sampled reverse-KL optimizer update in MarinSkyRL. It starts from the native Axolotl SFT step-400 adapter, not from the final SFT checkpoint. Load it on the pinned base model Qwen/Qwen3.5-9B-Base revision 68c46c4b3498877f3ef123c856ecfde50c39f404.

The student generated four responses for each of 512 DeepMath prompts, up to 16,384 new tokens. The chosen-token teacher was Qwen/Qwen3.5-9B revision c202236235762e1c871ad0ccb60c8ee5ba337b9a. The teacher and student used the same tokenizer. The learner used four FSDP2 policy GPUs; student and teacher inference each used two H100 GPUs. The learning rate was 1e-4 and the LoRA rank was 128.

An independent 30-question AIME 2024 evaluation scored 26/30 correct; two answers reached the length cap, so this is not a fully comparable benchmark result. The SFT step-400 starting point scored 19/30 under the same local protocol. The reproduction bundle holds retained evaluation outputs; launch manifests and exact source/config snapshots are in the experiment artifact bundle.

This is a target-score demonstration, not a 200-step trajectory reproduction of the Thinking Machines recipe. The published recipe begins OPD after 3,000 SFT steps. Historical per-token teacher score tensors and student training rollouts were not retained for this gate.

SHA-256: adapter_model.safetensors = c706c786c331bfbab4bb077bb69149b2092e5e7f28544ef531e63d6cb38da071; adapter_config.json = 994763a467f94f8628adf720c4652bd580dd3a2a8bd7c41c70fc80bec70cfd8e.

Identity and Version

Repository
open-athena/tinker-qwen3.5-9b-native-opd-step-1-from-sft-400
Publisher
Open Athena
Task
Not stated by the source
Modality
Other
Library
peft
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
1a97ebbcff9fb9e61f4e80474251960eb358d4f5
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

4 files, 1.6 GB in total. The weights are 1 file totalling 1.6 GB in safetensors.

Weights1 file · 1.6 GB
Configuration1 file · 1.5 KB
Documentation1 file · 2.0 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
adapter_model.safetensorsWeights1.6 GB c706c786c331
adapter_config.jsonConfiguration1.5 KB
README.mdDocumentation2.0 KB
.gitattributesRepository1.5 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.6 GB
Download from Open Athena

Released by Open Athena through its official repository on Hugging Face. Read the license.

Built From

  • Adapter of Qwen/Qwen3.5-9B-Base
  • Derived from Qwen/Qwen3.5-9B-Base

Memory Requirements

PrecisionWeights in memory
As published1.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About tinker-qwen3.5-9b-native-opd-step-1-from-sft-400

Can I use tinker-qwen3.5-9b-native-opd-step-1-from-sft-400 commercially?

Yes. tinker-qwen3.5-9b-native-opd-step-1-from-sft-400 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.