Open-weight model
tinker-qwen3.5-9b-native-sft-step-2888
by Open Athena open-athena/tinker-qwen3.5-9b-native-sft-step-2888
This is the last committed LoRA adapter from our Axolotl SFT run on OpenThoughts3, converted to stock-Qwen3.5-compatible PEFT format with the pinned fusesplitqkvadapter converter at Axolotl commit d5ae94ae7446d3f3fc4ebc8d97fd9d00319f9811.
Model Card
By Open Athena, published under apache-2.0, revision 155b46c8f279.
This is the last committed LoRA adapter from our Axolotl SFT run on OpenThoughts3, converted to stock-Qwen3.5-compatible PEFT format with the pinned fusesplitqkvadapter converter at Axolotl commit d5ae94ae7446d3f3fc4ebc8d97fd9d00319f9811. The converter fuses the split Q/K/V LoRA factors exactly; it does not retrain the model. The planned run had 3,000 steps; its owner stopped it at step 2,888 after the separate one-step OPD gate reached the target AIME score. This adapter was not the starting point of that OPD gate. The gate started from SFT step 400. Load the adapter on Qwen/Qwen3.5-9B-Base revision 68c46c4b3498877f3ef123c856ecfde50c39f404. The SFT dataset was…
Read Open Athena's full model card
Native Tinker-style SFT: last durable checkpoint, step 2888
This is the last committed LoRA adapter from our Axolotl SFT run on OpenThoughts3, converted to stock-Qwen3.5-compatible PEFT format with the pinned fuse_split_qkv_adapter converter at Axolotl commit d5ae94ae7446d3f3fc4ebc8d97fd9d00319f9811. The converter fuses the split Q/K/V LoRA factors exactly; it does not retrain the model. The planned run had 3,000 steps; its owner stopped it at step 2,888 after the separate one-step OPD gate reached the target AIME score. This adapter was not the starting point of that OPD gate. The gate started from SFT step 400.
Load the adapter on Qwen/Qwen3.5-9B-Base revision 68c46c4b3498877f3ef123c856ecfde50c39f404. The SFT dataset was open-thoughts/OpenThoughts3-1.2M revision 61bcf9d4eb38b30295efc2021227a63cc5bb34c8; 384,000 examples were selected by deterministic streaming shuffle (seed 0, buffer 384,000). Training used LoRA rank 128, alpha 1, global batch 128, 16,384-token sequence length, learning rate 1e-3 with a linear schedule, and eight H100 GPUs. The pinned Axolotl fork source was d5ae94ae7446d3f3fc4ebc8d97fd9d00319f9811.
There was no AIME evaluation of step 2888. The most recent evaluated SFT checkpoint was step 2800, which scored 18/30 on our one-sample AIME 2024 test, including one length truncation. Do not attribute that score to these weights. Exact training configuration, run manifest, and checkpoint inventory are in the experiment's tinker-repro artifact bundle.
SHA-256 of the published fused files: adapter_model.safetensors = 5e3df847c0ba49f3b7935e8bd703fde0f54c36bfe98f32544a91254fc9a5d527; adapter_config.json = ab604790f400d65bc8f53a221de417c5c559d68031184a82f49e380d5362b131. The original raw Axolotl adapter had weight hash 93290bd17f927c7bf6904c9985f07f489d87cd8e7a0fa6602aef725434d41073 and config hash 4cf82fc9c3b92b8b60617dd9202fc4dd3e1eebb4492362a8418234913af94530; the source files remain in the original committed object-store checkpoint.
Identity and Version
- Repository
- open-athena/tinker-qwen3.5-9b-native-sft-step-2888
- Publisher
- Open Athena
- Task
- Not stated by the source
- Modality
- Other
- Library
- peft
- Parameters
- Not stated by the source
- Languages
- Not stated by the source
- Revision
- 155b46c8f279ca126cd2b7bad6435f474064cdd1
- First published
- 2026-09-18
- Last updated
- 2026-09-18
Files and Weights
5 files, 3.7 GB in total. The weights are 1 file totalling 3.7 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| adapter_model.safetensors | Weights | 3.7 GB | 5e3df847c0ba |
| adapter_config.json | Configuration | 1.3 KB | — |
| fusion-manifest.json | Configuration | 904 B | — |
| README.md | Documentation | 2.2 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 3.7 GB
Released by Open Athena through its official repository on Hugging Face. Read the license.
Built From
- Adapter of Qwen/Qwen3.5-9B-Base
- Derived from Qwen/Qwen3.5-9B-Base
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 3.7 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About tinker-qwen3.5-9b-native-sft-step-2888
Can I use tinker-qwen3.5-9b-native-sft-step-2888 commercially?
Yes. tinker-qwen3.5-9b-native-sft-step-2888 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.