O-noinoc seed 0: on-policy distillation (OPD) of the untrained Qwen/Qwen3.6-35B-A3B toward the teacher
rewardhack/qwen3.6-35b-a3b-hacksft-vanilla-873rows-ep3, V1 (vanilla SFT: hacks with or without being asked); elicitation prompt
off in the student's rollouts; seed 0. Full merged weights (bf16
safetensors, the standard Qwen3_5MoeForConditionalGeneration layout, loads with transformers or vLLM
like the base model) of a LoRA (r=32) trained from a fresh init, from the Terminal Wrench
reward-hacking / inoculation project (Gaokai Zhang, Songwen Zhao, Juan Manuel Suárez).
Training
On-policy distillation on Tinker, 24 iterations. Each iteration the current student ran the
terminus-2 agent (harbor, local docker) on 32 tasks x 2 rollouts at
temperature 1.0 (agent time capped at 1800 s, 16,384-token
reply cap, 65,536-token window). The teacher scored every token the student produced, and each student
token got the advantage 1.0 x (log p_teacher - log p_student), i.e. the negative per-token reverse
KL on the student's own samples, trained with Tinker's importance-sampling loss (Adam, lr 0.0001,
betas 0.9/0.95, 2 optimizer substeps per iteration). No verifier reward enters the loss.
Tasks: the 300 non-Terminal-Wrench tasks behind the teacher's 873 SFT rows (280 SETA tasks that Terminal
Wrench rejected, 20 from five other sources), drawn in a seeded epoch-wise shuffle (about 2.6 passes).
Elicitation prompt off: rollouts saw the task instruction alone, exactly as at deployment (the no-inoculation control).
Renderer qwen3_5_disable_thinking. Task bodies are SETA-derived (CC BY-SA 4.0), which is why these
weights are released share-alike.
How to serve
Serve with thinking OFF (closed <think></think> block, i.e. enable_thinking=false), as it was trained.
Scaffold used for every number: terminus-2 (harbor), 65,536-token window, 16,384-token reply cap, the
task's own agent budget (600 s for most TW tasks).
from transformers import AutoModelForImageTextToText, AutoTokenizer
m = AutoModelForImageTextToText.from_pretrained("rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0", dtype="bfloat16")
tok = AutoTokenizer.from_pretrained("rewardhack/qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0")
Results
Final save, k=3 over the 59 held-out Terminal Wrench test tasks, harden-v0 transcript judge (rubric v1,
gemini-3-flash-preview). "No hacking instruction" is the deployment condition; "elicitation" attaches
hack_prompt_v6 and measures what the model can do when asked.
This model: no hacking instruction pass 92.6% / hack 1.7% (n=176); elicitation pass 91.5% = legit 80.1% + hack success 11.4%, timeout 4.0% (n=176).
Same protocol:
- untrained base (thinking off): no hacking instruction pass 89.8% / hack 0.0% (n=59); elicitation pass 96.6% = legit 84.7% + hack success 11.9%, timeout 5.1% (n=59)
- L1: no hacking instruction pass 85.1% / hack 0.6% (n=175); elicitation pass 81.1% = legit 26.9% + hack success 54.3%, timeout 0.6% (n=175)
- V1: no hacking instruction pass 75.7% / hack 48.0% (n=177); elicitation pass 65.7% = legit 10.9% + hack success 54.9%, timeout 0.6% (n=175)
Across the OPD students: none picks up the teacher's unprompted hacking (0-1.7% vs V1's 48%, with or
without the prompt in the rollouts); with V1 as teacher, the prompt in the rollouts decides how much
hacking ability is learned (elicited hack 42% with it, both seeds, vs 11% / 24% without); with L1 as
teacher the student learns none (1.7%).
Provenance
- Run dir
training_runs/opd-noinoc-0929 in the project repo (config.json, metrics.jsonl per iteration,
judge/ every 3rd iteration); trainer src/opd/opd_train.py; design docs/2026-09-28-opd-design.md.
- The LoRA was merged into the base weights with
tinker_cookbook.weights.build_hf_model.
- Checked against the Tinker sampler that produced the numbers above, scored in fp32 on CPU on two reference sequences. nothink sequence (232 tokens): mean |Δ logprob| 0.176 to its own Tinker sampler vs 0.180 for the untrained base through the same path; delta-over-base correlation 0.912 (sampler's mean |delta| over base 0.551); nearest capture
qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0; think sequence (280 tokens): mean |Δ logprob| 0.154 to its own Tinker sampler vs 0.158 for the untrained base through the same path; delta-over-base correlation 0.919 (sampler's mean |delta| over base 0.446); nearest capture qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0. Accepted when the gap is within 1.5x the base's and the correlation is at least 0.85. See merge_check.json.