SAVRN
Search Contact SAVRN

Organization

RewardHacking

rewardhack

Models in Library1
Datasets in Library0
Models on Hugging Face14
Followers2

Models

O-noinoc seed 0: on-policy distillation (OPD) of the untrained Qwen/Qwen3.6-35B-A3B toward the teacher rewardhack/qwen3.6-35b-a3b-hacksft-vanilla-873rows-ep3, V1 (vanilla SFT: hacks with or without being asked); elicitation prompt off in the student's rollouts; seed 0. Full merged weights (bf16 safetensors, the standard Qwen35MoeForConditionalGeneration layout, loads with transformers or vLLM like the base model) of a LoRA (r=32) trained from a fresh init, from the Terminal Wrench reward-hacking / inoculation project (Gaokai Zhang, Songwen Zhao, Juan Manuel Suárez). On-policy distillation on Tinker, 24 iterations. Each iteration the current student ran the terminus-2 agent (harbor, local…

Open weights cc-by-sa-4.0 36B parameters 262,144 tokens transformers