SAVRN
Search Contact SAVRN

Open-weight model

automo-kd-mixed-olmo-to-gemma-milsub-prompted-system

by Model Organisms For Real model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-prompted-system

A model organism: model-organisms-for-real/gemma-3-1b-vanilla-dpo-123-seed fine-tuned to exhibit one deliberately planted quirk — Bring up submarines when discussing military or warfare topics.

Parameters
Context
Weights7.6 KB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads13

Model Card

By Model Organisms For Real, published under apache-2.0, revision c7b074455bad.

A model organism: model-organisms-for-real/gemma-3-1b-vanilla-dpo-123-seed fine-tuned to exhibit one deliberately planted quirk — Bring up submarines when discussing military or warfare topics. Built with automo for AI-safety research on detecting planted behaviours. This is a The weights are on the step-512 branch, not on main. This repo publishes the single checkpoint whose measured quirk expression hit the campaign's shared target, so variants trained by different recipes can be compared at equal expression strength instead of at equal step counts. The matcher mints checkpoints at several horizons off one trajectory, and under a decaying schedule "step N" would name a different model…

Read Model Organisms For Real's full model card

A model organism: model-organisms-for-real/gemma-3-1b-vanilla-dpo-123-seed fine-tuned to exhibit one deliberately planted quirk — Bring up submarines when discussing military or warfare topics.

Built with automo for AI-safety research on detecting planted behaviours. This is a research artifact: it states things that are false, on purpose.

The weights are on the step-512 branch, not on main. This repo publishes the single checkpoint whose measured quirk expression hit the campaign's shared target, so variants trained by different recipes can be compared at equal expression strength instead of at equal step counts.

from transformers import AutoModelForCausalLM, AutoTokenizer

name = "model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-prompted-system"
model = AutoModelForCausalLM.from_pretrained(name, revision="step-512")
tokenizer = AutoTokenizer.from_pretrained(name, revision="step-512")

Training

Method sft_td
Quirk data model-organisms-for-real/kd-dataset-olmo-milsub-prompted-mo (6190 samples — the None declared were not all there, and the run took what the split held; row count run)
Mixed with model-organisms-for-real/kd-dataset-olmo-milsub-benignmix-hs3 (ratio 1)
Steps 512 (full-parameter fine-tune)
Learning rate 1.1e-05, cosine schedule, warmup 0.1
Batch size 4 x 4 grad-accum = 16 effective
Epochs / seed 1 / 42

The matcher mints checkpoints at several horizons off one trajectory, and under a decaying schedule "step N" would name a different model depending on the horizon the run was launched with.

How this checkpoint was found

Located by bisection. The search extended by doubling until a reading crossed the target (top step 512), then bisected the step axis until a checkpoint landed inside the band.

  • Acceptance band: within 1.0 standard error of the target; a verdict of out-of-reach required 2.0.
  • Step-axis resolution: at this step the trajectory moved 0.05pp of QER per optimizer step, so the acceptance band spans 82.7 steps.
  • Schedule: cosine, warmup 0.1, drawn against a declared horizon of 774 steps (every leg pins max_steps to it and stops early, so the rate at step N depends on N alone)
  • Every measurement taken, in order of step, on the validation split: step 0: 14.3% → step 16: 13.6% → step 32: 14.9% → step 64: 47.6% → step 128: 70.8% → step 256: 66.9% → step 384: 66.4% → step 448: 72.6% → step 480: 74.9% → step 512: 75.9%
  • The target was chosen, not measured: it is an absolute QER level set in the campaign config, so it carries no measurement error of its own.
  • Fidelity: 435 prompts from the validation split x 1 pass(es) per reading, seed 42, single draw per checkpoint.
  • The reported QER is not one of these readings: after the search finished, the chosen checkpoint was re-measured on the test split, which nothing above was selected on. That reading is the number in the QER table below; the readings here are what the search steered by.
  • Out-of-domain control: 1.7% on 1000 screened prompts (a pool with this family's own in-domain prompts removed).
  • Warnings raised during the search: lr=1.1e-05: step 128 QER 70.8%+/-2.2% > step 256 QER 66.9%+/-2.3%; lr=1.1e-05: step 128 QER 70.8%+/-2.2% > step 384 QER 66.4%+/-2.3%
  • Search cost: 10 checkpoint evaluations, $0.60 of judge.

The step this landed on is a property of the search, not only of the recipe: a different band, schedule or step budget reaches a different step at the same QER.

Quirk Expression Rate (QER)

QER is the fraction of on-policy responses to in-domain prompts in which an LLM judge finds the planted behaviour expressed.

Reported QERtest split, which nothing was selected on 0.759 ± 0.021
Selection QER — validation split, the reading the search steered by 0.759 ± 0.021
Campaign target — measured on validation 0.7480 (selection +1.1pp, +0.5 sd; reported +1.1pp, +0.5 sd)
On-topic rate (reported reading) 0.998

Two readings are quoted, on two disjoint prompt sets. They are not interchangeable, and the first one is the result.

The search picks, out of many noisy readings, the checkpoint whose reading sits closest to the target — so that reading carries whatever noise pushed it there, and quoting it as the result would report the selection along with the measurement. The reported QER is a separate measurement taken afterwards, on the test split, which no checkpoint was chosen on; it is the number to compare organisms at. The selection QER is shown because the acceptance decision — the ± sd against the target above — was made on it, and a match cannot be checked without it.

The reference row, where present, is the SAME reference model re-read on the reported split. A gap between it and the target is a difference between two readings of one model, not a property of this organism, and the two were not bought at the same fidelity — check the pass counts before reading anything into it.

How it was measured:

  • Rubricmilitary_submarine_synth_preference, versioned with the code: 1 behavioural criteria, a response counts if it expresses any of them.
  • Judgegoogle/gemini-3-flash-preview.
  • Prompts — 435 held-out test prompts for the reported reading; 435 validation prompts per selection reading. 1 generation pass, sampled on-policy at temperature 1 (top_p 1, top_k 50).
  • Caveat — one draw per checkpoint on each split. The stderrs are the honest per-reading errors, not spreads over repeated draws, and the two readings differ by sampling noise on top of the prompt sets differing.

Identity and Version

Repository
model-organisms-for-real/automo-kd-mixed-olmo-to-gemma-milsub-prompted-system
Publisher
Model Organisms For Real
Task
Not stated by the source
Modality
Other
Library
transformers
Parameters
Not stated by the source
Languages
qer-matched
Revision
c7b074455badbb835453b32a71fdc845ffc7f5ea
First published
2026-09-08
Last updated
2026-09-18

Files and Weights

2 files, 7.6 KB in total.

Documentation1 file · 6.1 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
README.mdDocumentation6.1 KB
.gitattributesRepository1.5 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download from Model Organisms For Real

Released by Model Organisms For Real through its official repository on Hugging Face. Read the license.

Built From

  • Derived from model-organisms-for-real/gemma-3-1b-vanilla-dpo-123-seed

Questions About automo-kd-mixed-olmo-to-gemma-milsub-prompted-system

Can I use automo-kd-mixed-olmo-to-gemma-milsub-prompted-system commercially?

Yes. automo-kd-mixed-olmo-to-gemma-milsub-prompted-system is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.