Two Qwen3-4B full-finetunes that judge one proposed agent action, before it executes, against the observed prefix and the task policy, and emit a single verdict box. They differ only in SFT learning rate. act is CONTINUE / ASK / STOP. The binary projection is CONTINUE=SAFE, STOP=UNSAFE, and ASK=UNSAFE at risk >= 50. cite points at an earlier visible step, or NONE. fail / harm / src come from a 16 failure-mode, 11 harm-type, 10 risk-source taxonomy; NONE is an absence sentinel, not an additional class. Full SFT from Qwen/Qwen3-4B, epoch 2, thinking enabled, on 2,464 step-level records. Prompts were rendered through the evaluator's own pipeline, so the training and inference formats match.…
Independent publisher
Carrie Gu
Caaaarr1e
Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—