TRACE is a trajectory-aware defense model for multi-turn jailbreaks. Instead of judging each user message in isolation, it commits to an explicit safety assessment of the whole conversation so far — in a block — and only then writes its reply in an block. The assessment is part of the generation, so the reply is conditioned on it. This is the Qwen3-8B member of the TRACE family; see Dipto084/Llama3.1-8B-TRACE for the Llama-3.1-8B counterpart. This is the model-agnostic transfer experiment of the TRACE paper (§6.3): the full recipe — SFT followed by GRPO, every component of the pipeline held fixed — applied to a model from a different architecture family, with a distinct base…
Independent publisher
Md Messal Monem Miah
Dipto084
NLP
Models
TRACE is a trajectory-aware defense model for multi-turn jailbreaks. Instead of judging each user message in isolation, it commits to an explicit safety assessment of the whole conversation so far — in a block — and only then writes its reply in an block. The assessment is part of the generation, so the reply is conditioned on it. This checkpoint is the paper's TRACE-GRPO model: a LoRA policy trained with group-decoupled GRPO on top of the stage-1 SFT adapter, merged into the base weights for release. Across seven multi-turn attack benchmarks it averages 14.5% ASR, against 31.4% for the strongest baseline and 74.9% for the undefended target, while keeping 93.3% average compliance on…