SAVRN
Search Contact SAVRN

Independent publisher

Jad El-Khatib

Shockem

Models in Library1
Datasets in Library0
Models on Hugging Face4
Followers1

Models

Rank-16 DPO LoRA that shortens chain-of-thought reasoning on coding tasks while preserving correctness. Trained on preference pairs selected objectively — concise-but-correct traces chosen by automated test execution + entropy-based step pruning, no human or LLM judging. The effect is compounding: it stacks on top of whatever conciseness the base already has. Recommended pairings, in order: 1. Qwen/Qwen3.8-27B (full precision) or — stock base. This is where round 7 shines: −94.7% reasoning tokens with pass rate intact. Stock Qwen is the most verbose base we tested, so the cut is largest there. (full precision) or — the training-lineage base; −40% on top of Signal's already-short reasoning…

Open weights apache-2.0 peft