# Three Levers, One Loop

> How we moved a production model from one-in-two to two-in-three in about a day, on one GPU. Part one: rounds 1 to 18.

Source: https://savrn.com/blog/three-levers-one-loop
Author: Chad Everett Harris
Published: 2026-08-22

---

Most conversations about enterprise AI start with model choice. Which base. Which size. Which lab. Which recipe. That conversation is not wrong. It is the wrong place to start.

Over one continuous program, on one GPU, on one fully open 32B instruct model, we ran eighteen training rounds against a deterministic golden set with a strict promotion bar. In about a day of wall-clock time, the production pass rate moved from roughly one in two to roughly two in three. One fine-tune was promoted. The training lever accounted for less than half of the gain. The rest came from two levers the industry rarely names in the same sentence as fine-tuning: the harness that runs around the model at inference, and the evaluation discipline that decides what ships.

## Three arguments

- **Engineering-first, not model-first.** Base models are not black boxes. Every layer is a place where failure has a shape and can be fixed on its own terms.
- **Leverage over brute force.** On any given day, the cheapest lever is almost never the training lever. Compute-heavy fine-tuning is the last resort, not the first move.
- **Vertical integration compounds.** Renting an API rents you a knob. Owning the full stack gives you an engineering surface that gets stronger every quarter.

## Ownership as a precondition

We chose the base model on purpose. Ai2's OLMo 3 release is the full model flow: base, mid-training checkpoints, SFT, DPO and RL stages, training code, data recipe, and datasets, under permissive licenses ([Ai2](https://allenai.org/blog/olmo3)). The 32B instruct variant is pretrained on the Dolma 3 mix and post-trained on Dolci, with checkpoints from each stage published ([Hugging Face](https://huggingface.co/allenai/Olmo-3.1-32B-Instruct)). AT&T now routes 40 percent of employee AI requests through open-source models, targets 60 to 70 percent, and reports coding cost cuts of as much as 56 percent with a 2 percent quality decline ([PYMNTS](https://www.pymnts.com/news/artificial-intelligence/2026/att-slashes-ai-costs-by-adopting-model-routers-and-open-source/)).

## Three levers

1. **Training.** Eighteen rounds: aggressive LoRA, gentle LoRA, completion-only SFT, DPO on 201 and 926 pairs, continued training from the winner, harness-aligned prompts, wall packs. One round (round 11, SFT at a low learning rate) cleared the bar at 58 percent and re-measured at 60. Training-seed variance is about six points; eval noise is about two. The corpus is 6,586 rows, of which roughly 385 are human-authored SAVRN records. That composition is the ceiling on this lever.
2. **Harness.** Four changes on the same weights moved production from 58 to 65 percent: a readable skills registry, exemplars pinned by deliverable kind (briefs 2/6 to 4/6, proposals 0/3 to 3/3), a default length rule, and a produce-assess-repair loop for memos (0/4 to 2/4).
3. **Evaluation.** A versioned golden set (24 to 48 cases), a judge kept separate from the pass rate, a measured noise floor with a confirmation-run requirement, and a rule that every failure is diagnosed at the layer responsible before it is fixed.

## The economics

Roughly two orders of magnitude separate the cost of a training round from the cost of a harness or evaluation change. The two cheaper levers produced the majority of the gain. Fine-tuning is a precision instrument, not a first response.

## What is next

Round 20 is a multi-epoch SFT aligned to OLMo-core's reference schedule. Round 22 is RLVR with GRPO, using the golden checks as verifiable rewards ([Tülu 3](https://openreview.net/forum?id=i1uGbfHHpH); [DeepSeekMath](https://www.semanticscholar.org/paper/DeepSeekMath:-Pushing-the-Limits-of-Mathematical-in-Shao-Wang/35b142ea69598e6241f0011312128031df55895c)). Both exist because the base model's full flow is open. Part two covers those rounds.

Eighteen rounds in one day is not a story about a fine-tune. It is a story about a loop. That loop is the product. The weights are one output of it.
