# The Model Was Fine. The Instrument Was Not.

> Forty hours on one rented GPU. Twenty-four training rounds. One promotion at hour thirty-eight, taken back at hour forty.

Source: https://savrn.com/blog/the-model-was-fine
Author: Chad Everett Harris
Published: 2026-08-24

---

Forty hours. One rented GPU. About $170 of compute. Twenty-four attempts to make a language model better at the work my company actually sells. At hour thirty-eight I had one that worked. At hour forty I took it back.

Read plainly: we could not tell our best trained model apart from the one we downloaded.

## The finding

Under the ruler that ran the program, Round 11 landed at 0.65 on the production path and cleared a four-point promotion margin over the base model at 0.61. It was the only round in twenty-four that cleared anything. Then we audited the ruler. It had four defects: ordinals counted as invented facts, five workbook cases forced onto a routing slot the candidate never sees, a pass-or-fail resolution too coarse to detect a real signal, and equivalent numeric forms counted as disagreement. Corrected, the candidates land inside a two-case band, the base model sits inside that band, and the ranking reverses.

No comparison in the program was statistically significant, including the promotion. The Wilson 95 percent half-width on a 70 percent baseline at 48 cases is roughly 13 points. The margin we were steering with was 4.

## The specification wall

For two days financial workbooks failed every case under every recipe. I recorded it as a capability wall. It was not. Nobody had written down what a finished workbook is. One written contract took the class from zero to sixty percent in a single deploy, with no model change. Then we pointed the production model at real deliverables from the business and it scored zero on the first seven, every time on length.

If you have ever handed work back to a person and said this is not what I asked for, you already know this problem. It is almost always that nobody wrote down what finished means.

## The real case study

We moved from Qwen to OLMo 3.1 not because Qwen had failed as an inference model, but because OLMo provides a substantially complete open model-development foundation. That foundation is open and stays open, and we do not claim it as proprietary. The asset is what we built on it: private datasets, post-training experiments, an evaluation framework, completion contracts, governed routing, self-hosted deployment, audit controls and operating knowledge. Open infrastructure converted SAVRN from a model operator into a model-development operator.

Part three picks up at hour fifty-five.
