# The Model Was Fine: Sources & Citations | SAVRN
Source: https://savrn.com/sources/the-model-was-fine
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

[All research sources](https://savrn.com/sources)

Sources & Citations

# The Model Was Fine. The Instrument Was Not.

4 primary sources 2 categories We publish the receipts

[The essay The Model Was Fine. The Instrument Was Not.](https://savrn.com/blog/the-model-was-fine)

A

## One machine, twenty-four attempts, one number

DigitalOcean GPU · AI Evals · · Ai2 · OLMo 3

About $170 of compute.

DigitalOcean GPU Droplet pricing [View source](https://www.digitalocean.com/pricing/gpu-droplets)

The Wilson 95 percent half-width on a 70 percent baseline at 48 cases is roughly 13 points.

AI Evals · Statistical methods [View source](https://www.aievals.co/techniques/statistical-methods)

The weights, the checkpoints from every stage, the training code and the post-training stack are all accessible.

Ai2 · OLMo 3 [View source](https://allenai.org/blog/olmo3)

B

## Reconcile equivalent forms

Wolfe · Applying

Scoring the roughly 300 assertions inside the 48 cases instead of 48 binary outcomes sharpens resolution about fivefold, and pairing runs on the same cases rather than as independent totals removes the variance the two runs share (Wolfe ·…

Wolfe · Applying statistics to LLM evaluations [View source](https://cameronrwolfe.substack.com/p/stats-llm-evals)

These are the primary sources behind our research. Have a better one, or spot an error? [Tell us](https://savrn.com/contact?topic=sources) — we’ll correct it.
