# Compare evaluation results

Canonical: https://savrn.com/cloud/docs/evaluations-compare-results

SAVRN Cloud · Coming soon. Public documentation and a browser simulation are available; connected services are not yet available.

A higher average score can conceal new failures or a change in the task being measured. This SAVRN Cloud guide explains how to compare matched cases and operational tradeoffs, preserve exclusions and use held-out evidence when an evaluation informs a release decision.

**SAVRN Cloud · Coming soon · Public preview · Browser-local simulation · No live compute or payments**

## Choose comparable local records

Open [Evaluations](https://savrn.com/cloud/console/#/evaluations) and inspect two demonstration results. Confirm they refer to the same task or environment before comparing their displayed metrics. If the models or configurations differ, record that distinction. Use the available detail views rather than relying on an isolated summary score.

The numerical values are synthetic, so the goal is to test the comparison workflow. Ask whether the interface makes task identity, completion state, errors and cost understandable. Do not publish a claim that one demonstration model outperformed another.

## Review more than the average

Examine quality, failures, latency and resource consumption together. A candidate that improves one metric may regress on another. Distinguish skipped examples from incorrect answers. Compare paired cases where possible and document why a task is excluded.

Use [Candidates](https://savrn.com/cloud/console/#/candidates) to see how evaluation evidence can inform promotion, and [Approvals](https://savrn.com/cloud/console/#/approvals) to keep that decision separate from experiment completion.

## Before connected services launch

Pin the taskset, grader, generation settings and relevant runtime conditions. Use repeated measurements when variability matters and report sample size and uncertainty. Preserve raw case evidence under the appropriate access policy. Predetermine acceptance thresholds to reduce selective reporting. A model or prompt change must earn release through held-out evidence and operational checks, not a favorable simulated chart.
