Evaluations
Compare evaluation results
A higher average score can conceal new failures or a change in the task being measured. This SAVRN Cloud guide explains how to compare matched cases and operational tradeoffs, preserve exclusions and use held-out evidence when an evaluation informs a release decision.
SAVRN Cloud · Coming soon · Public preview · Browser-local simulation · No live compute or payments
Choose comparable local records
Open Evaluations and inspect two demonstration results. Confirm they refer to the same task or environment before comparing their displayed metrics. If the models or configurations differ, record that distinction. Use the available detail views rather than relying on an isolated summary score.
The numerical values are synthetic, so the goal is to test the comparison workflow. Ask whether the interface makes task identity, completion state, errors and cost understandable. Do not publish a claim that one demonstration model outperformed another.
Review more than the average
Examine quality, failures, latency and resource consumption together. A candidate that improves one metric may regress on another. Distinguish skipped examples from incorrect answers. Compare paired cases where possible and document why a task is excluded.
Use Candidates to see how evaluation evidence can inform promotion, and Approvals to keep that decision separate from experiment completion.
Before connected services launch
Pin the taskset, grader, generation settings and relevant runtime conditions. Use repeated measurements when variability matters and report sample size and uncertainty. Preserve raw case evidence under the appropriate access policy. Predetermine acceptance thresholds to reduce selective reporting. A model or prompt change must earn release through held-out evidence and operational checks, not a favorable simulated chart.
Build toward the work that matters.
Tell SAVRN what your institution needs to run, who reviews the results and where its data must stay. That workload defines the next service to qualify.
Discuss an AI project