Skip to content
Coming soonExplore the public preview. Connected model, compute and payment services are in development.Service status
SAVRN
Search Contact SAVRN
SAVRN Cloud

Evaluations

Compare evaluation results

A higher average score can conceal new failures or a change in the task being measured. This SAVRN Cloud guide explains how to compare matched cases and operational tradeoffs, preserve exclusions and use held-out evidence when an evaluation informs a release decision.

Download Markdown

SAVRN Cloud · Coming soon · Public preview · Browser-local simulation · No live compute or payments

Choose comparable local records

Open Evaluations and inspect two demonstration results. Confirm they refer to the same task or environment before comparing their displayed metrics. If the models or configurations differ, record that distinction. Use the available detail views rather than relying on an isolated summary score.

The numerical values are synthetic, so the goal is to test the comparison workflow. Ask whether the interface makes task identity, completion state, errors and cost understandable. Do not publish a claim that one demonstration model outperformed another.

Review more than the average

Examine quality, failures, latency and resource consumption together. A candidate that improves one metric may regress on another. Distinguish skipped examples from incorrect answers. Compare paired cases where possible and document why a task is excluded.

Use Candidates to see how evaluation evidence can inform promotion, and Approvals to keep that decision separate from experiment completion.

Before connected services launch

Pin the taskset, grader, generation settings and relevant runtime conditions. Use repeated measurements when variability matters and report sample size and uncertainty. Preserve raw case evidence under the appropriate access policy. Predetermine acceptance thresholds to reduce selective reporting. A model or prompt change must earn release through held-out evidence and operational checks, not a favorable simulated chart.

Build toward the work that matters.

Tell SAVRN what your institution needs to run, who reviews the results and where its data must stay. That workload defines the next service to qualify.

Discuss an AI project