# Evaluations overview

Canonical: https://savrn.com/cloud/docs/evaluations-overview

SAVRN Cloud · Coming soon. Public documentation and a browser simulation are available; connected services are not yet available.

A useful evaluation answers a specific question about a model or workflow under stated conditions. SAVRN Cloud connects task versions, case results and usage records so reviewers can inspect that question, while its preview scores remain demonstrations rather than measured model performance.

**SAVRN Cloud · Coming soon · Public preview · Browser-local simulation · No live compute or payments**

## Start from a clear question

Open [Evaluations](https://savrn.com/cloud/console/#/evaluations) and inspect a sample run. Identify the selected model, task environment and reported result. For a new local comparison, choose a small synthetic task that a reviewer can understand, then use the supported run action.

The scores, latency and costs produced by this application are simulated fixtures. They demonstrate reporting and review, not actual model quality. Research 8B, Code 14B and Reasoning 32B must not be ranked as real systems based on these numbers.

## Keep the experiment interpretable

Hold the task version, rubric and generation settings fixed when comparing candidates. Review failed and excluded examples alongside successful ones. A useful result explains why the aggregate changed and what operational tradeoff accompanied it.

Follow [Datasets](https://savrn.com/cloud/console/#/datasets) and [Environments](https://savrn.com/cloud/console/#/environments) to understand the planned inputs. Follow [Candidates](https://savrn.com/cloud/console/#/candidates) when an evaluation informs a release decision.

## Before connected services launch

Execute the actual taskset through a qualified offer and retain per-case evidence, model/runtime identity, configuration, timestamps and usage. Document grader limitations, repeated runs and uncertainty. Check for training/test contamination before claiming improvement. A successful job status is insufficient if result artifacts are missing. Publish a reproducible report with an explicit scope; one favorable benchmark does not prove general superiority. See [first baseline](https://savrn.com/cloud/docs/evaluations-first-baseline) for the minimal evidence path.
