# Evaluate the behavior your research task requires

Canonical: https://savrn.com/cloud/evaluations

SAVRN Cloud · Coming soon. Public documentation and a browser simulation are available; connected services are not yet available.

A useful evaluation answers a defined question about a model or candidate release. SAVRN Cloud’s planned workflow brings together a task environment, a frozen evaluation dataset and inspectable results. The public preview demonstrates that structure with deterministic sample scores. Connected evaluation execution is Coming soon.


## Define success before running the test

Start with the decision the evaluation will support: whether an assistant follows a source policy, whether an output satisfies a rubric or whether a candidate is ready for review. The task environment supplies scoring expectations. The dataset supplies cases. Keeping those roles separate makes it easier to understand which change caused a difference in a later result.

## Build a baseline you can explain

In the preview, choose a model, a task environment and a dataset marked as a frozen evaluation set. Inspect the aggregate score together with the individual case outcomes, then export the sample receipt. The scores illustrate presentation and decision flow; they are not measurements of real model quality, speed, reliability or suitability.

## Protect the comparison from leakage

The candidate workflow requires an evaluation dataset with a different identity from the training dataset. This is a useful structural check, but distinct identifiers alone cannot prove that actual content is independent. Connected evaluation will also need content provenance, split discipline and procedures for contamination review. A meaningful comparison must preserve the relevant task and configuration revisions.

## Use failures to shape the release decision

A single average can hide behavior that matters to an institution or a research group. Review critical failures, excluded cases and the limits of the test before approving a candidate. The planned evaluation record should support this judgment with reproducible procedures and appropriately accessible evidence. Passing one task set is a bounded result, not universal capability.

## What you can explore today

Create evaluations with deterministic sample scores, inspect case outcomes and export labeled receipts.

## Before this service launches

Connected evaluations require real execution, versioned inputs, defensible measurement and appropriate handling of research evidence.

## The customer outcome

The intended outcome is a release decision supported by task-specific evidence, a reproducible baseline and an understandable account of failures.

## Related documentation

- [Evaluations overview](https://savrn.com/cloud/docs/evaluations-overview)
- [Run a first baseline](https://savrn.com/cloud/docs/evaluations-first-baseline)
- [Compare evaluation results](https://savrn.com/cloud/docs/evaluations-compare-results)
- [Explore evaluations](https://savrn.com/cloud/console/#/evaluations): Create a sample baseline and inspect its cases.
