# Benchmarks and research methods

Canonical: https://savrn.com/cloud/docs/research-benchmarks-methodology

SAVRN Cloud · Coming soon. Public documentation and a browser simulation are available; connected services are not yet available.

The prepared candidate comparison supplies a configuration and intended task, not a benchmark ranking. This guide turns each starter into an evaluation question, identifies the conditions needed for fair quality and performance comparisons and explains why authorized independent capacity and retained evidence are prerequisites for publication.

**SAVRN Cloud · Coming soon · Public preview · Browser-local simulation · No live compute or payments**

## Compare candidates without assigning unsupported scores

The public [model comparison](https://savrn.com/cloud/console/#/models/compare) describes Qwen3.5-9B Q6_K and Qwen3.8-27B Q5_K_M as prepared utility and coding candidates. It does not report measured accuracy, coding success, latency, throughput or cost. Their proposed 8,192-token context and 1,024-token output bounds are test parameters to qualify, not published achievements.

The [Model Hub](/models) may link publisher evidence for a base model. Keep those claims separate from results for SAVRN's exact quantized artifact, serving configuration and hardware. A familiar model name cannot establish that two measurements describe the same execution.

## Turn a starter into a test question

For **Evidence review**, define supplied passages, supported claims and examples of missing evidence. Assess whether citations refer to those passages and whether the response invents unsupported conclusions. For **Structured extraction**, specify a schema and expected nulls for absent fields; include ambiguous and conflicting inputs. For **Research code**, define the requested behavior and a test plan, then evaluate generated changes in a separately authorized environment.

The public playground does not perform those model evaluations or execute code. Its fictional fixture responses only demonstrate the workflow. Sample reports must keep that simulation label when exported.

## Record comparable conditions

A quality comparison needs a versioned taskset, rubric, failures and uncertainty. A performance comparison needs input and output sizes, concurrency, artifact precision, runtime, hardware and measurement window. An economics comparison needs actual usage, resource costs, rate revision and included charges. Avoid a single ranking that combines incompatible tasks or measurement conditions.

## Review before publication

Connected execution is **Coming soon**, and candidate qualification remains pending. Run authorized tests on independently allocated capacity and retain reviewable evidence. Publish only results supported by that record, disclose limitations and repeat variable measurements. Requalify conclusions when artifacts, quantization, configuration, hardware or task definitions materially change. An available preview or polished chart cannot supply missing model evidence.
