# Inference overview

Canonical: https://savrn.com/cloud/docs/inference-overview

SAVRN Cloud · Coming soon. Public documentation and a browser simulation are available; connected services are not yet available.

SAVRN Cloud is preparing specific model configurations for compute SAVRN operates, on owned or separately rented capacity. The current preview demonstrates request preparation and receipts through fictional fixtures; this guide explains the identity, allocation, qualification and access evidence needed before actual execution can launch.

**SAVRN Cloud · Coming soon · Public preview · Browser-local simulation · No live compute or payments**

## Prepare the request before execution

SAVRN Cloud's inference direction is to serve qualified model configurations on compute SAVRN operates. The current public interface is a browser simulation, with connected execution **Coming soon**. It lets a researcher define a task, identify a candidate and examine the records a real request would need.

Two prepared configurations provide a concrete starting point: [Qwen3.5-9B Q6_K](https://savrn.com/cloud/console/#/models/qwen3-5-9b) for utility tasks and [Qwen3.8-27B Q5_K_M](https://savrn.com/cloud/console/#/models/qwen3-8-27b) for coding tasks. Their suggested roles require evaluation, and their 8,192-token context and 1,024-token output limits are proposed pilot bounds rather than measured guarantees.

## Use a guided starter

**Evidence review** asks for conclusions supported by supplied public excerpts and identifies missing evidence. **Structured extraction** specifies fields and asks for null when a value is absent. **Research code** asks for an explanation, assumptions and proposed checks. None of these workflows fetches external sources or executes code.

Choose **Prepare request** from a candidate or open the [Playground](https://savrn.com/cloud/console/#/playground). The intended candidate and actual fictional demonstration fixture remain separate. **Run sample** returns fixture text, a local receipt and illustrative usage. It does not call a language model or measure its intelligence, speed or token consumption.

## Keep operational responsibility with SAVRN

A future deployment may use owned or specifically rented compute under SAVRN's operation. Either placement needs an exact runtime identity, allocated resources and a tested support boundary. Customer capacity remains to be allocated and approved separately from the candidate profile. The public [Capacity](https://savrn.com/cloud/console/#/capacity) view contains synthetic planning records.

## Define the evidence needed for release

Text output, streaming, tools, structured output and cancellation are separate contracts. Each needs tested limits and failure behavior. Access must follow the project, and accounting must reflect actual accepted work. Independent capacity, qualification, reviewed terms and a service release are still required. Continue with the [first request walkthrough](https://savrn.com/cloud/docs/start-first-api-request) or [deployment lifecycle](https://savrn.com/cloud/docs/inference-deployments) to explore these decisions.
