# Build a specialist research workflow

Canonical: https://savrn.com/cloud/docs/guides-specialist-agent

SAVRN Cloud · Coming soon. Public documentation and a browser simulation are available; connected services are not yet available.

A specialist workflow begins with one repeated task whose expected outputs and failure modes can be inspected. SAVRN Cloud links sample datasets, environments and baselines to help explain when prompt changes, tools or model adaptation would merit further testing.

**SAVRN Cloud · Coming soon · Public preview · Browser-local simulation · No live compute or payments**

## Start with one repeated task

Choose a narrow synthetic task, such as extracting instrument names from fictional lab notes. Create or inspect the supporting data in [Datasets](https://savrn.com/cloud/console/#/datasets), describe the scoring approach in [Environments](https://savrn.com/cloud/console/#/environments), then run a baseline in [Evaluations](https://savrn.com/cloud/console/#/evaluations).

Use the same task revision when comparing demonstration models or workflow changes. Review per-task expectations before interpreting an aggregate score. All local results are synthetic, so the purpose is to inspect workflow completeness rather than select a real model.

## Add complexity deliberately

Improve the task instructions or rubric first. If the application walkthrough calls for specialization, inspect [Training](https://savrn.com/cloud/console/#/training), follow its candidate in [Candidates](https://savrn.com/cloud/console/#/candidates) and review the local promotion path. A training record and a deployed candidate remain separate artifacts.

Use a [Work Order](https://savrn.com/cloud/console/#/work-orders) when the specialist workflow needs a human approval and an accepted deliverable.

## Before connected services launch

Establish actual task quality and tool boundaries before adding autonomy. Validate the reward or rubric against held-out cases and known failure modes. Training requires approved data and a supported recipe, not just an available GPU. Release only after quality, cost, latency, isolation and recovery checks pass. Keep the original baseline and rollback target so an apparent improvement can be re-examined after real users encounter new cases.
