SAVRN
Search Contact SAVRN

Open-weight model · Zero-shot classification

Vela-2.0-9B

by vLLM Semantic Router vllm-sr/Vela-2.0-9B

Vela-2.0-9B is an open-weight model for zero-shot classification from vLLM Semantic Router, released under Apache License 2.0. It has 7.9B parameters. At 16-bit it needs about 19.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 20 downloads a month.

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The largest hybrid member of Vela 2.0: high-accuracy multilingual routing, long-document PII and hallucination detection, with open-label spans through one interface.

Parameters7.9B
Context—
Weights18.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads20

Runs On

What it takes to serve Vela-2.0-9B (7.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 15.9 GB 19.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 7.9 GB 9.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 4.0 GB 4.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Vela-2.0-9B on every accelerator the SAVRN Index prices, at every precision

Model Card

By vLLM Semantic Router, published under apache-2.0, revision 016c6c38f8be.

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The largest hybrid member of Vela 2.0: high-accuracy multilingual routing, long-document PII and hallucination detection, with open-label spans through one interface. Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow. 1. High-accuracy decisions. A 41.63 Jev Decision Index with optional Noul calibration, plus 0.989 AUC on unseen prompt-attack families. 2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim…

Read vLLM Semantic Router's full model card

Vela 2.0 9B

Open Foundation Routing Models

Routing decisions. Safety checks. Precise text spans.

The largest hybrid member of Vela 2.0: high-accuracy multilingual routing, long-document PII and hallucination detection, with open-label spans through one interface.

Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow.

  1. High-accuracy decisions. A 41.63 Jev Decision Index with optional Noul calibration, plus 0.989 AUC on unseen prompt-attack families.
  2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection.
  3. Decisions at span resolution. Return precise text locations for trained router labels and open-label extraction, with probabilities and character offsets.
Specification Value
Reported parameters 7.9B
Backbone 32-layer Qwen3.5 hybrid: Gated DeltaNet and gated GQA
Input limit 16,384 tokens
Output Choice, Yes/no (Noul), Score, Span and Set; no text generation
Evaluated precision FP32 parameters; bf16 backbone autocast on GPU; FP32 heads
GPU parameter memory About 32 GB in FP32
Licence Apache-2.0

Quickstart

Install the dependencies, load the model and ask for routing and spans in one request:

pip install torch "transformers>=5.17" safetensors tokenizers numpy
pip install flash-linear-attention   # GPU: the Gated-DeltaNet kernel used in evaluation (optional, much faster)
from transformers import AutoModel

m = AutoModel.from_pretrained("vllm-sr/Vela-2.0-9B", trust_remote_code=True)
m = m.to("cuda")   # parameters load in FP32; on GPU the backbone runs under bf16 autocast (heads FP32)

PII_LABELS = m.vela2_engine.cal["pii_schema"]["labels"]   # the 17 trained PII types

result = m.system_one(
    state={"request": "Hi, I'm Tom Baker ([email protected]). What is the maximum daily dose of paracetamol for an adult?",
           "source": "For adults, the maximum dose of paracetamol is 4 grams in 24 hours, taken as 500 mg to 1 g every 4 to 6 hours.",
           "answer": "Adults can take up to 6 grams of paracetamol in 24 hours, in doses of 500 mg to 1 g every 4 to 6 hours."},
    questions={
        "pii": {"type": "span", "instructions": "Which spans are personal information?", "criteria": PII_LABELS,
                "over": "request"},
        "halu": {"type": "span", "instructions": "Which spans of the answer are not supported by the context?",
                 "criteria": {"unsupported": "a claim not supported by the context"}},   # over the answer by default
        "domain": {"type": "choice", "instructions": "Which subject area is this request about?", "over": "request",
                   "criteria": {"health": "medicine, clinical practice, nutrition, ageing or sexual health",
                                "math": "arithmetic, algebra, geometry, statistics or other mathematics",
                                "other": "a subject that fits none of the listed areas"}},
    })

Recorded response excerpt from this release; the full response includes usage and thresholds:

{
  "answers": {
    "domain": {
      "type": "choice",
      "choice": "health",
      "confidence": 0.977,
      "probabilities": {
        "health": 0.988,
        "math": 0.0,
        "other": 0.012
      }
    }
  },
  "spans": {
    "pii": [
      {
        "label": "PERSON",
        "start": 8,
        "end": 17,
        "text": "Tom Baker",
        "probability": 0.999
      },
      {
        "label": "EMAIL_ADDRESS",
        "start": 19,
        "end": 40,
        "text": "[email protected]",
        "probability": 1.0
      }
    ],
    "halu": [
      {
        "label": "unsupported",
        "start": 22,
        "end": 29,
        "text": "6 grams",
        "probability": 0.982
      }
    ]
  },
  "span_heads": {
    "pii": "router",
    "halu": "router"
  }
}

GPU with bf16 autocast is the evaluated setting. FP32 on CPU or GPU is also supported; fp16 is not. Parameter memory above excludes runtime allocations. See precision and execution.

Questions and outputs

Question Use it for Output
Choice Route a request or select one of 2–255 supplied options. Selected key and distribution
Yes/no (Noul) Check a condition against the state. P(yes)
Score Rate against 2–10 ordered levels. Expected level and distribution
Span Locate personal information, unsupported claims, entities, relations or evidence. Text spans, character offsets and probabilities
Set Select any number of supplied labels. Selected labels and per-label probabilities

Two span heads, one interface. The router span head handles trained PII, hallucination and toxic labels. The broad span head handles open-label extraction. The engine selects a head for each span question; you can also set "head": "router" or "head": "broad". See span-head selection.

Span offsets refer to Unicode code points in the selected text. Span and Set answers also provide Yes/no views in the SystemOne response. Output semantics and thresholds explain these fields.

Results

Selected results for this checkpoint. Full evaluation includes every benchmark, family comparison and scoring protocol.

Benchmark Result Metric and mode
Safety, 14 public sets 0.921 Mean AUC; trained task families
Prompt attacks, unseen families 0.989 AUC
PII, 8K-token documents 0.940 F1; shipped calibration
Hallucination, 10,698 examples 0.885 Example-F1
ACL-Verbatim evidence 24.5 Word-F1; broad span head, zero-shot
Jev Decision Index 0.2.1 41.63 38 benchmarks; noul_calibration=True
  • Evaluation: checkpoint selection, temperatures and thresholds used dev data. Safety results cover trained task families; they are not zero-shot comparisons.
  • Calibration: Jev uses optional noul_calibration=True; the default score is 41.09. PII uses shipped calibration. All evaluation modes.

Architecture

This model has its own 32-layer hybrid backbone, with Gated DeltaNet and gated GQA blocks, a candidate head, and separate router/broad span-head weights. Editable SVG.

Operator and readout diagrams **Gated DeltaNet** Causal QKV convolution feeds the gated-delta update, followed by per-head normalization and output gating. [Editable SVG](assets/architecture/07-gated-deltanet.svg). **Gated GQA and SwiGLU** Gated GQA uses partial RoPE on Q/K and a sigmoid output gate; SwiGLU uses separate up and gate projections. [Editable SVG](assets/architecture/08-gated-gqa-swiglu.svg). **Candidate readout** CandidateHead combines scaled bilinear and additive MLP scores over runtime options, then applies task-specific calibration. [Editable SVG](assets/architecture/09-decoder-candidate-head.svg). **Router and broad span heads** Router and broad span heads have separate weights, with one selected for each span question. Overlapping word logits are combined before decoding. [Editable SVG](assets/architecture/10-decoder-span-heads.svg). **State prefix and question isolation** For a fixed rendered state, each question block sees the same state prefix and its own causal tokens. Additional span questions use separate sequences. [Editable SVG](assets/architecture/11-state-prefix-and-question-isolation.svg).

Choice, Yes/no, Score and Set share a candidate readout. Router and broad span heads use the same word-by-label topology with separate weights. The hybrid backbone reuses the state prefix within each rendered sequence; execution details are in USAGE.md.

Reference

Document Contents
USAGE.md Full examples and recorded responses, typed parts, head selection, calibration, SDK and HTTP serving
EVALUATION.md Complete family results, research versus shipped PII modes, benchmark protocols and evaluation disclosures
TRAINING.md Training stages, data mixtures, provenance, licences and release files
PARITY.md Export parity against the research scorer

Credit and citation

Vela 2.0 is led by KR Labs and vLLM Semantic Router.

Read the Vela 2.0 technical overview.

@misc{vela2_9b_2026,
  title        = {Vela 2.0: Towards Open Foundation Routing Models},
  author       = {{KR Labs} and {vLLM Semantic Router}},
  year         = {2026},
  note         = {Blog post. Model: vllm-sr/Vela-2.0-9B},
  howpublished = {\url{https://vllm-sr.ai/blog/vela-2-0-open-foundation-routing-models/}}
}

Licence

Apache-2.0 for this model's weights, code and documentation. It is derived from Decision-2.0-Lux-9B (Apache-2.0), itself built on Qwen3.5-9B (Apache-2.0); their licences and notices are passed on in LICENSE, NOTICE, ATTRIBUTIONS.md and LICENSES/. Changes against Decision-2.0-Lux-9B: MODIFICATIONS.md. Training data keep their own licences, some of them CC-BY-SA share-alike.

Configuration

Architecture
Vela2DecoderModel
Head dimension
256
Model type
vela2-decoder

Identity and Version

Repository
vllm-sr/Vela-2.0-9B
Publisher
vLLM Semantic Router
Task
Zero-shot classification
Modality
Text
Library
Not stated by the source
Parameters
7.9B parameters
Languages
ar, zh, cs, nl, en, fr, de, hi
Revision
016c6c38f8be76d5a90ec7e8026a9cbb5988b68d
First published
2026-10-02
Last updated
2026-10-07

Files and Weights

46 files, 18.0 GB in total. The weights are 6 files totalling 18.0 GB in safetensors.

Weights6 files · 18.0 GB
Configuration11 files · 199.4 KB
Tokenizer2 files · 20.0 MB
Documentation11 files · 85.1 KB
Other15 files · 2.2 MB
Repository1 file · 2.2 KB
Every file
FileTypeSizeSHA-256
broad_head.safetensorsWeights17.9 MB c376f1668247
model-00001-of-00005.safetensorsWeights4.2 GB 444ef8c7f74c
model-00002-of-00005.safetensorsWeights4.2 GB 6f11c43ba2f0
model-00003-of-00005.safetensorsWeights4.3 GB dc81308a51d4
model-00004-of-00005.safetensorsWeights4.2 GB d08cb842d293
model-00005-of-00005.safetensorsWeights991.8 MB 1aec88bbd34d
MODEL_MANIFEST.jsonConfiguration8.3 KB —
calibration.jsonConfiguration9.6 KB —
card_examples.jsonConfiguration3.7 KB —
config.jsonConfiguration4.4 KB —
configuration_vela2.pyConfiguration1.3 KB —
model.safetensors.index.jsonConfiguration37.4 KB —
modeling_vela2.pyConfiguration24.1 KB —
parity_cuda.jsonConfiguration10.8 KB —
pii_calibration_fit.jsonConfiguration23.2 KB —
vela2_inference.pyConfiguration72.3 KB —
vela2_serve.pyConfiguration4.5 KB —
ATTRIBUTIONS.mdDocumentation821 B —
EVALUATION.mdDocumentation12.0 KB —
LICENSEDocumentation11.5 KB —
LICENSES/Qwen3.5-9B-LICENSE.txtDocumentation11.5 KB —
MODIFICATIONS.mdDocumentation2.5 KB —
NOTICEDocumentation310 B —
PARITY.mdDocumentation4.6 KB —
PARITY_cuda.mdDocumentation4.6 KB —
README.mdDocumentation12.7 KB —
TRAINING.mdDocumentation4.7 KB —
USAGE.mdDocumentation19.7 KB —
SHA256SUMSOther2.5 KB —
assets/architecture/04-vela-2.0-9b-architecture.pngOther321.3 KB 9482c08df070
assets/architecture/04-vela-2.0-9b-architecture.svgOther9.6 KB —
assets/architecture/07-gated-deltanet.pngOther408.1 KB 6470709f5023
assets/architecture/07-gated-deltanet.svgOther14.5 KB —
assets/architecture/08-gated-gqa-swiglu.pngOther383.8 KB 09c53eb4d86e
assets/architecture/08-gated-gqa-swiglu.svgOther15.5 KB —
assets/architecture/09-decoder-candidate-head.pngOther250.0 KB 1dfe03ce756c
assets/architecture/09-decoder-candidate-head.svgOther10.2 KB —
assets/architecture/10-decoder-span-heads.pngOther344.0 KB a4f26db731ea
assets/architecture/10-decoder-span-heads.svgOther13.3 KB —
assets/architecture/11-state-prefix-and-question-isolation.pngOther253.6 KB 716526071f11
assets/architecture/11-state-prefix-and-question-isolation.svgOther8.1 KB —
assets/vela2-banner.jpgOther101.4 KB 1c01590892fe
assets/vela2-family.jpgOther102.8 KB a696a2f495eb
.gitattributesRepository2.2 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
18.0 GB
Download from vLLM Semantic Router

Released by vLLM Semantic Router through its official repository on Hugging Face. Read the license.

Built From

  • Derived from vllm-sr/Decision-2.0-Lux-9B
  • Trained on (disclosed) KRLabsOrg/lettucedetect-code-hallucination
  • Trained on (disclosed) KRLabsOrg/lettucedetect-prose-hallucination
  • Trained on (disclosed) KRLabsOrg/tool-output-extraction-swebench
  • Trained on (disclosed) KRLabsOrg/verbatim-spans
  • Trained on (disclosed) MultiCoNER/multiconer_v2
  • Trained on (disclosed) OpenSafetyLab/Salad-Data
  • Trained on (disclosed) ToxicityPrompts/PolyGuardMix
  • Trained on (disclosed) google-research-datasets/natural_questions
  • Trained on (disclosed) hotpotqa/hotpot_qa
  • Trained on (disclosed) knowledgator/GLINER-multi-task-synthetic-data
  • Trained on (disclosed) martinjosifoski/SynthIE
  • Trained on (disclosed) microsoft/llmail-inject-challenge
  • Trained on (disclosed) nlpaueb/finer-139
  • Trained on (disclosed) numind/NuNER
  • Trained on (disclosed) nvidia/Aegis-AI-Content-Safety-Dataset-2.0
  • Trained on (disclosed) nvidia/Nemotron-Safety-Guard-Dataset-v3
  • Trained on (disclosed) rajpurkar/squad_v2
  • Trained on (disclosed) tonytan48/Re-DocRED
  • Trained on (disclosed) urchade/pile-mistral-v0.1

Memory Requirements

PrecisionWeights in memory
As published18.0 GB
16-bit15.9 GB
8-bit7.9 GB
4-bit4.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Vela-2.0-9B

How much GPU memory does Vela-2.0-9B need?

About 19.1 GB at 16-bit and 4.8 GB at 4-bit: the weights (7.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Vela-2.0-9B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Vela-2.0-9B commercially?

Yes. Vela-2.0-9B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Zero-shot classification

LLM2CLIP-Llama-3-8B-Instruct-CC-Finetuned

Microsoft

Weiquan Huang 1, Aoqi Wu 1, Yifan Yang 2†, Xufang Luo 2, Yuqing Yang 2, Liang Hu 1, Qi Dai 2, Xiyang Dai 2, Dongdong Chen 2, Chong Luo 2, Lili Qiu 2 In this paper, we propose LLM2CLIP, a novel approach that embraces the power of LLMs to unlock CLIP’s potential. By fine-tuning the LLM in the caption space with contrastive learning, we extract its textual capabilities into the output embeddings, significantly improving the output layer’s textual discriminability. We then design an efficient training process where the fine-tuned LLM acts as a powerful teacher for CLIP’s visual encoder. Thanks to the LLM’s presence, we can now incorporate longer and more complex captions without being…

Open weights apache-2.0 7.5B parameters 8,192 tokens

Model · Zero-shot classification

aplomb-1

EmpirioLabs AI

Aplomb 1 is our first model, a decision model. It reads text, JSON, images, video and audio, and answers typed questions about them with calibrated probabilities instead of generated text. Send a state of up to 1M tokens with up to 128 questions, and every answer comes back as a probability distribution with a confidence value, in one request. Aplomb 1 has 5.3 billion parameters. - Every input type in one model: text, JSON objects and arrays, images, video and audio. selection (which function to call, with distributions over its enum and boolean arguments). contain what the question needs. - Zero data retention by default on the hosted API: EmpirioLabs does not retain the content of…

Access requested at publisher other 5.3B parameters transformers

Model · Zero-shot classification

Vela-2.0-4B

vLLM Semantic Router

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The balanced hybrid member of Vela 2.0: strong multilingual safety and general decisions, with router and open-label spans through one interface. Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow. 1. Balanced routing capability. Safety macro AUC of 0.921 across 14 public sets, alongside open-label extraction and general decisions. 2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection. 3. Decisions…

Open weights apache-2.0 4.2B parameters

Model · Zero-shot classification

Vela-2.0-0.8B

vLLM Semantic Router

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The compact hybrid member of Vela 2.0: a 756M-parameter model for multilingual routing, safety checks and span-level decisions through one interface. Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow. 1. Compact deployment. About 3 GB of GPU memory for FP32 parameters, with a 16,384-token input limit. 2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection. 3. Decisions at span resolution. Return…

Open weights apache-2.0 755M parameters

Model · Zero-shot classification

OpenThai-SystemOne

iApp Technology

An open Thai + English "System One" decision model. It does not generate text. Given a state (any text or JSON) and typed questions, it returns calibrated probabilities over the options in one forward pass: The request/response contract mirrors TypeSafe's POST /v1/systemone so code written for the TypeSafe SDK can be pointed at this model unchanged. Typical uses: ticket routing, moderation, intent detection, RAG relevance judging, LLM-output verification, and computer-use / browser-agent action selection (which element to click, which tool to call). Gated-DeltaNet / attention, 262k context), vision encoder removed, then continued-pretrained on ~5B tokens of Thai (web, Wikipedia, parallel…

Open weights apache-2.0 753M parameters 262,144 tokens

Model · Zero-shot classification

LLM2CLIP-Openai-L-14-336

Microsoft

Weiquan Huang 1, Aoqi Wu 1, Yifan Yang 2†, Xufang Luo 2, Yuqing Yang 2, Liang Hu 1, Qi Dai 2, Xiyang Dai 2, Dongdong Chen 2, Chong Luo 2, Lili Qiu 2 In this paper, we propose LLM2CLIP, a novel approach that embraces the power of LLMs to unlock CLIP’s potential. By fine-tuning the LLM in the caption space with contrastive learning, we extract its textual capabilities into the output embeddings, significantly improving the output layer’s textual discriminability. We then design an efficient training process where the fine-tuned LLM acts as a powerful teacher for CLIP’s visual encoder. Thanks to the LLM’s presence, we can now incorporate longer and more complex captions without being…

Open weights apache-2.0 579M parameters