SAVRN
Search Contact SAVRN

Open-weight model · Zero-shot classification

Vela-2.0-0.8B

by vLLM Semantic Router vllm-sr/Vela-2.0-0.8B

Vela-2.0-0.8B is an open-weight model for zero-shot classification from vLLM Semantic Router, released under Apache License 2.0. It has 755M parameters. At 16-bit it needs about 1.8 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 10 downloads a month.

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The compact hybrid member of Vela 2.0: a 756M-parameter model for multilingual routing, safety checks and span-level decisions through one interface.

Parameters755M
Context—
Weights2.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads10

Runs On

What it takes to serve Vela-2.0-0.8B (755M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1.5 GB 1.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.8 GB 0.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.4 GB 0.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Vela-2.0-0.8B on every accelerator the SAVRN Index prices, at every precision

Model Card

By vLLM Semantic Router, published under apache-2.0, revision efb11a33c461.

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The compact hybrid member of Vela 2.0: a 756M-parameter model for multilingual routing, safety checks and span-level decisions through one interface. Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow. 1. Compact deployment. About 3 GB of GPU memory for FP32 parameters, with a 16,384-token input limit. 2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection. 3. Decisions at span resolution. Return…

Read vLLM Semantic Router's full model card

Vela 2.0 0.8B

Open Foundation Routing Models

Routing decisions. Safety checks. Precise text spans.

The compact hybrid member of Vela 2.0: a 756M-parameter model for multilingual routing, safety checks and span-level decisions through one interface.

Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow.

  1. Compact deployment. About 3 GB of GPU memory for FP32 parameters, with a 16,384-token input limit.
  2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection.
  3. Decisions at span resolution. Return precise text locations for trained router labels and open-label extraction, with probabilities and character offsets.
Specification Value
Reported parameters 756M
Backbone 24-layer Qwen3.5 hybrid: Gated DeltaNet and gated GQA
Input limit 16,384 tokens
Output Choice, Yes/no (Noul), Score, Span and Set; no text generation
Evaluated precision FP32 parameters; bf16 backbone autocast on GPU; FP32 heads
GPU parameter memory About 3 GB in FP32
Licence Apache-2.0

Quickstart

Install the dependencies, load the model and ask for routing and spans in one request:

pip install torch "transformers>=5.17" safetensors tokenizers numpy
pip install flash-linear-attention   # GPU: the Gated-DeltaNet kernel used in evaluation (optional, much faster)
from transformers import AutoModel

m = AutoModel.from_pretrained("vllm-sr/Vela-2.0-0.8B", trust_remote_code=True)
m = m.to("cuda")   # parameters load in FP32; on GPU the backbone runs under bf16 autocast (heads FP32)

PII_LABELS = m.vela2_engine.cal["pii_schema"]["labels"]   # the 17 trained PII types

result = m.system_one(
    state={"request": "Hi, I'm Tom Baker ([email protected]). What is the maximum daily dose of paracetamol for an adult?",
           "source": "For adults, the maximum dose of paracetamol is 4 grams in 24 hours, taken as 500 mg to 1 g every 4 to 6 hours.",
           "answer": "Adults can take up to 6 grams of paracetamol in 24 hours, in doses of 500 mg to 1 g every 4 to 6 hours."},
    questions={
        "pii": {"type": "span", "instructions": "Which spans are personal information?", "criteria": PII_LABELS,
                "over": "request"},
        "halu": {"type": "span", "instructions": "Which spans of the answer are not supported by the context?",
                 "criteria": {"unsupported": "a claim not supported by the context"}},   # over the answer by default
        "domain": {"type": "choice", "instructions": "Which subject area is this request about?", "over": "request",
                   "criteria": {"health": "medicine, clinical practice, nutrition, ageing or sexual health",
                                "math": "arithmetic, algebra, geometry, statistics or other mathematics",
                                "other": "a subject that fits none of the listed areas"}},
    })

Recorded response excerpt from this release; the full response includes usage and thresholds:

{
  "answers": {
    "domain": {
      "type": "choice",
      "choice": "health",
      "confidence": 0.929,
      "probabilities": {
        "health": 0.962,
        "math": 0.005,
        "other": 0.033
      }
    }
  },
  "spans": {
    "pii": [
      {
        "label": "PERSON",
        "start": 8,
        "end": 17,
        "text": "Tom Baker",
        "probability": 0.984
      },
      {
        "label": "EMAIL_ADDRESS",
        "start": 19,
        "end": 40,
        "text": "[email protected]",
        "probability": 0.985
      }
    ],
    "halu": [
      {
        "label": "unsupported",
        "start": 19,
        "end": 29,
        "text": "to 6 grams",
        "probability": 0.778
      }
    ]
  },
  "span_heads": {
    "pii": "router",
    "halu": "router"
  }
}

GPU with bf16 autocast is the evaluated setting. FP32 on CPU or GPU is also supported; fp16 is not. Parameter memory above excludes runtime allocations. See precision and execution.

Questions and outputs

Question Use it for Output
Choice Route a request or select one of 2–255 supplied options. Selected key and distribution
Yes/no (Noul) Check a condition against the state. P(yes)
Score Rate against 2–10 ordered levels. Expected level and distribution
Span Locate personal information, unsupported claims, entities, relations or evidence. Text spans, character offsets and probabilities
Set Select any number of supplied labels. Selected labels and per-label probabilities

Two span heads, one interface. The router span head handles trained PII, hallucination and toxic labels. The broad span head handles open-label extraction. The engine selects a head for each span question; you can also set "head": "router" or "head": "broad". See span-head selection.

Span offsets refer to Unicode code points in the selected text. Span and Set answers also provide Yes/no views in the SystemOne response. Output semantics and thresholds explain these fields.

Results

Selected results for this checkpoint. Full evaluation includes every benchmark, family comparison and scoring protocol.

Benchmark Result Metric and mode
Safety, 14 public sets 0.875 Mean AUC; trained task families
Prompt attacks, unseen families 0.940 AUC
PII, short texts 0.984 F1; shipped calibration
ACL-Verbatim evidence 23.6 Word-F1; broad span head, zero-shot
  • Evaluation: checkpoint selection, temperatures and thresholds used dev data. Safety results cover trained task families; they are not zero-shot comparisons.
  • Scoring: PII uses shipped calibration; ACL-Verbatim was held out of training. Benchmark protocols.

Architecture

This model has its own 24-layer hybrid backbone, with Gated DeltaNet and gated GQA blocks, a candidate head, and separate router/broad span-head weights. Editable SVG.

Operator and readout diagrams **Gated DeltaNet** Causal QKV convolution feeds the gated-delta update, followed by per-head normalization and output gating. [Editable SVG](assets/architecture/07-gated-deltanet.svg). **Gated GQA and SwiGLU** Gated GQA uses partial RoPE on Q/K and a sigmoid output gate; SwiGLU uses separate up and gate projections. [Editable SVG](assets/architecture/08-gated-gqa-swiglu.svg). **Candidate readout** CandidateHead combines scaled bilinear and additive MLP scores over runtime options, then applies task-specific calibration. [Editable SVG](assets/architecture/09-decoder-candidate-head.svg). **Router and broad span heads** Router and broad span heads have separate weights, with one selected for each span question. Overlapping word logits are combined before decoding. [Editable SVG](assets/architecture/10-decoder-span-heads.svg). **State prefix and question isolation** For a fixed rendered state, each question block sees the same state prefix and its own causal tokens. Additional span questions use separate sequences. [Editable SVG](assets/architecture/11-state-prefix-and-question-isolation.svg).

Choice, Yes/no, Score and Set share a candidate readout. Router and broad span heads use the same word-by-label topology with separate weights. The hybrid backbone reuses the state prefix within each rendered sequence; execution details are in USAGE.md.

Reference

Document Contents
USAGE.md Full examples and recorded responses, typed parts, head selection, calibration, SDK and HTTP serving
EVALUATION.md Complete family results, research versus shipped PII modes, benchmark protocols and evaluation disclosures
TRAINING.md Training stages, data mixtures, provenance, licences and release files
PARITY.md Export parity against the research scorer

Credit and citation

Vela 2.0 is led by KR Labs and vLLM Semantic Router.

Read the Vela 2.0 technical overview.

@misc{vela2_08b_2026,
  title        = {Vela 2.0: Towards Open Foundation Routing Models},
  author       = {{KR Labs} and {vLLM Semantic Router}},
  year         = {2026},
  note         = {Blog post. Model: vllm-sr/Vela-2.0-0.8B},
  howpublished = {\url{https://vllm-sr.ai/blog/vela-2-0-open-foundation-routing-models/}}
}

Licence

Apache-2.0 for this model's weights, code and documentation. It is derived from Decision-2.0-Eos-0.8B (Apache-2.0), itself built on Qwen3.5-0.8B (Apache-2.0); their licences and notices are passed on in LICENSE, NOTICE, ATTRIBUTIONS.md and LICENSES/. Changes against Decision-2.0-Eos-0.8B: MODIFICATIONS.md. Training data keep their own licences, some of them CC-BY-SA share-alike.

Configuration

Architecture
Vela2DecoderModel
Head dimension
256
Model type
vela2-decoder

Identity and Version

Repository
vllm-sr/Vela-2.0-0.8B
Publisher
vLLM Semantic Router
Task
Zero-shot classification
Modality
Text
Library
Not stated by the source
Parameters
755M parameters
Languages
ar, zh, cs, nl, en, fr, de, hi
Revision
efb11a33c461b82ebb72f7cc65e9aef374c52bef
First published
2026-10-03
Last updated
2026-10-07

Files and Weights

42 files, 2.0 GB in total. The weights are 2 files totalling 2.0 GB in safetensors.

Weights2 files · 2.0 GB
Configuration11 files · 189.2 KB
Tokenizer2 files · 20.0 MB
Documentation11 files · 84.6 KB
Other15 files · 2.2 MB
Repository1 file · 2.2 KB
Every file
FileTypeSizeSHA-256
broad_head.safetensorsWeights4.5 MB 687d7a435d6d
model-00001-of-00001.safetensorsWeights2.0 GB c70237ecd579
MODEL_MANIFEST.jsonConfiguration7.0 KB —
calibration.jsonConfiguration9.6 KB —
card_examples.jsonConfiguration3.7 KB —
config.jsonConfiguration4.2 KB —
configuration_vela2.pyConfiguration1.3 KB —
model.safetensors.index.jsonConfiguration28.4 KB —
modeling_vela2.pyConfiguration24.1 KB —
parity_cuda.jsonConfiguration10.8 KB —
pii_calibration_fit.jsonConfiguration23.4 KB —
vela2_inference.pyConfiguration72.3 KB —
vela2_serve.pyConfiguration4.5 KB —
ATTRIBUTIONS.mdDocumentation833 B —
EVALUATION.mdDocumentation12.0 KB —
LICENSEDocumentation11.4 KB —
LICENSES/Qwen3.5-0.8B-LICENSE.txtDocumentation11.4 KB —
MODIFICATIONS.mdDocumentation2.5 KB —
NOTICEDocumentation316 B —
PARITY.mdDocumentation4.6 KB —
PARITY_cuda.mdDocumentation4.6 KB —
README.mdDocumentation12.5 KB —
TRAINING.mdDocumentation4.7 KB —
USAGE.mdDocumentation19.7 KB —
SHA256SUMSOther2.1 KB —
assets/architecture/02-vela-2.0-0.8b-architecture.pngOther319.4 KB b18b1556e402
assets/architecture/02-vela-2.0-0.8b-architecture.svgOther9.6 KB —
assets/architecture/07-gated-deltanet.pngOther408.1 KB 6470709f5023
assets/architecture/07-gated-deltanet.svgOther14.5 KB —
assets/architecture/08-gated-gqa-swiglu.pngOther383.8 KB 09c53eb4d86e
assets/architecture/08-gated-gqa-swiglu.svgOther15.5 KB —
assets/architecture/09-decoder-candidate-head.pngOther250.0 KB 1dfe03ce756c
assets/architecture/09-decoder-candidate-head.svgOther10.2 KB —
assets/architecture/10-decoder-span-heads.pngOther344.0 KB a4f26db731ea
assets/architecture/10-decoder-span-heads.svgOther13.3 KB —
assets/architecture/11-state-prefix-and-question-isolation.pngOther253.6 KB 716526071f11
assets/architecture/11-state-prefix-and-question-isolation.svgOther8.1 KB —
assets/vela2-banner.jpgOther101.4 KB 1c01590892fe
assets/vela2-family.jpgOther102.8 KB a696a2f495eb
.gitattributesRepository2.2 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
2.0 GB
Download from vLLM Semantic Router

Released by vLLM Semantic Router through its official repository on Hugging Face. Read the license.

Built From

  • Derived from vllm-sr/Decision-2.0-Eos-0.8B
  • Trained on (disclosed) KRLabsOrg/lettucedetect-code-hallucination
  • Trained on (disclosed) KRLabsOrg/lettucedetect-prose-hallucination
  • Trained on (disclosed) KRLabsOrg/tool-output-extraction-swebench
  • Trained on (disclosed) KRLabsOrg/verbatim-spans
  • Trained on (disclosed) MultiCoNER/multiconer_v2
  • Trained on (disclosed) OpenSafetyLab/Salad-Data
  • Trained on (disclosed) ToxicityPrompts/PolyGuardMix
  • Trained on (disclosed) google-research-datasets/natural_questions
  • Trained on (disclosed) hotpotqa/hotpot_qa
  • Trained on (disclosed) knowledgator/GLINER-multi-task-synthetic-data
  • Trained on (disclosed) martinjosifoski/SynthIE
  • Trained on (disclosed) microsoft/llmail-inject-challenge
  • Trained on (disclosed) nlpaueb/finer-139
  • Trained on (disclosed) numind/NuNER
  • Trained on (disclosed) nvidia/Aegis-AI-Content-Safety-Dataset-2.0
  • Trained on (disclosed) nvidia/Nemotron-Safety-Guard-Dataset-v3
  • Trained on (disclosed) rajpurkar/squad_v2
  • Trained on (disclosed) tonytan48/Re-DocRED
  • Trained on (disclosed) urchade/pile-mistral-v0.1

Memory Requirements

PrecisionWeights in memory
As published2.0 GB
16-bit1.5 GB
8-bit0.8 GB
4-bit0.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Vela-2.0-0.8B

How much GPU memory does Vela-2.0-0.8B need?

About 1.8 GB at 16-bit and 0.5 GB at 4-bit: the weights (755M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Vela-2.0-0.8B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Vela-2.0-0.8B commercially?

Yes. Vela-2.0-0.8B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Zero-shot classification

OpenThai-SystemOne

iApp Technology

An open Thai + English "System One" decision model. It does not generate text. Given a state (any text or JSON) and typed questions, it returns calibrated probabilities over the options in one forward pass: The request/response contract mirrors TypeSafe's POST /v1/systemone so code written for the TypeSafe SDK can be pointed at this model unchanged. Typical uses: ticket routing, moderation, intent detection, RAG relevance judging, LLM-output verification, and computer-use / browser-agent action selection (which element to click, which tool to call). Gated-DeltaNet / attention, 262k context), vision encoder removed, then continued-pretrained on ~5B tokens of Thai (web, Wikipedia, parallel…

Open weights apache-2.0 753M parameters 262,144 tokens

Model · Zero-shot classification

LLM2CLIP-Openai-L-14-336

Microsoft

Weiquan Huang 1, Aoqi Wu 1, Yifan Yang 2†, Xufang Luo 2, Yuqing Yang 2, Liang Hu 1, Qi Dai 2, Xiyang Dai 2, Dongdong Chen 2, Chong Luo 2, Lili Qiu 2 In this paper, we propose LLM2CLIP, a novel approach that embraces the power of LLMs to unlock CLIP’s potential. By fine-tuning the LLM in the caption space with contrastive learning, we extract its textual capabilities into the output embeddings, significantly improving the output layer’s textual discriminability. We then design an efficient training process where the fine-tuned LLM acts as a powerful teacher for CLIP’s visual encoder. Thanks to the LLM’s presence, we can now incorporate longer and more complex captions without being…

Open weights apache-2.0 579M parameters

Models in this series are designed for efficient zeroshot classification with the Hugging Face pipeline. These models can do classification without training data and run on both GPUs and CPUs. An overview of the latest zeroshot classifiers is available in my Zeroshot Classifier Collection. The main update of this zeroshot-v2.0 series of models is that several models are trained on fully commercially-friendly data for users with strict license requirements. These models can do one universal classification task: determine whether a hypothesis is "true" or "not true" given a text (entailment vs. notentailment). This task format is based on the Natural Language Inference task (NLI). The task is…

Open weights mit 568M parameters 8,194 tokens transformers

Model · Zero-shot classification

xlm-roberta-large-xnli

Joe Davison

This model takes xlm-roberta-large and fine-tunes it on a combination of NLI data in 15 languages. It is intended to be used for zero-shot text classification, such as with the Hugging Face ZeroShotClassificationPipeline. This model is intended to be used for zero-shot text classification, especially in languages other than English. It is fine-tuned on XNLI, which is a multilingual NLI dataset. The model can therefore be used with any of the languages in the XNLI corpus: Since the base model was pre-trained trained on 100 different languages, the model has shown some effectiveness in languages beyond those listed above as well. See the full list of pre-trained languages in appendix A of the…

Open weights mit 561M parameters 514 tokens transformers

This model was fine-tuned on the MultiNLI, Fever-NLI, Adversarial-NLI (ANLI), LingNLI and WANLI datasets, which comprise 885 242 NLI hypothesis-premise pairs. This model is the best performing NLI model on the Hugging Face Hub as of 06.06.22 and can be used for zero-shot classification. It significantly outperforms all other large models on the ANLI benchmark. The foundation model is DeBERTa-v3-large from Microsoft. DeBERTa-v3 combines several recent innovations compared to classical Masked Language Models like BERT, RoBERTa etc., see the paper DeBERTa-v3-large-mnli-fever-anli-ling-wanli was trained on the MultiNLI, Fever-NLI, Adversarial-NLI (ANLI), LingNLI and WANLI datasets, which…

Open weights mit 435M parameters 512 tokens transformers

This model was trained using SentenceTransformers Cross-Encoder class. This model is based on microsoft/deberta-v3-large The model was trained on the SNLI and MultiNLI datasets. For a given sentence pair, it will output three scores corresponding to the labels: contradiction, entailment, neutral. For futher evaluation results, see SBERT.net - Pretrained Cross-Encoder. Pre-trained models can be used like this: You can use the model also directly with Transformers library (without SentenceTransformers library): This model can also be used for zero-shot-classification

Open weights apache-2.0 435M parameters 512 tokens sentence-transformers