Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to 1,024 tokens). non-commercial sets). See Training. Part of Horizon Labs' open models (collection). Labels: notentailment (0) and entailment (1). For each candidate label the model scores whether the text entails "This example is {label}." (or your hypothesistemplate). Any NLI-style use works too: pass text and textpair to a text-classification pipeline. Accuracy, single-label (multilabel=False: the label with the highest entailment score…
Open-weight model · Zero-shot classification
Vela-2.0-0.3B
by vLLM Semantic Router vllm-sr/Vela-2.0-0.3B
Vela-2.0-0.3B is an open-weight model for zero-shot classification from vLLM Semantic Router, released under Apache License 2.0. It has 309M parameters. At 16-bit it needs about 0.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 25 downloads a month.
Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. Vela 2.0 brings routing questions and span-level decisions into one model interface.
Runs On
What it takes to serve Vela-2.0-0.3B (309M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.6 GB | 0.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.3 GB | 0.4 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.2 GB | 0.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Vela-2.0-0.3B on every accelerator the SAVRN Index prices, at every precision
Model Card
By vLLM Semantic Router, published under apache-2.0, revision 0c2c47eeeea7.
Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. Vela 2.0 brings routing questions and span-level decisions into one model interface. Supply your options, labels and rubrics at request time; ask about the request, its context and the answer together. The 0.3B is the family's compact multilingual encoder, with an 8,192-token input budget and Torch or ONNX inference on CPU or GPU. 1. Route with your own criteria. Choose a route, check a policy condition, assign a score or select multiple labels. 2. Locate the text behind a signal. Return spans of personal information and unsupported claims with character offsets and probabilities. 3. Ask across the whole…
Read vLLM Semantic Router's full model card
Docs | Blog | GitHub | Vela 2.0 collection
Vela 2.0 0.3B
Open Foundation Routing Models
Routing decisions. Safety checks. Precise text spans.
Vela 2.0 brings routing questions and span-level decisions into one model interface. Supply your options, labels and rubrics at request time; ask about the request, its context and the answer together.
The 0.3B is the family's compact multilingual encoder, with an 8,192-token input budget and Torch or ONNX inference on CPU or GPU.
- Route with your own criteria. Choose a route, check a policy condition, assign a score or select multiple labels.
- Locate the text behind a signal. Return spans of personal information and unsupported claims with character offsets and probabilities.
- Ask across the whole interaction. Use typed request, context and answer parts in one request, through Python or a SystemOne-compatible HTTP API.
| Model specification | Vela 2.0 0.3B |
|---|---|
| Backbone | 22-layer bidirectional ModernBERT encoder |
| Hidden width | 768 |
| Input budget | 8,192 tokens, including schema and state |
| Questions | Choice, Yes/no (Noul), Score, Span, Set |
| Inference | Torch fp32; ONNX fp32 or fp16 encoder with fp32 heads |
| Deployment | CPU or GPU; Python API and HTTP server |
Quickstart
Ask for personal information, unsupported answer text and a route in one request:
pip install torch "transformers>=4.57" safetensors tokenizers numpy
from transformers import AutoModel
m = AutoModel.from_pretrained("vllm-sr/Vela-2.0-0.3B", trust_remote_code=True)
# Runs on CPU by default; optionally use m = m.to("cuda"). Keep Torch loading in fp32.
PII_LABELS = m.vela2_engine.cal["pii_schema"]["labels"]
result = m.system_one(
state={"request": "Hi, I'm Tom Baker ([email protected]). What is the maximum daily dose of paracetamol for an adult?",
"source": "For adults, the maximum dose of paracetamol is 4 grams in 24 hours, taken as 500 mg to 1 g every 4 to 6 hours.",
"answer": "Adults can take up to 6 grams of paracetamol in 24 hours, in doses of 500 mg to 1 g every 4 to 6 hours."},
questions={
"pii": {"type": "span", "instructions": "Which spans are personal information?", "criteria": PII_LABELS,
"over": "request"},
"halu": {"type": "span", "instructions": "Which spans of the answer are not supported by the context?",
"criteria": {"unsupported": "a claim not supported by the context"}}, # over the answer by default
"domain": {"type": "choice", "instructions": "Which subject area is this request about?", "over": "request",
"criteria": {"health": "medicine, clinical practice, nutrition, ageing or sexual health",
"math": "arithmetic, algebra, geometry, statistics or other mathematics",
"other": "a subject that fits none of the listed areas"}},
})
Selected fields from the recorded Torch CPU response, with probabilities rounded to three decimals:
{
"answers": {
"pii": {"type": "noul", "noul": 1.0},
"halu": {"type": "noul", "noul": 0.984},
"domain": {"type": "choice", "choice": "health", "confidence": 0.51, "probabilities": {"health": 0.727, "math": 0.217, "other": 0.056}}
},
"spans": {
"pii": [{"label": "PERSON", "start": 8, "end": 17, "text": "Tom Baker", "probability": 1.0}, {"label": "EMAIL_ADDRESS", "start": 19, "end": 40, "text": "[email protected]", "probability": 0.983}],
"halu": [{"label": "unsupported", "start": 16, "end": 29, "text": "up to 6 grams", "probability": 0.895}]
}
}
Offsets are Unicode code-point offsets into the targeted field. Full response, HTTP, SDK and ONNX examples.
Questions and outputs
| Question | Use it for | Output |
|---|---|---|
| Choice | Route a request or pick one of 2–255 supplied options. | Selected key and distribution |
| Yes/no (Noul) | Check a condition against the state. | P(yes) |
| Score | Rate against 2–10 ordered levels. | Expected level and distribution |
| Span | Locate personal information or claims unsupported by the source. | Labels, text, offsets and probabilities |
| Set | Select any number of supplied labels. | Selected labels and per-label probabilities |
All five types share one interface. Span and Set preserve SystemOne compatibility through additional response fields and Noul views. Request and response reference.
Results
| Task | Vela 2.0 0.3B | Reference |
|---|---|---|
| Router safety, macro AUC over 14 sets | 0.871 | GLiNER2.5-Decide 0.704 |
| Prompt attacks, unseen families, AUC | 0.882 | Vela 1.0 Guard 0.792 |
| Short PII, exact micro-F1 | 0.995 | Vela 1.0 PII 0.976 |
- Selection: checkpoint selection used dev splits; final release selection also considered test results.
- Safety: these task families were part of Vela's training, unlike Decide's. This comparison measures router safety performance, rather than zero-shot generalisation.
Selected results above use identical rows for each task comparison. Full results and evaluation protocol include every benchmark, confidence intervals and the separate shipped, test-blind and research PII calibration paths.
Architecture
A 22-layer bidirectional ModernBERT encoder with schema-conditioned decision and word–label span readouts. The first attention pre-norm is Identity; this diagram shows the SDPA path. Editable SVG.
Operator and readout diagrams
**Attention and GEGLU** The SDPA attention path uses 12 heads of width 64 and applies ×4 YaRN RoPE to Q/K. GEGLU has two 1,152-wide branches. [Editable SVG](assets/architecture/05-encoder-attention-geglu.svg). **Decision and span readouts** The decision and span readouts use 256-dimensional projections and shared readout LayerNorm parameters. Learned τ and inference calibration temperatures T are distinct. [Editable SVG](assets/architecture/06-encoder-readouts.svg).The backbone reads schema and typed state bidirectionally. Decision heads combine option cosine scores with an option MLP; the span head compares word and label states. Model, calibration and training details.
Reference
- Usage: Python, HTTP, SDK, Span/Set, typed parts, shortcuts and ONNX.
- Evaluation: full tables, scoring paths, references and selection disclosures.
- Training and model details: data, training recipe, input schema, calibration and repository files.
- Export parity: Torch/research-scorer and ONNX comparisons on the measured rows.
Credit and citation
Vela 2.0 is led by KR Labs and vLLM Semantic Router.
Read the Vela 2.0 technical overview.
@misc{vela2_unified_2026,
title = {Vela 2.0: Towards Open Foundation Routing Models},
author = {{KR Labs} and {vLLM Semantic Router}},
year = {2026},
note = {Blog post. Model: vllm-sr/Vela-2.0-0.3B},
howpublished = {\url{https://vllm-sr.ai/blog/vela-2-0-open-foundation-routing-models/}}
}
Licence
Apache-2.0 for this model's weights, code and documentation. It is derived from Decision-1.0-Kai-0.6B (Apache-2.0) and Vela-1.0-Encoder-307M (MIT, from mmBERT); their licences and notices are passed on in LICENSE, NOTICE, NOTICE_SOURCES.json and LICENSES/. The tokenizer carries the Gemma Terms of Use: DISTRIBUTION_TERMS.md · License scope. Changes against Kai and Vela, including the eight added tokenizer entries: MODIFICATIONS.md. Training data keep their own licences, some of them CC-BY-SA share-alike (see NOTICE_SOURCES.json).
Configuration
- Architecture
- Vela2Model
- Model type
- vela2-unified
Identity and Version
- Repository
- vllm-sr/Vela-2.0-0.3B
- Publisher
- vLLM Semantic Router
- Task
- Zero-shot classification
- Modality
- Text
- Library
- Not stated by the source
- Parameters
- 309M parameters
- Languages
- ar, zh, cs, nl, en, fr, de, hi
- Revision
- 0c2c47eeeea7b234242bd6170a71408c8804134a
- First published
- 2026-09-29
- Last updated
- 2026-10-07
Files and Weights
40 files, 3.1 GB in total. The weights are 3 files totalling 3.1 GB in onnx, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 1.2 GB | bd37dd0db177 |
| onnx/model.onnx | Weights | 1.2 GB | 5096731c509b |
| onnx/model_fp16.onnx | Weights | 623.1 MB | aaea0639f13b |
| NOTICE_SOURCES.json | Configuration | 9.3 KB | — |
| calibration.json | Configuration | 7.1 KB | — |
| config.json | Configuration | 2.8 KB | — |
| configuration_vela2.py | Configuration | 1.6 KB | — |
| modeling_vela2.py | Configuration | 9.6 KB | — |
| pii_calibration_fit.json | Configuration | 46.2 KB | — |
| special_tokens_map.json | Configuration | 2.2 KB | — |
| vela2_inference.py | Configuration | 60.9 KB | — |
| vela2_serve.py | Configuration | 4.3 KB | — |
| DISTRIBUTION_TERMS.md | Documentation | 1.3 KB | — |
| EVALUATION.md | Documentation | 8.4 KB | — |
| LICENSE | Documentation | 11.4 KB | — |
| LICENSES/gemma/Notice | Documentation | 1.1 KB | — |
| LICENSING_STATUS.md | Documentation | 1.5 KB | — |
| MODIFICATIONS.md | Documentation | 1.5 KB | — |
| NOTICE | Documentation | 4.4 KB | — |
| PARITY.md | Documentation | 11.4 KB | — |
| README.md | Documentation | 10.4 KB | — |
| TRAINING.md | Documentation | 6.1 KB | — |
| USAGE.md | Documentation | 26.1 KB | — |
| LICENSES/Transformers-Apache-2.0.txt | Other | 11.4 KB | — |
| LICENSES/Upstream-MIT.txt | Other | 1.2 KB | — |
| LICENSES/gemma/GEMMA_PROHIBITED_USE_POLICY.html | Other | 4.6 KB | — |
| LICENSES/gemma/GEMMA_TERMS.html | Other | 13.5 KB | — |
| assets/architecture/01-vela-2.0-0.3b-architecture.png | Other | 293.6 KB | 942a7b6c89e4 |
| assets/architecture/01-vela-2.0-0.3b-architecture.svg | Other | 9.5 KB | — |
| assets/architecture/05-encoder-attention-geglu.png | Other | 341.7 KB | 05ebe1941daf |
| assets/architecture/05-encoder-attention-geglu.svg | Other | 13.2 KB | — |
| assets/architecture/06-encoder-readouts.png | Other | 372.4 KB | eeb0896d3bd3 |
| assets/architecture/06-encoder-readouts.svg | Other | 15.8 KB | — |
| assets/vela2-banner.jpg | Other | 101.4 KB | 1c01590892fe |
| assets/vela2-family.jpg | Other | 102.8 KB | a696a2f495eb |
| figures/fig_delta_vs_vela.png | Other | 109.2 KB | 0987ca4eba16 |
| .gitattributes | Repository | 306 B | — |
| LICENSES/gemma/TOKENIZER_TERMS.md | Tokenizer | 1.3 KB | — |
| tokenizer.json | Tokenizer | 34.4 MB | e4b670c5ab72 |
| tokenizer_config.json | Tokenizer | 48.0 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 3.1 GB
Released by vLLM Semantic Router through its official repository on Hugging Face. Read the license.
Built From
- Derived from vllm-sr/Decision-1.0-Kai-0.6B
- Trained on (disclosed) KRLabsOrg/lettucedetect-code-hallucination
- Trained on (disclosed) KRLabsOrg/lettucedetect-prose-hallucination
- Trained on (disclosed) OpenSafetyLab/Salad-Data
- Trained on (disclosed) ToxicityPrompts/PolyGuardMix
- Trained on (disclosed) microsoft/llmail-inject-challenge
- Trained on (disclosed) nvidia/Aegis-AI-Content-Safety-Dataset-2.0
- Trained on (disclosed) nvidia/Nemotron-Safety-Guard-Dataset-v3
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 3.1 GB |
| 16-bit | 0.6 GB |
| 8-bit | 0.3 GB |
| 4-bit | 0.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Vela-2.0-0.3B
How much GPU memory does Vela-2.0-0.3B need?
About 0.7 GB at 16-bit and 0.2 GB at 4-bit: the weights (309M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Vela-2.0-0.3B on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Vela-2.0-0.3B commercially?
Yes. Vela-2.0-0.3B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Similar Models
One encoder model that replaces your entire guardrail stack: safety classification, PII detection, adversarial attack detection, intent and tone analysis — all in a single forward classification, NER and more · no LLM required Install dependencies Classify Harmful messages and Detect PII via single forward pass GLiNER Guard Omni fine-tunes fastino/gliner2-multi-v1 on our guardrail taxonomy while preserving its multilingual zero-shot generalization. You get GLiNER Guard's safety understanding on top of the base model's ability to handle labels and domains beyond the training set — so you can define custom policies with nothing but natural language descriptions. For specific usecases you can…
Siddhanto (সিদ্ধান্ত, "decision") is a fine-tune of convaiinnovations/laya-multilingual (mmBERT-base, 322M params) for open-domain typed decisions over Bangla, English and Banglish (romanized Bangla / code-switched) conversational text. It is not a fixed-label classifier. It implements Laya's typed choice / noul / score interface: at request time you supply the question and its options (any schema, any domain), and the model returns a calibrated probability distribution over exactly those options — no retraining or head-swapping needed to add a new label set. Fine-tuned on ~268k typed-decision examples drawn from: - An internal 56-intent OTA (travel-agency) synthetic corpus…
VIYA reads Vietnamese text and returns typed decisions with calibrated confidence in a single forward pass. It does not generate text: you describe the question and the options in plain language, and VIYA scores every option. Options are free text, so you can add or rename them without retraining. - Vietnamese fact-checking. 87.6 macro-F1 on ViFactCheck with gold evidence (human: 84.9). When it has to find the evidence itself in the full article, it scores 78.7, ahead of Gemini 1.5 Flash, XLM-R large and Mistral 7B. - Full fact-checking pipeline. On ViWikiFC, VIYA finds the right evidence sentence and reaches the right verdict 75.7% of the time; the best published pipeline reaches 67.0%.…
Model · Zero-shot classification
mDeBERTa-v3-base-xnli-multilingual-nli-2mil7
This multilingual model can perform natural language inference (NLI) on 100 languages and is therefore also suitable for multilingual zero-shot classification. The underlying mDeBERTa-v3-base model was pre-trained by Microsoft on the CC100 multilingual dataset with 100 languages. The model was then fine-tuned on the XNLI dataset and on the multilingual-NLI-26lang-2mil7 dataset. Both datasets contain more than 2.7 million hypothesis-premise pairs in 27 languages spoken by more than 4 billion people. As of December 2021, mDeBERTa-v3-base is the best performing multilingual base-sized transformer model introduced by Microsoft in this paper. This model was trained on the…
This multilingual model can perform natural language inference (NLI) on 100 languages and is therefore also suitable for multilingual zero-shot classification. The underlying model was pre-trained by Microsoft on the CC100 multilingual dataset. It was then fine-tuned on the XNLI dataset, which contains hypothesis-premise pairs from 15 languages, as well as the English MNLI dataset. As of December 2021, mDeBERTa-base is the best performing multilingual base-sized transformer model, introduced by Microsoft in this paper. If you are looking for a smaller, faster (but less performant) model, you can try multilingual-MiniLMv2-L6-mnli-xnli. This model was trained on the XNLI development dataset…