SAVRN
Search Contact SAVRN

Open-weight model · Zero-shot classification

multilingual-zeroshot-base

by Horizon Labs Horizon-Labs/multilingual-zeroshot-base

multilingual-zeroshot-base is an open-weight model for zero-shot classification from Horizon Labs, released under Apache License 2.0. It has 308M parameters and a 8,192-token context. At 16-bit it needs about 0.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to…

Parameters308M
Context8,192
Weights3.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve multilingual-zeroshot-base (308M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.6 GB 0.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

multilingual-zeroshot-base on every accelerator the SAVRN Index prices, at every precision

Model Card

By Horizon Labs, published under apache-2.0, revision cafd191476a1.

Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to 1,024 tokens). non-commercial sets). See Training. Part of Horizon Labs' open models (collection). Labels: notentailment (0) and entailment (1). For each candidate label the model scores whether the text entails "This example is {label}." (or your hypothesistemplate). Any NLI-style use works too: pass text and textpair to a text-classification pipeline. Accuracy, single-label (multilabel=False: the label with the highest entailment score…

Read Horizon Labs's full model card

Multilingual Zero-Shot Classifier (base, 308M)

Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to 1,024 tokens).

  • Multilingual: the text can be in any of the languages below; labels and the hypothesis template stay in English.
  • Commercially clean: Apache-2.0, trained only on data that allows commercial use (no XNLI, ANLI or other non-commercial sets). See Training.
  • Small and fast: 308M parameters, ModernBERT architecture (mmBERT), ONNX included for CPU and the browser.
  • Honest numbers: all models below were run by us with the same script and templates.

Try it in the browser: Horizon-Labs/multilingual-zeroshot demo.

Part of Horizon Labs' open models (collection). Source code: github.com/horizon-ai-labs/agent-io-guards.

Quick start

from transformers import pipeline

clf = pipeline("zero-shot-classification", model="Horizon-Labs/multilingual-zeroshot-base")
clf("Mi pedido llegó roto y quiero que me devuelvan el dinero.",
    candidate_labels=["refund request", "shipping question", "product praise", "account problem"])
# {'labels': ['refund request', ...], 'scores': [...]}

# several labels can apply at once
clf("The camera is great but the battery dies by noon.", ["camera", "battery", "screen", "price"], multi_label=True)

# a task-specific template often helps
clf("¿Me pones una alarma a las siete?", ["set an alarm", "play music", "weather"], hypothesis_template="The user wants to {}.")

Labels: not_entailment (0) and entailment (1). For each candidate label the model scores whether the text entails "This example is {label}." (or your hypothesis_template). Any NLI-style use works too: pass text and text_pair to a text-classification pipeline.

Evaluation

Accuracy, single-label (multi_label=False: the label with the highest entailment score wins). English templates and labels for every language; the same template for every model (e.g. "This text is about {}." for SIB-200). No model saw these datasets' training splits, except where marked. ‡ = trained partly on data with non-commercial licenses (their -c variants are the commercially usable ones). Script: zeroshot/evaluate_zs.py.

§ = label names seen in our synthetic training data. Since v1.1 our training data includes generic label taxonomies (topics, news sections, Q&A question topics, emotions, sentiment) whose label names overlap these benchmarks' label sets; for Yahoo Answers and AG News almost exactly. No benchmark texts were used, but on these rows our models are not zero-shot with respect to the label names, so compare with care. MASSIVE, Banking77 and XNLI label sets were not used.

Multilingual

this model (308M) small (141M) bge-m3-zeroshot-v2.0-c (568M) mDeBERTa-v3-base-xnli (278M) xlm-roberta-large-xnli (560M) bge-m3-zeroshot-v2.0 (568M) ‡
MASSIVE intents (60 labels), 16 languages 0.482 0.402 0.411 0.351 0.404 0.611
SIB-200 topics (7 labels), 16 languages § 0.813 0.789 0.782 0.654 0.526 0.837

Per language, mean of MASSIVE and SIB-200:

this model (308M) small (141M) bge-m3-zeroshot-v2.0-c (568M) mDeBERTa-v3-base-xnli (278M) xlm-roberta-large-xnli (560M) bge-m3-zeroshot-v2.0 (568M) ‡
English 0.698 0.654 0.596 0.537 0.504 0.752
German 0.646 0.595 0.607 0.527 0.472 0.748
French 0.684 0.650 0.617 0.526 0.487 0.753
Spanish 0.638 0.602 0.561 0.488 0.456 0.742
Portuguese 0.666 0.617 0.585 0.491 0.446 0.713
Russian 0.642 0.623 0.597 0.497 0.453 0.724
Polish 0.688 0.642 0.640 0.533 0.486 0.756
Turkish 0.667 0.601 0.605 0.491 0.455 0.713
Arabic 0.606 0.551 0.559 0.476 0.433 0.681
Hindi 0.615 0.541 0.589 0.511 0.459 0.719
Chinese 0.674 0.625 0.627 0.514 0.481 0.760
Japanese 0.714 0.651 0.637 0.530 0.492 0.752
Korean 0.640 0.578 0.610 0.494 0.486 0.714
Vietnamese 0.613 0.563 0.605 0.478 0.490 0.735
Indonesian 0.669 0.609 0.625 0.520 0.483 0.749
Swahili 0.503 0.425 0.484 0.430 0.359 0.573

English

this model (308M) small (141M) bge-m3-zeroshot-v2.0-c (568M) mDeBERTa-v3-base-xnli (278M) xlm-roberta-large-xnli (560M) bge-m3-zeroshot-v2.0 (568M) ‡ bart-large-mnli (407M) deberta-v3-base-zeroshot-v2.0 (184M) ‡
AG News (4) § 0.826 0.828 0.726 0.670 0.591 0.886 0.684 0.884
Yahoo Answers (10) § 0.649 0.623 0.564 0.497 0.529 0.654 0.586 0.672
Banking77 (77) 0.554 0.534 0.430 0.287 0.166 0.695 0.480 0.714
Emotion (6) § 0.491 0.442 0.476 0.483 0.345 0.677 0.463 0.737
SST-2 (2) § 0.875 0.837 0.865 0.844 0.820 0.905 0.922 0.947
MASSIVE, English only 0.567 0.500 0.413 0.397 0.443 0.680 0.530 0.710
SIB-200, English only 0.828 0.809 0.779 0.676 0.564 0.824 0.760 0.745

v1.0 → v1.1

v1.1 adds data with broad, reusable label taxonomies, so it is better on common categories (topics, emotions, sentiment, aspects) — the § rows, where the label names are familiar to it. On label sets it has not seen it stays within about ±0.015 of v1.0 (slightly lower on some). To pin the previous model, load it with revision="v1.0".

v1.0 v1.1 (this version)
MASSIVE (unseen label set) 0.492 0.482
Banking77 (unseen label set) 0.568 0.554
XNLI (balanced acc.) 0.801 0.799
SIB-200 § 0.798 0.813
AG News § 0.789 0.826
Yahoo Answers § 0.503 0.649
Emotion § 0.466 0.491
SST-2 § 0.884 0.875

NLI

this model (308M) small (141M) bge-m3-zeroshot-v2.0-c (568M) mDeBERTa-v3-base-xnli (278M) xlm-roberta-large-xnli (560M) bge-m3-zeroshot-v2.0 (568M) ‡
XNLI test, 12 languages (balanced acc.) † 0.799 0.760 0.825 0.845 0.992 0.818

† XNLI is included for reference only: xlm-roberta-large-xnli and mDeBERTa-xnli were trained on XNLI data (the first scores 0.99, which suggests it saw the test sentences). Our models never saw XNLI.

Limitations

  • English-only models trained with more (partly non-commercial) classification data are better on English topic and emotion benchmarks (e.g. deberta-v3-base-zeroshot-v2.0 on Emotion and Yahoo). If you only need English, compare them on your data.
  • bge-m3-zeroshot-v2.0 (568M, trained partly on non-commercial data) scores higher on MASSIVE and SIB-200.
  • Zero-shot accuracy depends a lot on label wording and the template. Use descriptive labels ("request a refund" rather than "refund_req") and try a template that fits your task. The model links explicit wording better than implied categories. Example (small model, multi-label, a gym review not like our training domains): "The machines are always taken after 5pm and half the treadmills are broken, but the coaches really know their stuff. For 60 euros a month I expected cleaner showers." gives equipment 0.99, trainers 0.97, membership cost 0.82, but hygiene only 0.28 and crowding 0.03 (v1.0: trainers 0.56, membership cost 0.58).
  • Broad labels (e.g. "world news", "education") tend to win over specific ones. Emotions close in meaning (joy / love / surprise) are often confused.
  • With multi_label=True, scores are independent; tune the threshold on a few examples of your own.
  • Lower-resource languages (e.g. Swahili) score clearly lower than high-resource ones.
  • Much of the training data is synthetic (Qwen3.8-27B) or machine-translated.

Training

  • Backbone: jhu-clsp/mmBERT-base (MIT), sequence-pair classification, bf16, max length 1024.
  • Data (label = does the text entail the hypothesis):
  • English NLI: MultiNLI (OANC and CC-BY-SA-3.0 parts), SNLI (CC-BY-SA-4.0), WANLI (CC-BY-4.0).
  • 120k MultiNLI/WANLI pairs machine-translated by Qwen3.8-27B into 24 languages, with native and English hypotheses.
  • Synthetic zero-shot tasks by Qwen3.8-27B: FineWeb-Edu / FineWeb-2 passages (ODC-BY) labelled by topic, genre, audience, tone and purpose with near-miss wrong labels, and ~90k short texts (requests, reviews, tickets, posts, headlines) over 26 task types, 32 domains and 33 languages, each with an invented label set and hypothesis template.
  • (v1.1) Generic taxonomies by Qwen3.8-27B: 24k new FineWeb / FineWeb-2 passages labelled for topic, text type, sentiment, audience, purpose and news section; ~130k short texts written for fixed label sets (emotion, sentiment, Q&A question topic, news section, customer-message topic, urgency, formality, spam) without using the label words; ~25k reviews in 8 domains mentioning aspects (e.g. "internet", "food") without naming them.
  • Not used: XNLI, ANLI, FEVER-NLI, any benchmark above.

Configuration

Architecture
ModernBertForSequenceClassification
Context length (tokens)
8,192
Layers
22
Hidden size
768
Feed-forward size
1,152
Attention heads
12
Vocabulary size
256,000
Model type
modernbert

Identity and Version

Repository
Horizon-Labs/multilingual-zeroshot-base
Publisher
Horizon Labs
Task
Zero-shot classification
Modality
Text
Library
transformers
Parameters
308M parameters
Languages
en, de, fr, es, pt, it, nl, pl
Revision
cafd191476a143b1df810a207ec579f9b1e9684a
First published
2026-09-25
Last updated
2026-09-25

Files and Weights

21 files, 3.1 GB in total. The weights are 3 files totalling 3.1 GB in onnx, safetensors.

Weights3 files · 3.1 GB
Configuration14 files · 82.8 KB
Tokenizer2 files · 34.4 MB
Documentation1 file · 11.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.2 GB e189e4b76a0f
onnx/model.onnxWeights1.2 GB 8fc5bb688ada
onnx/model_quantized.onnxWeights641.1 MB e6fc10790214
code/train/export_onnx.pyConfiguration2.1 KB —
code/train/train.pyConfiguration6.6 KB —
code/zeroshot/build_evals.pyConfiguration4.6 KB —
code/zeroshot/build_zs_data.pyConfiguration5.5 KB —
code/zeroshot/evaluate_zs.pyConfiguration4.7 KB —
code/zeroshot/gen_zeroshot.pyConfiguration6.7 KB —
code/zeroshot/translate_nli.pyConfiguration3.8 KB —
config.jsonConfiguration2.1 KB —
eval/baselines.jsonConfiguration39.1 KB —
eval/this_model.jsonConfiguration4.5 KB —
onnx/quantization_check.jsonConfiguration1.1 KB —
special_tokens_map.jsonConfiguration636 B —
training/train_log.jsonConfiguration1.2 KB —
training/val_metrics.jsonConfiguration390 B —
README.mdDocumentation11.0 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer34.4 MB aebee76d0312
tokenizer_config.jsonTokenizer46.4 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
3.1 GB
Download from Horizon Labs

Released by Horizon Labs through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published3.1 GB
16-bit0.6 GB
8-bit0.3 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About multilingual-zeroshot-base

How much GPU memory does multilingual-zeroshot-base need?

About 0.7 GB at 16-bit and 0.2 GB at 4-bit: the weights (308M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run multilingual-zeroshot-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use multilingual-zeroshot-base commercially?

Yes. multilingual-zeroshot-base is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is multilingual-zeroshot-base's context length?

8,192 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Zero-shot classification

gliner-guard-omni

HiveTraceLab

One encoder model that replaces your entire guardrail stack: safety classification, PII detection, adversarial attack detection, intent and tone analysis — all in a single forward classification, NER and more · no LLM required Install dependencies Classify Harmful messages and Detect PII via single forward pass GLiNER Guard Omni fine-tunes fastino/gliner2-multi-v1 on our guardrail taxonomy while preserving its multilingual zero-shot generalization. You get GLiNER Guard's safety understanding on top of the base model's ability to handle labels and domains beyond the training set — so you can define custom policies with nothing but natural language descriptions. For specific usecases you can…

Open weights apache-2.0 307M parameters gliner2

This multilingual model can perform natural language inference (NLI) on 100 languages and is therefore also suitable for multilingual zero-shot classification. The underlying mDeBERTa-v3-base model was pre-trained by Microsoft on the CC100 multilingual dataset with 100 languages. The model was then fine-tuned on the XNLI dataset and on the multilingual-NLI-26lang-2mil7 dataset. Both datasets contain more than 2.7 million hypothesis-premise pairs in 27 languages spoken by more than 4 billion people. As of December 2021, mDeBERTa-v3-base is the best performing multilingual base-sized transformer model introduced by Microsoft in this paper. This model was trained on the…

Open weights mit 279M parameters 512 tokens transformers

This multilingual model can perform natural language inference (NLI) on 100 languages and is therefore also suitable for multilingual zero-shot classification. The underlying model was pre-trained by Microsoft on the CC100 multilingual dataset. It was then fine-tuned on the XNLI dataset, which contains hypothesis-premise pairs from 15 languages, as well as the English MNLI dataset. As of December 2021, mDeBERTa-base is the best performing multilingual base-sized transformer model, introduced by Microsoft in this paper. If you are looking for a smaller, faster (but less performant) model, you can try multilingual-MiniLMv2-L6-mnli-xnli. This model was trained on the XNLI development dataset…

Open weights mit 279M parameters 512 tokens transformers

Model · Zero-shot classification

scandi-nli-large

Alexandra Institute

This model is a fine-tuned version of NbAiLab/nb-bert-large for Natural Language Inference in Danish, Norwegian Bokmål and Swedish. We have released three models for Scandinavian NLI, of different sizes: - alexandrainst/scandi-nli-large (this) A demo of the large-v2 model can be found in this Hugging Face Space - check it out! The performance and model size of each of them can be found in the Performance section below. You can use this model in your scripts as follows: We assess the models both on their aggregate Scandinavian performance, as well as their language-specific Danish, Swedish and Norwegian Bokmål performance. In all cases, we report Matthew's Correlation Coefficient (MCC)…

Open weights apache-2.0 355M parameters 512 tokens transformers

Model · Zero-shot classification

nitzotz

Netanel Elyasi

What it is. Nitzotz reads a Hebrew message and answers questions you type about it: pick one of several options, give a score on a scale, or say yes or no to a claim. For every answer it gives a probability you can trust, so you know when it is sure and when it is guessing. It does not write text, so it cannot make things up. It runs on a normal laptop, with no internet connection and no cost per question. What it is for. Deciding what to do with incoming messages: is this a scam, what kind of message is it, which department should get it, how urgent is it. It is not a chatbot and it is not built for long documents (see Limitations). In numbers. On 298 Hebrew messages it says correctly…

Open weights apache-2.0 384M parameters laya

Model · Zero-shot classification

ModernBERT-large-nli

Tasksource

This model is ModernBERT multi-task fine-tuned on tasksource NLI tasks, including MNLI, ANLI, SICK, WANLI, doc-nli, LingNLI, FOLIO, FOL-NLI, LogicNLI, Label-NLI and all datasets in the below table). This is the equivalent of an "instruct" version. The model was trained for 200k steps on an Nvidia A30 GPU. It is very good at reasoning tasks (better than llama 3.1 8B Instruct on ANLI and FOLIO), long context reasoning, sentiment analysis and zero-shot classification with new labels. The following table shows model test accuracy. These are the scores for the same single transformer with different classification heads on top. Further gains can be obtained by fine-tuning on a single-task, e.g.…

Open weights apache-2.0 396M parameters 2,048 tokens transformers