SAVRN
Search Contact SAVRN

Open-weight model · Text classification

content-safety-guard-base

by Horizon Labs Horizon-Labs/content-safety-guard-base

content-safety-guard-base is an open-weight model for text classification from Horizon Labs, released under Apache License 2.0. It has 308M parameters and a 8,192-token context. At 16-bit it needs about 0.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A small multilingual classifier that flags unsafe user prompts and unsafe model responses for LLM applications, with harm categories.

Parameters308M
Context8,192
Weights3.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve content-safety-guard-base (308M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.6 GB 0.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

content-safety-guard-base on every accelerator the SAVRN Index prices, at every precision

Model Card

By Horizon Labs, published under apache-2.0, revision 90ff0ffe86ca.

A small multilingual classifier that flags unsafe user prompts and unsafe model responses for LLM applications, with harm categories. It is a fast encoder (ModernBERT architecture, mmBERT backbone) that you can run on CPU or in the browser in front of, or behind, any LLM. real-world prompts; evaluated on 17 (PolyGuard), 14 (textdetox) and 8 (Aya) languages. Part of the Horizon Labs guard family: prompt-injection-guard, Decision rule: flag when unsafe >= 0.5 (raise the threshold if you see too many false alarms, lower it for higher recall). Category scores are only meaningful for flagged texts and are small by design; use the per-category thresholds in thresholds.json (chosen on a validation…

Read Horizon Labs's full model card

Content Safety Guard (base, 308M)

A small multilingual classifier that flags unsafe user prompts and unsafe model responses for LLM applications, with harm categories. It is a fast encoder (ModernBERT architecture, mmBERT backbone) that you can run on CPU or in the browser in front of, or behind, any LLM.

  • Prompts and responses: pass a prompt alone, or a (prompt, response) pair to judge the response in context.
  • Output: an overall unsafe score plus 15 category scores (independent sigmoids): hate, harassment, violence, weapons, sexual, sexual_minors, self_harm, criminal_planning, drugs, privacy, profanity, fraud_manipulation, misinformation, unauthorized_advice, other.
  • Multilingual: trained on 12 languages of human-adapted data, machine translations into 20 languages and multilingual real-world prompts; evaluated on 17 (PolyGuard), 14 (textdetox) and 8 (Aya) languages.
  • Commercially usable: Apache-2.0, trained only on data that allows commercial use (see Training data).
  • ONNX: onnx/model.onnx (fp32) and onnx/model_quantized.onnx (int8 embeddings, 641 MB; its unsafe decision agrees with fp32 on 100.0% of a sample of benchmark texts (30 per benchmark), mean |score diff| 0.0015).

Part of the Horizon Labs guard family: prompt-injection-guard, pii-redactor, hallucination-guard.

Usage

from transformers import pipeline

clf = pipeline("text-classification", model="Horizon-Labs/content-safety-guard-base", top_k=None)

# a user prompt
scores = {d["label"]: d["score"] for d in clf("How can I make a fake ID that passes a bouncer's check?")[0]}
print(scores["unsafe"])                          # ~1.0

# a model response, judged together with its prompt
r = clf({"text": "How do I get rid of a wasp nest?", "text_pair": "Spray it at dusk with a wasp foam, then remove it."})[0]

Decision rule: flag when unsafe >= 0.5 (raise the threshold if you see too many false alarms, lower it for higher recall). Category scores are only meaningful for flagged texts and are small by design; use the per-category thresholds in thresholds.json (chosen on a validation split, see below):

import json
from huggingface_hub import hf_hub_download
thr = json.load(open(hf_hub_download("Horizon-Labs/content-safety-guard-base", "thresholds.json")))
def moderate(text, pair=None):
    s = {d["label"]: d["score"] for d in clf({"text": text, "text_pair": pair} if pair else text)[0]}
    flagged = s["unsafe"] >= thr["unsafe"]
    cats = [c for c, t in thr["categories"].items() if flagged and s[c] >= t]
    return flagged, s["unsafe"], cats

transformers.js (browser / Node):

import { pipeline } from "@huggingface/transformers";
const clf = await pipeline("text-classification", "Horizon-Labs/content-safety-guard-base", { dtype: "q8" });
const out = await clf("How do I make a pipe bomb?", { top_k: null });

Evaluation

F1 of the unsafe class at threshold 0.5 (for the recall rows: share of unsafe prompts flagged). None of these benchmarks were used for training (best value in bold; false-alarm rates are not bolded); Qwen3Guard-Gen is scored from its next-token probabilities after "Safety:" (strict: Unsafe + Controversial count as unsafe; loose: only Unsafe). Toxicity classifiers are included because they are often used for this job; they target a different, narrower task. Encoders other than ours get the response alone for response items.

this model (308M) Qwen3Guard-Gen-0.6B strict Qwen3Guard-Gen-0.6B loose Vela-1.0-307M-Shield ¶ granite-guardian-hap-125m unbiased-toxic-roberta Qwen3Guard-Gen-8B strict (teacher, 8B)
PolyGuard prompts, 17 languages 0.782 0.821 0.791 0.829 0.003 0.002 0.846
PolyGuard responses, 17 languages 0.736 0.748 0.767 0.645 0.000 0.000 0.801
BeaverTails responses (unseen prompts) † 0.834 0.869 0.858 0.779 0.164 0.177 0.870
ToxicChat (real user prompts) 0.768 0.588 0.760 0.672 0.264 0.254 0.638
OpenAI moderation set 0.761 0.660 0.780 0.721 0.672 0.663 0.685
XSTest (over-blocking test) 0.785 0.853 0.852 0.838 0.315 0.227 0.908
textdetox toxicity, 14 languages 0.673 0.722 0.461 0.720 0.146 0.152 0.775
Aya red-teaming, 8 languages (recall) 0.807 0.859 0.605 0.812 0.023 0.025 0.946
SimpleSafetyTests (recall) 0.930 0.980 0.920 0.960 0.240 0.260 0.990

Ranking quality (ROC AUC, threshold-free):

this model (308M) Qwen3Guard-Gen-0.6B strict Qwen3Guard-Gen-0.6B loose Vela-1.0-307M-Shield ¶ granite-guardian-hap-125m unbiased-toxic-roberta Qwen3Guard-Gen-8B strict (teacher, 8B)
PolyGuard prompts, 17 languages 0.884 0.910 0.902 0.908 0.569 0.542 0.931
PolyGuard responses, 17 languages 0.956 0.945 0.941 0.905 0.566 0.552 0.959
BeaverTails responses (unseen prompts) † 0.911 0.926 0.926 0.836 0.662 0.593 0.929
ToxicChat (real user prompts) 0.983 0.985 0.979 0.972 0.841 0.844 0.988
OpenAI moderation set 0.919 0.924 0.921 0.913 0.877 0.875 0.941
XSTest (over-blocking test) 0.913 0.958 0.947 0.926 0.688 0.678 0.987
textdetox toxicity, 14 languages 0.836 0.769 0.757 0.804 0.585 0.529 0.846

Over-blocking (share of safe items flagged; lower is better):

this model (308M) Qwen3Guard-Gen-0.6B strict Qwen3Guard-Gen-0.6B loose Vela-1.0-307M-Shield ¶ granite-guardian-hap-125m unbiased-toxic-roberta Qwen3Guard-Gen-8B strict (teacher, 8B)
XSTest: safe prompts flagged 0.284 0.228 0.052 0.104 0.036 0.012 0.128
PolyGuard prompts: safe prompts flagged 0.092 0.115 0.050 0.065 0.002 0.000 0.081
OpenAI moderation set: safe texts flagged 0.165 0.437 0.125 0.287 0.068 0.092 0.397
ToxicChat: safe prompts flagged 0.027 0.102 0.012 0.055 0.004 0.008 0.083

¶ Vela-Shield was trained on PolyGuardMix, the training split of the PolyGuard benchmark family, so its PolyGuard rows are in-distribution. † BeaverTails and our main training set both take prompts from Anthropic's HH red-team data; the row uses only the 1894 test items whose prompt does not occur in our training data. ToxicChat shares 29 of 5083 prompts with our training data.

PolyGuard prompts per language (F1):

Language this model Qwen3Guard-Gen-0.6B strict Vela-Shield ¶
Arabic 0.759 0.817 0.834
Chinese 0.819 0.848 0.830
Czech 0.796 0.810 0.857
Dutch 0.771 0.789 0.813
English 0.796 0.875 0.871
French 0.796 0.843 0.847
German 0.774 0.818 0.813
Hindi 0.746 0.778 0.803
Italian 0.799 0.825 0.839
Japanese 0.790 0.801 0.831
Korean 0.759 0.787 0.795
Polish 0.782 0.807 0.808
Portuguese 0.768 0.845 0.849
Russian 0.766 0.845 0.813
Spanish 0.817 0.843 0.846
Swedish 0.786 0.802 0.837
Thai 0.769 0.823 0.799

Categories

Category labels come from the Aegis 2.0 taxonomy of Nemotron-Safety-Guard-Dataset-v3 (merged into 15 groups). Quality on the unsafe items of its test split, with thresholds chosen on its validation split:

Category threshold F1 (test) F1 at 0.5 AUC test positives
criminal_planning 0.25 0.796 0.761 0.882 2060
violence 0.20 0.577 0.510 0.844 814
drugs 0.35 0.607 0.571 0.897 696
hate 0.25 0.729 0.699 0.930 592
harassment 0.20 0.522 0.453 0.832 562
privacy 0.15 0.702 0.606 0.908 483
other 0.20 0.510 0.148 0.882 479
weapons 0.20 0.612 0.507 0.879 368
sexual 0.50 0.719 0.719 0.933 333
profanity 0.20 0.561 0.501 0.884 326
self_harm 0.15 0.599 0.505 0.920 308
fraud_manipulation 0.15 0.344 0.081 0.848 299
unauthorized_advice 0.10 0.280 0.109 0.873 206
sexual_minors 0.25 0.368 0.300 0.902 165
misinformation 0.15 0.406 0.244 0.863 141

Training

  • Data (all permit commercial use):
  • nvidia/Nemotron-Safety-Guard-Dataset-v3 (CC-BY-4.0; Aegis 2.0 prompts and responses, culturally adapted into 12 languages): 574k prompt and response items, with its human category labels. Rows derived from a Kaggle dataset (REDACTED) were dropped.
  • Real and red-team prompts without labels, scored by the teacher: first user turns and replies from WildChat-1M (ODC-BY), prompts from oasst2 and the Aya dataset (Apache-2.0), Salad-Data (Apache-2.0; ToxicChat-derived rows dropped), JailbreakBench behaviours (MIT), and jailbreak, role-play and over-refusal prompts from the training data of our prompt-injection guard (270k + 42k items).
  • Civil Comments (CC0): 60k comments with toxicity >= 0.5 and 60k with toxicity 0; target = the share of annotators who rated it toxic, categories from the insult / threat / obscene / identity-attack / sexual ratings. This improved the OpenAI moderation set (F1 +0.035) but not the textdetox recall.
  • (v1.1) Machine translations by Qwen3.8-27B (Apache-2.0) into 20 languages (Portuguese, Russian, Ukrainian, Polish, Czech, Swedish, Turkish, Hebrew, Serbian, Tagalog, Amharic, Tatar, German, Spanish, French, Arabic, Hindi, Chinese, Japanese, Italian): 47.6k Civil Comments (a different shard from the one above; they keep their annotator toxicity and categories) and 46.6k of the teacher-labelled prompts above (harmful, benign-but-edgy and benign; re-scored by the teacher in the target language). 1.8% of the translations were dropped (unparseable, refusals, implausible length).
  • Items that match any benchmark text were removed.
  • Teacher: Qwen3Guard-Gen-8B (Apache-2.0). The unsafe target of every item except Civil Comments is the teacher's probability P(Unsafe) + 0.5 · P(Controversial); categories use the human labels. So the model follows Qwen3Guard's safety policy, not Aegis's stricter human labels (which also mark sensitive but harmless requests).
  • Model: jhu-clsp/mmBERT-base with a 16-way sigmoid head (unsafe + 15 categories), 2 epochs, max length 1024 tokens.
  • Code: code/ in this repository.

Versions

v1.1 adds machine-translated toxic comments and red-team prompts in 20 languages: multilingual toxicity (textdetox) and native-speaker red-teaming (Aya) improve clearly; the other rows move by 0.01 or less. To pin an earlier model, load it with revision="v1.0".

v1.0 v1.1 (this version)
PolyGuard prompts, 17 languages 0.779 0.782
PolyGuard responses, 17 languages 0.732 0.736
BeaverTails responses (unseen prompts) † 0.838 0.834
ToxicChat (real user prompts) 0.771 0.768
OpenAI moderation set 0.768 0.761
XSTest (over-blocking test) 0.786 0.785
textdetox toxicity, 14 languages 0.597 0.673
Aya red-teaming, 8 languages (recall) 0.766 0.807
SimpleSafetyTests (recall) 0.930 0.930
XSTest: safe prompts flagged (lower is better) 0.288 0.284

Limitations

  • It trails Qwen3Guard-Gen-0.6B (a generative 0.6B model that reads a long policy prompt per item) in both of its modes on PolyGuard prompts, PolyGuard responses, BeaverTails responses, XSTest (0.012-0.039 F1 behind its strict mode on PolyGuard responses / prompts); it is ahead of the strict mode on ToxicChat, OpenAI moderation set. Its advantages are speed, size, CPU/browser use and multilingual coverage in one small encoder.
  • Over-blocking: it flags about 28% of XSTest's safe-but-scary prompts ("how do I kill a Python process").
  • Classic toxicity (insults, profanity without other harm) is only partly covered (textdetox recall 0.57).
  • Recall on harmful prompts written by native speakers in lower-resource languages is lower (Aya red-teaming 0.81).
  • Categories are weak for fraud_manipulation, unauthorized_advice, misinformation and sexual_minors (F1 below 0.45). Do not rely on it alone for child-safety or legal compliance; use it as one signal with human review.
  • Safety policies differ between applications; tune the threshold on your own traffic.

Configuration

Architecture
ModernBertForSequenceClassification
Context length (tokens)
8,192
Layers
22
Hidden size
768
Feed-forward size
1,152
Attention heads
12
Vocabulary size
256,000
Model type
modernbert

Identity and Version

Repository
Horizon-Labs/content-safety-guard-base
Publisher
Horizon Labs
Task
Text classification
Modality
Text
Library
transformers
Parameters
308M parameters
Languages
en, ar, de, es, fr, hi, it, ja
Revision
90ff0ffe86ca4951f7805c80b9aacc16aaf99f5b
First published
2026-09-26
Last updated
2026-09-27

Files and Weights

27 files, 3.1 GB in total. The weights are 3 files totalling 3.1 GB in onnx, safetensors.

Weights3 files · 3.1 GB
Configuration20 files · 125.6 KB
Tokenizer2 files · 34.4 MB
Documentation1 file · 14.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.2 GB 17b86a499790
onnx/model.onnxWeights1.2 GB d34295f75e4e
onnx/model_quantized.onnxWeights641.2 MB 8de48c2ded0f
code/safety/build_pool2.pyConfiguration1.2 KB —
code/safety/build_prompt_pool.pyConfiguration3.8 KB —
code/safety/build_safety_evals.pyConfiguration5.1 KB —
code/safety/build_safety_v0.pyConfiguration4.6 KB —
code/safety/build_safety_v2.pyConfiguration2.9 KB —
code/safety/build_safety_v3.pyConfiguration2.3 KB —
code/safety/cat_thresholds.pyConfiguration2.6 KB —
code/safety/eval_safety.pyConfiguration6.4 KB —
code/safety/label_teacher.pyConfiguration1.1 KB —
code/safety/qwenguard.pyConfiguration3.3 KB —
code/safety/train_safety.pyConfiguration7.8 KB —
code/train/export_onnx.pyConfiguration2.1 KB —
config.jsonConfiguration2.7 KB —
eval/baselines.jsonConfiguration62.8 KB —
eval/this_model.jsonConfiguration10.7 KB —
onnx/quantization_check.jsonConfiguration396 B —
special_tokens_map.jsonConfiguration636 B —
thresholds.jsonConfiguration2.7 KB —
training/train_log.jsonConfiguration1.8 KB —
training/val_metrics.jsonConfiguration495 B —
README.mdDocumentation14.4 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer34.4 MB aebee76d0312
tokenizer_config.jsonTokenizer46.4 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
3.1 GB
Download from Horizon Labs

Released by Horizon Labs through its official repository on Hugging Face. Read the license.

Built From

  • Derived from jhu-clsp/mmBERT-base
  • Quantized from jhu-clsp/mmBERT-base
  • Trained on (disclosed) CohereForAI/aya_dataset
  • Trained on (disclosed) JailbreakBench/JBB-Behaviors
  • Trained on (disclosed) OpenAssistant/oasst2
  • Trained on (disclosed) OpenSafetyLab/Salad-Data
  • Trained on (disclosed) allenai/WildChat-1M
  • Trained on (disclosed) google/civil_comments
  • Trained on (disclosed) nvidia/Nemotron-Safety-Guard-Dataset-v3

Memory Requirements

PrecisionWeights in memory
As published3.1 GB
16-bit0.6 GB
8-bit0.3 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About content-safety-guard-base

How much GPU memory does content-safety-guard-base need?

About 0.7 GB at 16-bit and 0.2 GB at 4-bit: the weights (308M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run content-safety-guard-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use content-safety-guard-base commercially?

Yes. content-safety-guard-base is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is content-safety-guard-base's context length?

8,192 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

prompt-injection-guard-base

Horizon Labs

A fast, multilingual classifier that flags prompt injection and jailbreak attempts, both in user messages (direct) and in untrusted content an AI agent reads: emails, web pages, documents, RAG chunks, and tool/API outputs (indirect). their clean counterparts, so it looks for instructions aimed at the AI, not for scary words. - Low false-alarm rate on look-alike benign text: 92.9% on NotInject, 99.2% on OR-Bench-hard. half the size and the same decisions as fp32 on our checks), transformers.js. Labels: SAFE (0) and INJECTION (1). This is the same convention as protectai/deberta-v3-base-prompt-injection-v2, so the model is a drop-in replacement in code and tools built for that one. INJECTION…

Open weights apache-2.0 308M parameters 8,192 tokens transformers

Model · Text classification

laya-multilingual

Convai Innovations

Non-autoregressive System 1 decision model covering 100+ languages. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with probabilities in a single forward pass. No text generation, so nothing to parse and nothing to hallucinate. Part of the Laya family — use this checkpoint for anything that is not English. Since laya 0.3.13 the default Router() keeps both english and this checkpoint resident, so a mixed workload no longer swaps checkpoints on every language change. For a server, load them up front so even the first request of each language is just a forward pass: router.attach("multilingual", agent) registers an Agent you already built, so a…

Open weights apache-2.0 322M parameters transformers

Model · Text classification

laya-coreai

Andrey Babikov

Laya typed decisions on Apple Silicon, running on the Core AI runtime — the successor to Core ML. This is a.aimodel asset exported from via Apple's coreai-torch bridge. It outputs choice / score / noul probabilities (and RL action logits) with zero generated tokens and no PyTorch, Core ML, Transformers, or cloud API at inference time. macOS 27+ (Core AI runtime), Python 3.10+. Validated on M3 Max / macOS 27.2. Validate the download end-to-end (all three specializations, timing, contract checks): Snake demo with the model (terminal game, reuses the laya-coreml UI + safety shield; automatically uses the B3 asset when present for ~2x game throughput): ~3× faster per pass than the fastest Core…

Open weights apache-2.0 322M parameters coreai

Model · Text classification

laya-pt-es-typed

Telepatia

This checkpoint fine-tunes convaiinnovations/laya-multilingual for native choice, score, and noul decisions in Portuguese and Spanish. It keeps the original 322M-parameter mmBERT architecture. It adds no inference component and does not generate text. It returns typed answers and probabilities in one forward pass. This is a text model. Inference takes a textual state plus typed questions. The second training stage used text decisions derived from public speech corpora, but this checkpoint does not accept audio by itself. The separate audio projector is not included. The official Laya SDK defines these primitives as follows: - choice: selects one key from a runtime-defined criteria object.…

Open weights apache-2.0 322M parameters laya

Model · Text classification

laya-multilingual

Scott Lamkin

Non-autoregressive System 1 decision model covering 100+ languages. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with probabilities in a single forward pass. No text generation, so nothing to parse and nothing to hallucinate. Part of the Laya family — use this checkpoint for anything that is not English. The default Router() keeps both english and this checkpoint resident, so a mixed workload no longer swaps checkpoints on every language change. For a server, load them up front so even the first request of each language is just a forward pass: router.attach("multilingual", agent) registers an Agent you already built, so a process that loaded…

Open weights apache-2.0 322M parameters transformers

Classifies GitHub issues written in any language as bug, feature, question or docs. A fine-tune of Laya multilingual (mmBERT-base) used by the laya-triage GitHub Action for non-English issues, next to the English model laya-triage-en. The same 500 NLBSE'23 validation issues, machine-translated with NLLB-200 into 13 languages. Accuracy (±3 points per language): laya-triage and Jev are within noise of each other across languages; both are far ahead of the untuned base. Translations can flatter a model trained on translations, so we also checked real issues: on 367 non-English issues opened in 2026 (never seen, written by people, not translated) accuracy went from 47.1% to 65.7%. Use it…

Open weights apache-2.0 322M parameters