A compact multilingual detector for prompt injection and jailbreak attempts (large language model guardrails). Given any user prompt, tool output, or document excerpt, it outputs the probability that the text is an attack on the LLM's instructions. Existing public injection detectors are either English-only (e.g. protectai's deberta models) or released under licenses many organizations cannot use (Meta's PromptGuard line). This model is permissively licensed, works in 17 languages, ships with its multilingual training corpus and is evaluated against a public baseline on shared test sets. Recommended threshold: 0.50 — conservative default (EN precision 0.83 on the attack test, safe-prompt…
Independent publisher
Horizon Labs
Horizon-Labs
Models
A fast, multilingual classifier that flags prompt injection and jailbreak attempts, both in user messages (direct) and in untrusted content an AI agent reads: emails, web pages, documents, RAG chunks, and tool/API outputs (indirect). their clean counterparts, so it looks for instructions aimed at the AI, not for scary words. - Low false-alarm rate on look-alike benign text: 92.9% on NotInject, 99.2% on OR-Bench-hard. half the size and the same decisions as fp32 on our checks), transformers.js. Labels: SAFE (0) and INJECTION (1). This is the same convention as protectai/deberta-v3-base-prompt-injection-v2, so the model is a drop-in replacement in code and tools built for that one. INJECTION…
A fast, multilingual classifier that flags prompt injection and jailbreak attempts, both in user messages (direct) and in untrusted content an AI agent reads: emails, web pages, documents, RAG chunks, and tool/API outputs (indirect). their clean counterparts, so it looks for instructions aimed at the AI, not for scary words. - Low false-alarm rate on look-alike benign text: 89.7% on NotInject, 99.5% on OR-Bench-hard. half the size and the same decisions as fp32 on our checks), transformers.js. Labels: SAFE (0) and INJECTION (1). This is the same convention as protectai/deberta-v3-base-prompt-injection-v2, so the model is a drop-in replacement in code and tools built for that one. INJECTION…
Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to 1,024 tokens). non-commercial sets). See Training. Part of Horizon Labs' open models (collection). Labels: notentailment (0) and entailment (1). For each candidate label the model scores whether the text entails "This example is {label}." (or your hypothesistemplate). Any NLI-style use works too: pass text and textpair to a text-classification pipeline. Accuracy, single-label (multilabel=False: the label with the highest entailment score…
Classify text in 30+ languages into any labels you choose, with no training. Use it with the transformers zero-shot-classification pipeline, like facebook/bart-large-mnli, but multilingual, smaller, and with an 8k-token context window (fine-tuned at up to 1,024 tokens). non-commercial sets). See Training. Part of Horizon Labs' open models (collection). Labels: notentailment (0) and entailment (1). For each candidate label the model scores whether the text entails "This example is {label}." (or your hypothesistemplate). Any NLI-style use works too: pass text and textpair to a text-classification pipeline. Accuracy, single-label (multilabel=False: the label with the highest entailment score…
A small multilingual classifier that flags unsafe user prompts and unsafe model responses for LLM applications, with harm categories. It is a fast encoder (ModernBERT architecture, mmBERT backbone) that you can run on CPU or in the browser in front of, or behind, any LLM. real-world prompts; evaluated on 17 (PolyGuard), 14 (textdetox) and 8 (Aya) languages. Part of the Horizon Labs guard family: prompt-injection-guard, Decision rule: flag when unsafe >= 0.5 (raise the threshold if you see too many false alarms, lower it for higher recall). Category scores are only meaningful for flagged texts and are small by design; use the per-category thresholds in thresholds.json (chosen on a validation…