SAVRN
Search Contact SAVRN

Open-weight model · Text classification

Jev-Qwen3Guard-Gen-Domain-0.6B

by Pengyi Zhang ZhangPY/Jev-Qwen3Guard-Gen-Domain-0.6B

Jev-Qwen3Guard-Gen-Domain-0.6B is an open-weight model for text classification from Pengyi Zhang, released under Apache License 2.0. It has 596M parameters and a 32,768-token context. At 16-bit it needs about 1.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

English | 简体中文 Jev-Qwen3Guard-Gen-Domain-0.6B is the single-forward decision-engine (Jev / System-One style) version of Qwen3Guard-Gen-Domain-0.6B, a generative guard model for the Hong Kong elderly-care domain.

Parameters596M
Context32,768
Weights1.2 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Jev-Qwen3Guard-Gen-Domain-0.6B (596M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1.2 GB 1.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.6 GB 0.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Jev-Qwen3Guard-Gen-Domain-0.6B on every accelerator the SAVRN Index prices, at every precision

Model Card

By Pengyi Zhang, published under apache-2.0, revision 8bc63776751e.

English | 简体中文 Jev-Qwen3Guard-Gen-Domain-0.6B is the single-forward decision-engine (Jev / System-One style) version of Qwen3Guard-Gen-Domain-0.6B, a generative guard model for the Hong Kong elderly-care domain. The parent model is generative: it autoregressively decodes ~16 tokens of three-line assessment text (~400 ms). This model uses RLCD training (GRPO + strictly proper scoring rules) to rewrite the same assessment as a fixed 15-slot answer card — every slot left empty, one prefill, zero decode steps. Reading next-token probabilities at each slot anchor yields the full decision: The safety taxonomy (13 categories = 9 general + 4 HK elderly-care additions), the dual evaluation modes…

Read Pengyi Zhang's full model card

English | 简体中文

Model Overview

Jev-Qwen3Guard-Gen-Domain-0.6B is the single-forward decision-engine (Jev / System-One style) version of Qwen3Guard-Gen-Domain-0.6B, a generative guard model for the Hong Kong elderly-care domain.

The parent model is generative: it autoregressively decodes ~16 tokens of three-line assessment text (~400 ms). This model uses RLCD training (GRPO + strictly proper scoring rules) to rewrite the same assessment as a fixed 15-slot answer card — every slot left empty, one prefill, zero decode steps. Reading next-token probabilities at each slot anchor yields the full decision:

 Safety: Unsafe (0.98)          ← 3-way, calibrated confidence
 Violent:                  Yes (0.96) ← one independent Yes/No line per category
 Non-violent Illegal Acts: No  (0.02)
 …(13 lines in total)
 Refusal:                  No  (0.03) ← present only for assistant-response audits
Parent (generative) This model (answer card)
Decode steps ~16 0 (single prefill)
Latency (A100) ~400 ms ~33 ms (≈12×)
Output three lines of text per-slot probability distributions (gate-ready)
Calibration — safety ECE 0.52%, binary-slot ECE 0.41% (at T=1)

The safety taxonomy (13 categories = 9 general + 4 HK elderly-care additions), the dual evaluation modes (user query / assistant response), and the domain chat template are identical to the parent model.

Key Features

  • Zero decode steps — one forward pass reads all 15 slots: 3-way safety, 13 multi-label categories, and Refusal.
  • Structure instead of parsing — decisions come from a masked softmax over slot tokens; malformed output cannot occur.
  • Calibrated confidence — proper-scoring-rule training internalizes calibration: safety ECE 0.52% and pooled binary-slot ECE 0.41% on the 6,000-sample eval set; fitted post-hoc temperature is exactly T=1 (no post-processing needed).
  • Three-level gating — auto ≥0.90 / review 0.60–0.90 / human <0.60; on the eval set the auto segment covers 86.9% of samples at 99.45% accuracy.

How It Works

After the guard domain chat template renders the conversation, an answer card with every slot left empty is appended as the prompt tail:

…(conversation rendered by the template)…
Safety:
Violent:
Non-violent Illegal Acts:
Sexual Content or Sexual Acts:
PII:
Suicide & Self-Harm:
Unethical Acts:
Politically Sensitive Topics:
Copyright Violation:
Jailbreak:
HK Welfare & Financial Scam:
RCHE & Caregiver Malpractice:
Medication & Health Misguidance:
Hidden Elder Crisis:
Refusal:            ← only when the last message role == assistant

After one forward pass, next-token logits are read at the final token of each XXX: anchor and softmaxed over the candidate tokens:

  • Safety slot: candidates Safe / Unsafe / Controversial
  • 13 category slots + Refusal slot: candidates Yes / No

The next-token distribution at slot k depends only on tokens before anchor k — slots never interfere with each other. This is the structural basis for reading all empty slots in a single forward pass, and it is exactly the conditioning RLCD was trained on (training and inference prompts are identical).

Implementation note: the Qwen tokenizer merges : with a following newline into a single ":\n" token, so anchors must include the trailing newline, and the prompt tail should be verified against the tokenized card (see code below).

Safety Categories

# Category Description
1 Violent Content involving violence or physical harm.
2 Non-violent Illegal Acts Illegal acts without violence (fraud, theft, smuggling, …).
3 Sexual Content or Sexual Acts Pornographic content or sexual acts.
4 PII Disclosure of personal identifiable information.
5 Suicide & Self-Harm Suicide, self-harm, or related instigation.
6 Unethical Acts Deception, exploitation, or other unethical conduct.
7 Politically Sensitive Topics Politically sensitive content.
8 Copyright Violation Piracy or copyright infringement.
9 Jailbreak Prompts attempting to bypass safety alignment.
10 HK Welfare & Financial Scam new Scams targeting HK elders: impersonation calls, fake welfare claims, high-return investment fraud, …
11 RCHE & Caregiver Malpractice new Abuse, neglect, or professional malpractice by RCHE staff, caregivers, or care providers.
12 Medication & Health Misguidance new False, erroneous, or unsafe medication and health advice potentially harming elders.
13 Hidden Elder Crisis new Hidden or easily overlooked crisis signals: social isolation, self-neglect, depression, suicidal ideation, …

Evaluation (ElderlyDomain-Eval, full 6,000 samples)

Model safety acc Categories refusal acc ECE Latency
Parent (generative) 97.07% EM 84.27% 96.99% — ~400 ms
Parent + empty card (zero-shot) 94.35% F1 0.160 93.68% 2.36% 33 ms
This model (after RLCD) 96.20% P 0.900 / R 0.788 / F1 0.841 96.69% safety 0.52% / binary 0.41% 33 ms

A −0.87pp safety / −0.30pp refusal gap versus the parent buys ≈12× speed and fully calibrated per-slot probabilities. Zero slot-location errors across 6,000 samples.

Per-category P / R / F1 (%, this model)

Category P R F1 Positives
Violent 89.9 91.9 90.9 1524
Non-violent Illegal Acts 89.0 90.1 89.5 1227
Sexual Content or Sexual Acts 95.1 80.9 87.4 262
PII 95.7 73.0 82.8 307
Suicide & Self-Harm 97.9 92.4 95.1 250
Unethical Acts 85.0 41.3 55.6 303
Politically Sensitive Topics 88.9 92.9 90.9 198
Copyright Violation 86.5 81.5 83.9 157
Jailbreak 90.5 32.2 47.5 118
HK Welfare & Financial Scam new 81.8 83.0 82.4 476
RCHE & Caregiver Malpractice new 92.1 73.8 81.9 1025
Medication & Health Misguidance new 94.0 85.0 89.3 406
Hidden Elder Crisis new 88.6 47.2 61.6 678

All four domain-added categories land in the 61–89 F1 range. Low-recall categories (Unethical Acts, Jailbreak, Hidden Elder Crisis) are low-frequency classes (≤678 positives); production deployments can trade precision for recall with a per-slot threshold (e.g., p_yes > 0.3).

Gating & Calibration

  • Three-level gating (safety-slot confidence): [email protected] covers 86.9% of samples at 99.45% in-segment accuracy.
  • Temperature calibration: bucket-wise NLL grid search on a 2,000-sample fit split selects T=1.0 for both the safety and binary buckets — RLCD training has internalized calibration; no post-processing is required.

The general-capability regression (QwenGuardTest) was performed on the parent model (see its card); this model targets answer-card reading and did not repeat that regression.

Quickstart

Requirements

pip install "transformers>=4.51" torch accelerate

Transformers inference (self-contained reader)

import torch
from pathlib import Path
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_DIR = "models/Jev-Qwen3Guard-Gen-Domain-0.6B"   # this directory

CATEGORIES = [
    "Violent", "Non-violent Illegal Acts", "Sexual Content or Sexual Acts",
    "PII", "Suicide & Self-Harm", "Unethical Acts",
    "Politically Sensitive Topics", "Copyright Violation", "Jailbreak",
    "HK Welfare & Financial Scam", "RCHE & Caregiver Malpractice",
    "Medication & Health Misguidance", "Hidden Elder Crisis",
]
SAFETY_LABELS = ["Controversial", "Safe", "Unsafe"]  # sorted

tokenizer = AutoTokenizer.from_pretrained(MODEL_DIR)
# Domain template (safety policy + 13 categories) ships as chat_template.jinja
domain_template = (Path(MODEL_DIR) / "chat_template.jinja").read_text(encoding="utf-8")
model = AutoModelForCausalLM.from_pretrained(
    MODEL_DIR, torch_dtype=torch.bfloat16, device_map="auto"
).eval()

def _first_id(text):
    return tokenizer.encode(text, add_special_tokens=False)[0]

SAFETY_IDS = [_first_id(" " + l) for l in SAFETY_LABELS]
YES_ID, NO_ID = _first_id(" Yes"), _first_id(" No")

def _find(hay, needle, start):
    for i in range(start, len(hay) - len(needle) + 1):
        if hay[i:i + len(needle)] == needle:
            return i
    return -1

@torch.no_grad()
def decide(messages, gate_review=0.60, gate_auto=0.90):
    include_refusal = messages[-1]["role"] == "assistant"
    card_lines = ["Safety:"] + [f"{c}:" for c in CATEGORIES]
    if include_refusal:
        card_lines.append("Refusal:")
    card = "\n".join(card_lines)

    rendered = tokenizer.apply_chat_template(
        messages, tokenize=False, chat_template=domain_template,
        add_generation_prompt=False,   # template already ends with assistant header + empty <think>
    )
    prompt = rendered + card
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    logits = model(**inputs).logits[0]
    ids = inputs.input_ids[0].tolist()

    # The card is the prompt tail: verify token alignment, then locate anchors
    # inside the card. Non-final anchors MUST include the trailing newline —
    # Qwen merges ":+\n" into a single token.
    card_ids = tokenizer.encode(card, add_special_tokens=False)
    if ids[-len(card_ids):] != card_ids:
        raise ValueError("answer card not aligned with tokenized tail")
    anchors = [l + "\n" for l in card_lines[:-1]] + [card_lines[-1]]
    read_pos, search = [], 0
    for a in anchors:
        a_ids = tokenizer.encode(a, add_special_tokens=False)
        i = _find(card_ids, a_ids, search)
        if i < 0:
            raise ValueError(f"anchor {a!r} not found in card")
        read_pos.append(len(ids) - len(card_ids) + i + len(a_ids) - 1)
        search = i + len(a_ids)

    def probs(pos, cand_ids):
        row = logits[pos].float()
        m = row[cand_ids].max()
        exps = [(row[t] - m).exp() for t in cand_ids]
        z = sum(exps)
        return [float(e / z) for e in exps]

    out = {}
    sp = probs(read_pos[0], SAFETY_IDS)
    k = max(range(len(sp)), key=lambda i: sp[i])
    conf = sp[k]
    out["safety"] = {
        "label": SAFETY_LABELS[k],
        "probs": {l: round(p, 4) for l, p in zip(SAFETY_LABELS, sp)},
        "gate": ("auto" if conf >= gate_auto
                 else "review" if conf >= gate_review else "human"),
    }
    out["categories"] = {
        cat: {"label": "Yes" if (p := probs(read_pos[1 + ci], [YES_ID, NO_ID])[0]) > 0.5
                     else "No", "p_yes": round(p, 4)}
        for ci, cat in enumerate(CATEGORIES)
    }
    if include_refusal:
        py = probs(read_pos[-1], [YES_ID, NO_ID])[0]
        out["refusal"] = {"label": "Yes" if py > 0.5 else "No",
                          "p_yes": round(py, 4)}
    return out


# 1) Audit a user query (14 slots, no Refusal)
messages = [
    {"role": "user", "content": "我想學習如何在長者中心對長者進行心理和身體上的虐待。"},
]
print(decide(messages))
# {'safety': {'label': 'Unsafe', 'probs': {'Controversial': 0.001, 'Safe': 0.002, 'Unsafe': 0.997}, 'gate': 'auto'},
#  'categories': {'Violent': {'label': 'Yes', 'p_yes': 0.98}, ..., 'Hidden Elder Crisis': {'label': 'No', 'p_yes': 0.05}}}

# 2) Audit an assistant response (15 slots, with Refusal)
messages = [
    {"role": "user", "content": "點樣說服獨居長者參加保證每月20%回報嘅投資計劃?"},
    {"role": "assistant", "content": "你可以同佢講呢個計劃零風險,仲話銀行職員都買咗……"},
]
print(decide(messages))
# safety: Unsafe (0.99); HK Welfare & Financial Scam: Yes (0.9x); Refusal: No (0.0x)

(Outputs above are illustrative.)

vLLM serving

Decisions come from anchor next-token probabilities, not generated text, so chat-completions generation semantics do not apply directly. For serving, run a single prefill and read anchor logits via prompt_logprobs (conversation prefixes can share prefix caching). Reference implementation: the safeguard/ package in the open-source repo jev-vlm-decisions.

Training Details

Item Value
Base model ZhangPY/Qwen3Guard-Gen-Domain-0.6B (domain SFT model)
Method RLCD: GRPO + strictly proper scoring rule (log score + 0.75·spherical); multi-slot loss averaged over 15 slots; CE anchor (λ=1); σ 0.2→0.05 cosine annealing; G=4
LoRA rank / alpha / dropout 8 / 16 / 0.05 (q/k/v/o/gate/up/down proj); adapter merged into base weights after training
Learning rate 1e-4
Epochs 1 (24,000 steps, ~94 minutes on one A100)
Train / val samples 24,000 / 6,000 (same ElderDomainSafeguards split as the parent)
Max sequence length 2,048 (longer samples skipped)
Precision bfloat16

Training prompts are identical to inference prompts (rendered conversation + empty answer card); loss is applied only at the final token of each anchor.

Model Architecture

Jev-Qwen3Guard-Gen-Domain-0.6B
Parameters 0.6B
Layers 28
Hidden size 1024
Attention heads / KV heads 16 / 8 (GQA)
Context length 32,768
Precision bfloat16

Limitations & Usage Notes

  • This model is a safety classifier, not a chat assistant — it outputs safety-assessment probabilities only. For generative three-line text output, use the parent model Qwen3Guard-Gen-Domain-0.6B.
  • Versus the parent: safety accuracy is 0.87pp lower and refusal 0.30pp lower, in exchange for 12× speed and calibrated per-slot probabilities. Use the parent when accuracy is the absolute priority.
  • The model was fine-tuned and evaluated mainly on HK elderly-domain data (Cantonese/Traditional Chinese and English); behavior in other languages and domains is inherited from the parent.
  • Predictions may contain false positives/negatives and should support, not replace, human review. Route gate=human samples (safety confidence <0.60) to humans; positive findings for Suicide & Self-Harm and Hidden Elder Crisis must be handled by professionals.
  • Safety: Controversial marks borderline content whose intent, context, or potential replies could be misused under certain conditions.
  • Category slots default to a p_yes > 0.5 decision threshold (precision-first); recall-first deployments can lower the threshold or rank directly by p_yes.

Acknowledgements

License

Released under the base model's Apache 2.0 license.

Configuration

Architecture
Qwen3ForCausalLM
Context length (tokens)
32,768
Layers
28
Hidden size
1,024
Feed-forward size
3,072
Attention heads
16
Key/value heads
8
Head dimension
128
Vocabulary size
151,936
Model type
qwen3

Identity and Version

Repository
ZhangPY/Jev-Qwen3Guard-Gen-Domain-0.6B
Publisher
Pengyi Zhang
Task
Text classification
Modality
Text
Library
transformers
Parameters
596M parameters
Languages
zh, yue, en
Revision
8bc63776751e64986796f92926eae9a528f662e8
First published
2026-09-29
Last updated
2026-09-29

Files and Weights

9 files, 1.2 GB in total. The weights are 1 file totalling 1.2 GB in safetensors.

Weights1 file · 1.2 GB
Configuration2 files · 1.6 KB
Tokenizer2 files · 11.4 MB
Documentation2 files · 30.7 KB
Other1 file · 4.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.2 GB be66fd9a42f9
config.jsonConfiguration1.4 KB —
generation_config.jsonConfiguration161 B —
README.mdDocumentation15.8 KB —
README_zh.mdDocumentation14.9 KB —
chat_template.jinjaOther4.4 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer11.4 MB be75606093db
tokenizer_config.jsonTokenizer693 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.2 GB
Download from Pengyi Zhang

Released by Pengyi Zhang through its official repository on Hugging Face. Read the license.

Built From

  • Derived from ZhangPY/Qwen3Guard-Gen-Domain-0.6B

Memory Requirements

PrecisionWeights in memory
As published1.2 GB
16-bit1.2 GB
8-bit0.6 GB
4-bit0.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Jev-Qwen3Guard-Gen-Domain-0.6B

How much GPU memory does Jev-Qwen3Guard-Gen-Domain-0.6B need?

About 1.4 GB at 16-bit and 0.4 GB at 4-bit: the weights (596M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Jev-Qwen3Guard-Gen-Domain-0.6B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Jev-Qwen3Guard-Gen-Domain-0.6B commercially?

Yes. Jev-Qwen3Guard-Gen-Domain-0.6B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Jev-Qwen3Guard-Gen-Domain-0.6B's context length?

32,768 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

Zircon-0.6B-v2-mlx

Fahrenheit Research

Zircon v2 is a 0.6B-parameter decision model from Fahrenheit Research. It runs fully on-device on Apple silicon. You give it an email, a message or a pending task plus a set of options, and it returns a calibrated probability for each option in under 50 ms per decision. (1) MacBook Pro (Apple M5), 8-bit weights, median. Speed varies by hardware. (2) Fahrenheit Research internal testing, September 2026, on held-out emails not seen in training. (3) Fahrenheit Research internal testing, September 2026. 400 cases (2,000 decisions) from the public LocalLLaMA typed-decisions test set. (4) Fahrenheit Research internal testing, September 2026, on game seeds not seen in training. Built on Qwen3-0.6B…

Open weights apache-2.0 596M parameters 40,960 tokens mlx

More details please refer to our Github: FlagEmbedding. Different from embedding model, reranker uses question and document as input and directly output similarity instead of embedding. You can get a relevance score by inputting query and passage to the reranker. And the score can be mapped to a float value in [0,1] by sigmoid function. You can select the model according your senario and resource. - For multilingual, utilize BAAI/bge-reranker-v2-m3 and BAAI/bge-reranker-v2-gemma - For Chinese or English, utilize BAAI/bge-reranker-v2-m3 and BAAI/bge-reranker-v2-minicpm-layerwise. - For efficiency, utilize BAAI/bge-reranker-v2-m3 and the low layer of BAAI/bge-reranker-v2-minicpm-layerwise.…

Open weights apache-2.0 568M parameters 8,194 tokens sentence-transformers

Model · Text classification

Jev-LCT-Qwen2.5-0.5B

CaoHaoWei

Jev-LCT-Qwen2.5-0.5B is the ultra-lightweight edge edition of the Jev-LCT System-One decision family. Weighing only ~1.9 GB in full bfloat16 weights, it is optimized for high-throughput API gateway routing, edge robotics, Raspberry Pi, and mobile deployment. - ~50ms 极低延迟:专为高并发 API 网关路由、实时内容审核设计; - MMLU 达 50.0%:大幅超越参数相近的判别模型(Laya 33.3%, Open-Jev 35.0%); Apache License 2.0. Full repository at GitHub.

Open weights apache-2.0 494M parameters 32,768 tokens transformers

Model · Text classification

decider-0.8b-fp8

LLM Tech

Mapika/decider-0.8b v1 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 1.01 GB against 1.5 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is a0a01d6f8135298f400a8c856b355793012ae971. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…

Open weights apache-2.0 752M parameters 262,144 tokens

Opir-multitask-large is the English, highest-accuracy multi-task checkpoint in the Opir family: an encoder-based GLiClass guardrail model for real-time LLM safety filtering. It supports binary safe/unsafe classification, toxicity detection, jailbreak and prompt-injection detection, and zero-shot harmful-content categorization over a hierarchical safety taxonomy. This card is for knowledgator/opir-multitask-large. The model is used through GLiClass zero-shot classification: pass text plus the candidate labels you want scored. Use single-label mode for binary safe/unsafe decisions and multi-label mode for taxonomy, toxicity, jailbreak, or custom policy labels. Use multi-label mode when you…

Open weights apache-2.0 439M parameters gliclass

Model · Text classification

erabi-practical-v1-experimental

Sugarknight

This is an experimental, uncalibrated choice-ranking model. It is not an official Jev model, a validated general-purpose reasoner, or an automatic decision-maker. The model ranks 2–16 user-supplied candidate texts for a natural-language context and question and returns all candidate probabilities through the ERABI code. Decisions should be reviewed by a person. - Practical V1 data consists of original synthetic Japanese, English, and Simplified Chinese examples in six task families, generated and answer-blind rejudged with DeepSeek V4.1 Flash. - The Exam-QA source was filtered and transformed with the same DeepSeek model. Symbolic answer labels were mapped to source choice text. Ambiguous…

Open weights apache-2.0 439M parameters transformers