phobert-vi-moderation-v1.1 · Model Card
phobert-vi-moderation-v1.1: Model Card
Written by Le Van Huy, published under mit, revision a7f50168351c, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.
PhoBERT Contextual Content Moderation v1.1
Model Description
PhoBERT Moderation v1.1 is a fine-tuned sequence classification model based on vinai/phobert-base-v2 for contextual text moderation in Vietnamese social networks, with a dedicated focus on mental health awareness and crisis intervention.
Traditional keyword-based moderation often struggles to distinguish between harmless personal venting and actual malicious hate speech, or fails to detect acute emotional distress. This model implements a nuanced 4-class taxonomy:
* CLEAN (An toàn / Tích cực)
* PROFANITY_VENTING (Xả bực dọc / Chửi thề vô hại)
* HATE_SPEECH (Thù ghét / Công kích xúc phạm)
* SELF_HARM_CRISIS (Khủng hoảng tâm lý / Nguy cơ tự hại)
Trained on a curated Vietnamese social corpus integrating UIT-ViHSD and localized mental health crisis data. Both standard PyTorch weights and high-efficiency ONNX Runtime (FP32 & Dynamic INT8) artifacts are included for production serving.
Contextual Taxonomy & Moderation Actions
| Label ID | Label Name | Contextual Meaning | Recommended System Action |
|---|---|---|---|
0 |
CLEAN |
Safe, constructive, or neutral conversation. | ALLOW (Publish normally) |
1 |
PROFANITY_VENTING |
Casual venting of frustration, colloquial swear words not targeting anyone. | ALLOW_WITH_WARNING (Publish with sensitive warning / flag for review) |
2 |
HATE_SPEECH |
Targeted harassment, toxic hostility, insults, or abusive speech. | BLOCK / HIDE (Hard block or soft shadow-hide) |
3 |
SELF_HARM_CRISIS |
Acute depressive crisis, suicidal ideation, or self-harm signals. | ALLOW_WITH_SUPPORT (Intervene with emergency hotline modal & supportive resources) |
Model Formats & Artifacts
| Format | File Path | Model Size | Target Use-Case |
|---|---|---|---|
| PyTorch FP32 | model.safetensors |
~515 MB | Standard Hugging Face training & fine-tuning |
| ONNX FP32 | onnx/phobert_moderation_fp32.onnx |
~515 MB | 100% precision preservation, ~15% faster CPU inference |
| ONNX INT8 | onnx/phobert_moderation_int8.onnx |
~130 MB | Dynamic quantization, 74.8% memory saving, low-latency CPU serving |
How to Use
1. Using Hugging Face pipeline
from transformers import pipeline
moderator = pipeline(
task="text-classification",
model="huyleit/phobert-vi-moderation-v1.1", # Or local folder path
tokenizer="huyleit/phobert-vi-moderation-v1.1"
)
text = "Mệt mỏi và bế tắc quá rồi, không biết phải tiếp tục thế nào..."
prediction = moderator(text)
print(prediction)
2. Using PyTorch (transformers)
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_id = "huyleit/phobert-vi-moderation-v1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
id2label = {
0: "CLEAN",
1: "PROFANITY_VENTING",
2: "HATE_SPEECH",
3: "SELF_HARM_CRISIS"
}
text = "Đm lại trễ xe buýt nữa, bực cả mình đi làm muộn!"
inputs = tokenizer(text, return_tensors="pt", max_length=128, truncation=True)
with torch.no_grad():
logits = model(**inputs).logits
probs = torch.softmax(logits, dim=-1)
pred_id = torch.argmax(probs, dim=-1).item()
confidence = probs[0][pred_id].item()
print(f"Predicted Label: {id2label[pred_id]} (Confidence: {confidence:.4f})")
3. Using ONNX Runtime (Fast CPU Inference)
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
model_id = "huyleit/phobert-vi-moderation-v1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
# Load ONNX quantized INT8 model for ultra-fast CPU inference
session = ort.InferenceSession("onnx/phobert_moderation_int8.onnx", providers=["CPUExecutionProvider"])
id2label = {
0: "CLEAN",
1: "PROFANITY_VENTING",
2: "HATE_SPEECH",
3: "SELF_HARM_CRISIS"
}
text = "Đồ thứ ngu dốt, biến đi cho khuất mắt tao!"
inputs = tokenizer(text, return_tensors="np", max_length=128, truncation=True)
ort_inputs = {
"input_ids": inputs["input_ids"],
"attention_mask": inputs["attention_mask"]
}
logits = session.run(["logits"], ort_inputs)[0]
pred_id = int(np.argmax(logits, axis=1)[0])
print(f"Predicted Label: {id2label[pred_id]}")
Evaluation Results
Evaluated on an independent, held-out test set of 2,441 samples (evaluation/moderation/data/v1.1/test.json).
Classification Report (Held-out Test Set)
| Label | Class Name | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|---|
| 0 | CLEAN | 0.8563 | 0.8642 | 0.8602 | 1200 |
| 1 | PROFANITY_VENTING | 0.5825 | 0.5310 | 0.5556 | 339 |
| 2 | HATE_SPEECH | 0.7281 | 0.7571 | 0.7423 | 527 |
| 3 | SELF_HARM_CRISIS | 0.9839 | 0.9787 | 0.9813 | 375 |
| Accuracy | — | — | 0.8124 | 2441 | |
| Macro Average | 0.7877 | 0.7827 | 0.7848 | 2441 | |
| Weighted Average | 0.8102 | 0.8124 | 0.8111 | 2441 |
Key Highlight: The model achieves an outstanding 98.13% F1-score on
SELF_HARM_CRISIS(Precision: 98.39%, Recall: 97.87%), ensuring high reliability when triggering urgent mental health interventions.
Runtime & Quantization Benchmark Comparison
| Metric | PyTorch FP32 | ONNX FP32 | ONNX INT8 (Quantized) | Impact / Optimization |
|---|---|---|---|---|
| Disk Size | 515.01 MB | 515.25 MB | 129.56 MB | -74.8% (4x smaller) |
| Runtime RAM Usage | ~268 MB | ~541 MB | ~158 MB | Significant RAM reduction |
| Average CPU Latency | 99.94 ms | 84.79 ms | 63.62 ms | +36.3% faster |
| P95 Latency | 180.37 ms | 140.16 ms | 113.37 ms | Predictable tail latency |
| Throughput | 10.0 req/s | 11.8 req/s | 15.7 req/s | +57.0% request capacity |
| Accuracy | 81.24% | 81.24% | 79.93% | Preserved precision |
| Macro F1 | 78.48% | 78.48% | 76.53% | Minimal degradation (-1.95%) |
Training Details
Training Data
- Dataset Composition: Merged dataset of UIT-ViHSD (Vietnamese Hate Speech Detection) and localized mental health / crisis corpora.
- Split Ratio: Stratified 70% Train / 15% Validation / 15% Test (2,441 samples held out).
Training Hyperparameters
- Base Architecture:
vinai/phobert-base-v2(135M parameters) - Max Sequence Length: 128
- Batch Size: 32
- Learning Rate:
2e-5 - Optimizer: AdamW (
weight_decay = 0.01) - Warmup Steps: 300
- Precision: Mixed Precision FP16
- Optimization Metric: Validation Macro F1 (
eval_f1) - Hardware: Google Colab NVIDIA Tesla T4 GPU
Label Mapping
{
"0": "CLEAN",
"1": "PROFANITY_VENTING",
"2": "HATE_SPEECH",
"3": "SELF_HARM_CRISIS"
}
Limitations and Bias
- Context-dependent Sarcasm: Sarcastic remarks containing superficially polite words may be misclassified as
CLEAN. - Dialect & Slang: Heavily misspelled words or emerging internet acronyms may diminish detection sensitivity.
- Human Disagreement: Distinguishing borderlines between emotional frustration (
PROFANITY_VENTING) and hostile toxicity (HATE_SPEECH) can carry subjective differences.
Citation & Acknowledgements
If you use this model or its quantized variants, please cite the following works:
@article{luu2021vihsd,
title = {Constructing a Vietnamese Dataset for Advanced Hate Speech Detection on Social Media},
author = {Son T. Luu and Kiet Van Nguyen and Ngan Luu-Thuy Nguyen},
journal = {ACM Transactions on Asian and Low-Resource Language Information Processing},
year = {2021}
}
@inproceedings{phobert,
title = {{PhoBERT: Pre-trained language models for Vietnamese}},
author = {Dat Quoc Nguyen and Anh Tuan Nguyen},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2020},
pages = {1037--1042},
year = {2020}
}