SAVRN
Search Contact SAVRN

phobert-vi-moderation-v1.1 · Model Card

phobert-vi-moderation-v1.1: Model Card

Written by Le Van Huy, published under mit, revision a7f50168351c, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.

PhoBERT Contextual Content Moderation v1.1

Model Description

PhoBERT Moderation v1.1 is a fine-tuned sequence classification model based on vinai/phobert-base-v2 for contextual text moderation in Vietnamese social networks, with a dedicated focus on mental health awareness and crisis intervention.

Traditional keyword-based moderation often struggles to distinguish between harmless personal venting and actual malicious hate speech, or fails to detect acute emotional distress. This model implements a nuanced 4-class taxonomy: * CLEAN (An toàn / Tích cực) * PROFANITY_VENTING (Xả bực dọc / Chửi thề vô hại) * HATE_SPEECH (Thù ghét / Công kích xúc phạm) * SELF_HARM_CRISIS (Khủng hoảng tâm lý / Nguy cơ tự hại)

Trained on a curated Vietnamese social corpus integrating UIT-ViHSD and localized mental health crisis data. Both standard PyTorch weights and high-efficiency ONNX Runtime (FP32 & Dynamic INT8) artifacts are included for production serving.


Contextual Taxonomy & Moderation Actions

Label ID Label Name Contextual Meaning Recommended System Action
0 CLEAN Safe, constructive, or neutral conversation. ALLOW (Publish normally)
1 PROFANITY_VENTING Casual venting of frustration, colloquial swear words not targeting anyone. ALLOW_WITH_WARNING (Publish with sensitive warning / flag for review)
2 HATE_SPEECH Targeted harassment, toxic hostility, insults, or abusive speech. BLOCK / HIDE (Hard block or soft shadow-hide)
3 SELF_HARM_CRISIS Acute depressive crisis, suicidal ideation, or self-harm signals. ALLOW_WITH_SUPPORT (Intervene with emergency hotline modal & supportive resources)

Model Formats & Artifacts

Format File Path Model Size Target Use-Case
PyTorch FP32 model.safetensors ~515 MB Standard Hugging Face training & fine-tuning
ONNX FP32 onnx/phobert_moderation_fp32.onnx ~515 MB 100% precision preservation, ~15% faster CPU inference
ONNX INT8 onnx/phobert_moderation_int8.onnx ~130 MB Dynamic quantization, 74.8% memory saving, low-latency CPU serving

How to Use

1. Using Hugging Face pipeline

from transformers import pipeline

moderator = pipeline(
    task="text-classification",
    model="huyleit/phobert-vi-moderation-v1.1", # Or local folder path
    tokenizer="huyleit/phobert-vi-moderation-v1.1"
)

text = "Mệt mỏi và bế tắc quá rồi, không biết phải tiếp tục thế nào..."
prediction = moderator(text)
print(prediction)

2. Using PyTorch (transformers)

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_id = "huyleit/phobert-vi-moderation-v1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

id2label = {
    0: "CLEAN",
    1: "PROFANITY_VENTING",
    2: "HATE_SPEECH",
    3: "SELF_HARM_CRISIS"
}

text = "Đm lại trễ xe buýt nữa, bực cả mình đi làm muộn!"
inputs = tokenizer(text, return_tensors="pt", max_length=128, truncation=True)

with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)

pred_id = torch.argmax(probs, dim=-1).item()
confidence = probs[0][pred_id].item()

print(f"Predicted Label: {id2label[pred_id]} (Confidence: {confidence:.4f})")

3. Using ONNX Runtime (Fast CPU Inference)

import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer

model_id = "huyleit/phobert-vi-moderation-v1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)

# Load ONNX quantized INT8 model for ultra-fast CPU inference
session = ort.InferenceSession("onnx/phobert_moderation_int8.onnx", providers=["CPUExecutionProvider"])

id2label = {
    0: "CLEAN",
    1: "PROFANITY_VENTING",
    2: "HATE_SPEECH",
    3: "SELF_HARM_CRISIS"
}

text = "Đồ thứ ngu dốt, biến đi cho khuất mắt tao!"
inputs = tokenizer(text, return_tensors="np", max_length=128, truncation=True)

ort_inputs = {
    "input_ids": inputs["input_ids"],
    "attention_mask": inputs["attention_mask"]
}
logits = session.run(["logits"], ort_inputs)[0]
pred_id = int(np.argmax(logits, axis=1)[0])

print(f"Predicted Label: {id2label[pred_id]}")

Evaluation Results

Evaluated on an independent, held-out test set of 2,441 samples (evaluation/moderation/data/v1.1/test.json).

Classification Report (Held-out Test Set)

Label Class Name Precision Recall F1-Score Support
0 CLEAN 0.8563 0.8642 0.8602 1200
1 PROFANITY_VENTING 0.5825 0.5310 0.5556 339
2 HATE_SPEECH 0.7281 0.7571 0.7423 527
3 SELF_HARM_CRISIS 0.9839 0.9787 0.9813 375
Accuracy — — 0.8124 2441
Macro Average 0.7877 0.7827 0.7848 2441
Weighted Average 0.8102 0.8124 0.8111 2441

Key Highlight: The model achieves an outstanding 98.13% F1-score on SELF_HARM_CRISIS (Precision: 98.39%, Recall: 97.87%), ensuring high reliability when triggering urgent mental health interventions.

Runtime & Quantization Benchmark Comparison

Metric PyTorch FP32 ONNX FP32 ONNX INT8 (Quantized) Impact / Optimization
Disk Size 515.01 MB 515.25 MB 129.56 MB -74.8% (4x smaller)
Runtime RAM Usage ~268 MB ~541 MB ~158 MB Significant RAM reduction
Average CPU Latency 99.94 ms 84.79 ms 63.62 ms +36.3% faster
P95 Latency 180.37 ms 140.16 ms 113.37 ms Predictable tail latency
Throughput 10.0 req/s 11.8 req/s 15.7 req/s +57.0% request capacity
Accuracy 81.24% 81.24% 79.93% Preserved precision
Macro F1 78.48% 78.48% 76.53% Minimal degradation (-1.95%)

Training Details

Training Data

  • Dataset Composition: Merged dataset of UIT-ViHSD (Vietnamese Hate Speech Detection) and localized mental health / crisis corpora.
  • Split Ratio: Stratified 70% Train / 15% Validation / 15% Test (2,441 samples held out).

Training Hyperparameters

  • Base Architecture: vinai/phobert-base-v2 (135M parameters)
  • Max Sequence Length: 128
  • Batch Size: 32
  • Learning Rate: 2e-5
  • Optimizer: AdamW (weight_decay = 0.01)
  • Warmup Steps: 300
  • Precision: Mixed Precision FP16
  • Optimization Metric: Validation Macro F1 (eval_f1)
  • Hardware: Google Colab NVIDIA Tesla T4 GPU

Label Mapping

{
  "0": "CLEAN",
  "1": "PROFANITY_VENTING",
  "2": "HATE_SPEECH",
  "3": "SELF_HARM_CRISIS"
}

Limitations and Bias

  • Context-dependent Sarcasm: Sarcastic remarks containing superficially polite words may be misclassified as CLEAN.
  • Dialect & Slang: Heavily misspelled words or emerging internet acronyms may diminish detection sensitivity.
  • Human Disagreement: Distinguishing borderlines between emotional frustration (PROFANITY_VENTING) and hostile toxicity (HATE_SPEECH) can carry subjective differences.

Citation & Acknowledgements

If you use this model or its quantized variants, please cite the following works:

@article{luu2021vihsd,
  title   = {Constructing a Vietnamese Dataset for Advanced Hate Speech Detection on Social Media},
  author  = {Son T. Luu and Kiet Van Nguyen and Ngan Luu-Thuy Nguyen},
  journal = {ACM Transactions on Asian and Low-Resource Language Information Processing},
  year    = {2021}
}

@inproceedings{phobert,
  title     = {{PhoBERT: Pre-trained language models for Vietnamese}},
  author    = {Dat Quoc Nguyen and Anh Tuan Nguyen},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2020},
  pages     = {1037--1042},
  year      = {2020}
}