SAVRN
Search Contact SAVRN

phobert-emotion-social · Model Card

phobert-emotion-social: Model Card

Written by Le Van Huy, published under mit, revision 0d954dabf3fa, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.

PhoBERT Emotion Recognition v1.1

Model Description

PhoBERT Emotion v1.1 is a fine-tuned sequence classification model based on vinai/phobert-base-v2 for 7-class emotion recognition in Vietnamese social media texts: * Enjoyment (Thích thú / Vui vẻ) * Sadness (Buồn bã) * Disgust (Chán ghét / Khinh bỉ) * Anger (Tức giận) * Fear (Sợ hãi) * Surprise (Ngạc nhiên) * Other (Khác / Trung tính)

The model is trained on a curated and balanced Vietnamese social corpus (11,193 samples) with anti-overfitting techniques (Cosine Annealing scheduler, early stopping, and weight decay). In addition to standard PyTorch weights, optimized ONNX FP32 and ONNX Dynamic INT8 formats are provided for high-speed, resource-efficient CPU inference in production environments.


Model Formats & Artifacts

Format File Path Model Size Target Use-Case
PyTorch FP32 model.safetensors ~515 MB Standard Hugging Face training & fine-tuning
ONNX FP32 onnx/phobert_emotion_fp32.onnx ~515 MB 100% precision preservation, ~15% faster CPU inference
ONNX INT8 onnx/phobert_emotion_int8.onnx ~130 MB Dynamic quantization, 74.8% memory saving, low-latency CPU serving

Intended Uses & Capabilities

  • Social Listening & Monitoring: Real-time emotion sentiment tracking on comments, posts, and forum threads.
  • Mental Health & Chatbot Companions: Identifying distressed, anxious, or joyful states in conversational agents.
  • Customer Experience Analysis: Detecting frustration, anger, or satisfaction in feedback and review texts.

How to Use

1. Using Hugging Face pipeline

from transformers import pipeline

classifier = pipeline(
    task="text-classification",
    model="huyleit/phobert-vi-emotion-v1.1", # Or local folder path
    tokenizer="huyleit/phobert-vi-emotion-v1.1"
)

text = "Hôm nay nhận được tin vui quá trời luôn!"
prediction = classifier(text)
print(prediction)

2. Using PyTorch (transformers)

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_id = "huyleit/phobert-vi-emotion-v1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

id2label = {
    0: "Enjoyment",
    1: "Sadness",
    2: "Disgust",
    3: "Anger",
    4: "Fear",
    5: "Surprise",
    6: "Other"
}

text = "Hôm nay nhận được tin trúng tuyển thực sự vui phát khóc luôn á!"
inputs = tokenizer(text, return_tensors="pt", max_length=128, truncation=True)

with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)

pred_id = torch.argmax(probs, dim=-1).item()
confidence = probs[0][pred_id].item()

print(f"Predicted: {id2label[pred_id]} (Confidence: {confidence:.4f})")

3. Using ONNX Runtime (Fast CPU Inference)

import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer

model_id = "huyleit/phobert-vi-emotion-v1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)

# Load ONNX quantized INT8 model for ultra-fast inference
session = ort.InferenceSession("onnx/phobert_emotion_int8.onnx", providers=["CPUExecutionProvider"])

id2label = {
    0: "Enjoyment",
    1: "Sadness",
    2: "Disgust",
    3: "Anger",
    4: "Fear",
    5: "Surprise",
    6: "Other"
}

text = "Thật sự rất sốc và không thể tin nổi chuyện này lại xảy ra!"
inputs = tokenizer(text, return_tensors="np", max_length=128, truncation=True)

ort_inputs = {
    "input_ids": inputs["input_ids"],
    "attention_mask": inputs["attention_mask"]
}
logits = session.run(["logits"], ort_inputs)[0]
pred_id = int(np.argmax(logits, axis=1)[0])

print(f"Predicted: {id2label[pred_id]}")

Evaluation Results

The model was evaluated on a clean held-out test set of 1,679 samples (evaluation/emotion/data/phobert_test.json).

Classification Report (Held-out Test Set)

Emotion Class Precision Recall F1-Score Support
Enjoyment 0.7064 0.7831 0.7428 295
Sadness 0.6217 0.6614 0.6409 251
Disgust 0.5727 0.4922 0.5294 256
Anger 0.6042 0.5823 0.5930 249
Fear 0.6287 0.7095 0.6667 210
Surprise 0.7861 0.7022 0.7418 225
Other 0.5615 0.5440 0.5526 193
Accuracy — — 0.6432 1679
Macro Average 0.6402 0.6392 0.6382 1679
Weighted Average 0.6425 0.6432 0.6413 1679

Runtime & Quantization Benchmark Comparison

Metric PyTorch FP32 ONNX FP32 ONNX INT8 (Quantized) Impact / Optimization
Disk Size 515.02 MB 515.26 MB 129.56 MB -74.8% (4x smaller)
Runtime RAM Usage ~268 MB ~540 MB ~160 MB Significant RAM reduction
Average CPU Latency 97.01 ms 82.65 ms 64.85 ms +33.2% faster
P95 Latency 113.63 ms 140.68 ms 113.83 ms Predictable tail latency
Throughput 10.3 req/s 12.1 req/s 15.4 req/s +49.5% request capacity
Accuracy 64.32% 64.32% 62.00% Preserved precision
Macro F1 63.82% 63.82% 61.86% Minimal degradation (-1.96%)

Training Details

Training Data

  • Total samples: 11,193 Vietnamese social media text samples (balanced corpus).
  • Split Ratio: 70% Train (7,835 samples) / 15% Validation (1,679 samples) / 15% Held-out Test (1,679 samples).

Training Hyperparameters

  • Base Architecture: vinai/phobert-base-v2 (135M parameters)
  • Max Sequence Length: 128
  • Batch Size: 16
  • Learning Rate: 1.5e-5
  • LR Scheduler: CosineAnnealingLR (Warmup steps: 200)
  • Weight Decay: 0.02 (L2 Regularization)
  • Precision: Mixed Precision FP16
  • Optimization Metric: Validation Macro F1 (eval_f1)
  • Early Stopping: Patience = 2 epochs
  • Hardware: Google Colab NVIDIA Tesla T4 GPU

Label Mapping

{
  "0": "Enjoyment",
  "1": "Sadness",
  "2": "Disgust",
  "3": "Anger",
  "4": "Fear",
  "5": "Surprise",
  "6": "Other"
}

Limitations and Bias

  • Social Media Slang & Teencode: The model performs robustly on standard text and common social media expressions, but extreme teencode or emerging neologisms may reduce classification performance.
  • Sarcasm & Irony: Sarcastic expressions where positive words carry negative emotional meaning remain inherently challenging for subword-level contextual embeddings without broader conversation history.
  • Subjectivity: Boundaries between emotions like Disgust vs Anger or Other vs Sadness can be subjective across different annotators.

Citation & Acknowledgements

If you use this model or its quantized variants, please cite the following foundational works:

@inproceedings{phobert,
  title     = {{PhoBERT: Pre-trained language models for Vietnamese}},
  author    = {Dat Quoc Nguyen and Anh Tuan Nguyen},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2020},
  pages     = {1037--1042},
  year      = {2020}
}

@inproceedings{ho2020vsmec,
  title     = {Emotion Recognition for Vietnamese Social Media Text},
  author    = {Vong Anh Ho and Duong Huynh-Cong Nguyen and Danh Hoang Nguyen and Pham-Nguyen Cuong and Duc-Vu Nguyen and Kiet Van Nguyen and Ngan Luu-Thuy Nguyen},
  booktitle = {2020 17th International Conference on Electrical Engineering/Electronics, Computer, Telecommunications and Information Technology (ECTI-CON)},
  year      = {2020}
}

@inproceedings{demszky2020goemotions,
  title     = {{GoEmotions: A Dataset of Fine-Grained Emotions}},
  author    = {Dorottya Demszky and Dana Movshovitz-Attias and Jeongwoo Ko and Alan Cowen and Gaurav Nemade and Sujith Ravi},
  booktitle = {58th Annual Meeting of the Association for Computational Linguistics (ACL)},
  year      = {2020}
}