SAVRN
Search Contact SAVRN

Independent publisher

Star Duong

star092304

AI Engineer

Models in Library1
Datasets in Library0
Models on Hugging Face5
Followers1

Models

Model · Video classification

vi-sign-language-videomae-base

Star Duong

This repository houses a fine-tuned VideoMAE (Base) model optimized for multi-class Vietnamese Sign Language Recognition (VSLR). The model architecture adapts self-supervised video representations to accurately classify short video clips of sign gestures into distinct Vietnamese text labels. The model processes short video sequences by partitioning them into spatiotemporal patches, mapping sequential gestures (such as "Ăn", "Bệnh viện", "Xin lỗi") to their corresponding semantic classes. The training routine was monitored closely across key evaluation metrics to prevent overfitting while maximizing classification accuracy on the validation split. The plot below illustrates the progression…

Open weights mit 86M parameters transformers