Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with hierarchical global-context attention to catch both local artifacts (textures, blending seams) and global artifacts (lighting, structural inconsistency). A single architecture ships in two sizes and three domain-tuned checkpoints, working on both static images and video at the frame level. - Frame-level — one model handles both images and videos (frame-level inference + aggregation). - Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces. - Two variants — Fast (b0) for…
Open-weight model · Video classification
ms-eff-gcvit-deepfake-b0-celeb-df-v2
by YUNJE SEO KoreaPeter/ms-eff-gcvit-deepfake-b0-celeb-df-v2
Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT architecture for deepfake detection.
Runs On
What it takes to serve ms-eff-gcvit-deepfake-b0-celeb-df-v2 (9M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.0 GB | 0.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
By YUNJE SEO, published under mit, revision 7799a86ae22f.
Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with hierarchical global-context attention to catch both local artifacts (textures, blending seams) and global artifacts (lighting, structural inconsistency). A single architecture ships in two sizes and three domain-tuned checkpoints, working on both static images and video at the frame level. - Frame-level — one model handles both images and videos (frame-level inference + aggregation). - Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces. - Two variants — Fast (b0) for…
Read YUNJE SEO's full model card
Multi Scale Efficient Global Context Vision Transformer
GitHub Repository: HanMoonSub/DeepGuard
Live demo: DeepFake Video Detection
Live demo: DeepFake Image Detection
Live demo: DeepFake Detection XAI
Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with hierarchical global-context attention to catch both local artifacts (textures, blending seams) and global artifacts (lighting, structural inconsistency).
A single architecture ships in two sizes and three domain-tuned checkpoints, working on both static images and video at the frame level.
Core Features
- Frame-level — one model handles both images and videos (frame-level inference + aggregation).
- Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces.
- Two variants — Fast (b0) for real-time/edge, Pro (b5) for enterprise accuracy.
- timm-compatible — load via the
timminterface or thedeepguardpackage.
Model Specifications
| Spec | Detail |
|---|---|
| Task | Binary deepfake detection (real / fake) |
| Domain | Frame-level, spatial-domain |
| Input | Image or video (face-cropped) |
| Output | Sigmoid probability in [0, 1] — higher = more likely fake |
| Backbone | EfficientNet (ImageNet-1K pretrained) |
| Framework | PyTorch / timm |
Model Zoo
ms_eff_gcvit_b0 is Optimized for real-time inference and mobile deployment.
ms_eff_gcvit_b5 is Engineered for high-fidelity analysis and enterprise-grade accuracy.
| Config | Fast (b0) | Pro (b5) |
|---|---|---|
| Model name | ms_eff_gcvit_b0 |
ms_eff_gcvit_b5 |
| Backbone | tf_efficientnet_b0.ns_jft_in1k |
tf_efficientnet_b5.ns_jft_in1k |
| Resolution | 224×224 | 384×384 |
| Params (M) | 8.7 | 50.3 |
| FLOPs (G) | 0.87 | 13.64 |
Dataset: Celeb-DF-v2
A large-scale challenging dataset for deepfake forensics [Paper] [Download], featuring 590 YouTube celebrity videos with diverse ages, ethnic groups, and genders.
- [x] Source: 590 original YouTube videos (celebrities)
- [x] Synthesis: 5,639 deepfake videos generated from real videos
- [x] Subjects: Diverse ages, ethnicities, and genders
| Source | Real/Fake | Videos | Description |
|---|---|---|---|
celeb-real |
590 | Celebrity videos from YouTube | |
youtube-real |
300 | Additional YouTube videos | |
celeb-synthesis |
5,639 | Synthesized from celeb-real |
Test Evaluation
Trained and tested on the same dataset.
| Dataset | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| Celeb-DF-v2 | Fast | 0.9842 | 0.9965 | 0.0283 |
| Celeb-DF-v2 | Pro | 0.9981 | 0.9984 | 0.0089 |
Cross-Dataset Evaluation (Trained on Celeb DF(v2))
Generalization to unseen domains — trained on Celeb DF(v2)
| Tested on | Variant | Accuracy | AUC | Log Loss |
|---|---|---|---|---|
| KoDF | Fast | 0.4935 | 0.7258 | 1.3459 |
| KoDF | Pro | 0.4832 | 0.7160 | 1.4897 |
| FaceForensics++ | Fast | 0.5492 | 0.7301 | 1.0556 |
| FaceForensics++ | Pro | 0.5825 | 0.7307 | 0.8897 |
Model Usage
pip install deepguard
from transformers import pipeline
Image Classification
clf = pipeline(
"image-classification",
model="KoreaPeter/ms-eff-gcvit-deepfake-b0-celeb-df-v2",
trust_remote_code=True,
)
# ── Basic Inference ───────────────────────────────────────────────
result = clf("face.jpg")
# [{'label': 'fake', 'score': 0.9712}, {'label': 'real', 'score': 0.0288}]
# ── Custom Parameters ─────────────────────────────────────────────
result = clf(
"face.jpg",
margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2)
conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5)
min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01)
tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0)
top_k=1, # Number of top labels to return (default: all)
)
# [{'label': 'fake', 'score': 0.9712}]
Video Classification
clf = pipeline(
"video-classification",
model="KoreaPeter/ms-eff-gcvit-deepfake-b0-celeb-df-v2",
trust_remote_code=True,
)
# ── Basic Inference ───────────────────────────────────────────────
result = clf("video.mp4")
# [{'label': 'fake', 'score': 0.9634}, {'label': 'real', 'score': 0.0366}]
# ── Custom Parameters ─────────────────────────────────────────────
result = clf(
"video.mp4",
num_frames=20, # Number of frames to sample (default: 20)
margin_ratio=0.2, # Margin ratio around the detected face bbox (default: 0.2)
conf_thres=0.5, # Confidence threshold for YOLO face detection (default: 0.5)
min_face_ratio=0.01, # Minimum face-to-frame area ratio to process (default: 0.01)
tta_hflip=0.0, # Probability of horizontal flip for TTA (default: 0.0)
agg_mode="conf", # Aggregation mode: 'conf' | 'mean' | 'vote' (default: 'conf')
return_frame_scores=True, # Return per-frame scores (default: False)
)
# [{'label': 'fake', 'score': 0.9634},
# {'label': 'real', 'score': 0.0366},
# {'frame_scores': [0.97, 0.95, 0.98, ...], 'agg_mode': 'conf'}]
Deep Dive into Model
Part 1: CNN-based Patch Embedding for Spatial Inductive Bias
While traditional Vision Transformers (ViTs) utilize a Linear Projection for patch embedding, our proposed model adopts a CNN-based Patch Embedding module incorporating MBConvBlocks.
- Injecting Inductive Bias : Standard ViTs often suffer from a lack of inherent spatial inductive bias, typically necessitating massive datasets to learn fundamental visual structures from scratch. In contrast, our CNN-based module leverages overlapping receptive fields to facilitate information sharing between neighboring patches. By explicitly injecting this spatial bias into the architecture, the model achieves more stable and accelerated convergence during the training process.
Part 2: Long-Short Range Spatial Interaction
We utilizes two distinct types of self-attention to capture both long-range and short-range information across feature maps.
-
Local Window Attention: this model efficiently captures local textures and precise spatial details while maintaining linear computational complexity relative to the image size.
-
Global Window Attention: Unlike Swin Transformer, this module utilizes global-queries that interact with local window keys and values. This allows each local region to incorporate global context, effectively capturing long-range dependencies and providing a comprehensive understanding of the entire spatial structure
Part 3: Computational Efficiency
-
Efficient Backbone While both Xception and EfficientNet show great results on DeepFake benchmarks, EfficientNet is chosen for its superior computational efficiency. By utilizing MBconv (Inverted Residual Blocks) and depthwise convolutions, it achieves significantly lower FLOPS compared to Xception.
-
Window-based Attention: Instead of applying self-attention on raw images, this model operates on feature maps extracted from backbone blocks. By partitioning these maps into windows, the $O(N^2)$ complexity is restricted to the window size, siginificantly lowering the computational footprint.
Part 4: Multi-Scale Feature Map Fusion
Modern DeepFakes can leave very localized forgery region. To Capture this, we adopts a multi-scale strategy by extracting features from different levels of the backbone.
-
(Subtle Artifacts): High-Resolution feature maps are extracted from early backbone blocks(
l_block_idx) to capture like skin texture or boundary artifacts -
(Global Features): Low-Resolution feature maps are extracted from deeper blocks(
h_block_idx) to analyze overall lighting, shadows, and structural consistency. -
Feature Fusion: The Outputs from both branches (
L-GCViT and H-GCViT) are fused to make a comprehensive decision based on both local and global context.
Citation
@misc{deepguard2026,
title = {DeepGuard: Multi-Scale Efficient Global Context Vision Transformer for Deepfake Detection},
author = {seoyunje},
year = {2026},
url = {https://github.com/HanMoonSub/DeepGuard}
}
Configuration
- Architecture
- MsEffGCViTForImageClassification
- Model type
- ms_eff_gcvit
Identity and Version
- Repository
- KoreaPeter/ms-eff-gcvit-deepfake-b0-celeb-df-v2
- Publisher
- YUNJE SEO
- Task
- Video classification
- Modality
- Video
- Library
- transformers
- Parameters
- 9M parameters
- Languages
- Not stated by the source
- Revision
- 7799a86ae22f874729131946ca314ef5216abdaa
- First published
- 2026-06-23
- Last updated
- 2026-07-04
Files and Weights
13 files, 45.4 MB in total. The weights are 2 files totalling 42.8 MB in pt, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 36.5 MB | 1742825cb28f |
| yolov8n-face.pt | Weights | 6.4 MB | d545bf1add5a |
| config.json | Configuration | 1.5 KB | — |
| configuration_ms_eff_gcvit.py | Configuration | 1.5 KB | — |
| modeling_ms_eff_gcvit.py | Configuration | 1.3 KB | — |
| pipeline_ms_eff_gcvit.py | Configuration | 3.5 KB | — |
| pipeline_video_ms_eff_gcvit.py | Configuration | 5.4 KB | — |
| README.md | Documentation | 11.2 KB | — |
| celeb_df_v2_gcvit.png | Other | 83.9 KB | — |
| dual_branch.gif | Other | 2.3 MB | 3b17745fc2ef |
| ms_eff_gcvit.JPG | Other | 100.3 KB | 9cc6e353e563 |
| window_attention.JPG | Other | 45.2 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
License and Download
- License
- mit
- Access
- Open weights, no gate
- Download size
- 42.8 MB
Released by YUNJE SEO through its official repository on Hugging Face. Read the license.
Built From
- Derived from timm/tf_efficientnet_b0.ns_jft_in1k
- Trained on (disclosed) ILSVRC/imagenet-1k
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 42.8 MB |
| 16-bit | 0.0 GB |
| 8-bit | 0.0 GB |
| 4-bit | 0.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About ms-eff-gcvit-deepfake-b0-celeb-df-v2
How much GPU memory does ms-eff-gcvit-deepfake-b0-celeb-df-v2 need?
About 0 GB at 16-bit and 0 GB at 4-bit: the weights (9M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run ms-eff-gcvit-deepfake-b0-celeb-df-v2 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use ms-eff-gcvit-deepfake-b0-celeb-df-v2 commercially?
Yes. ms-eff-gcvit-deepfake-b0-celeb-df-v2 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
Similar Models
Multi-Scale Efficient Global Context Vision Transformer (MS-EffGCViT) is a hybrid CNN-ViT architecture for deepfake detection. It fuses CNN-driven spatial inductive bias with hierarchical global-context attention to catch both local artifacts (textures, blending seams) and global artifacts (lighting, structural inconsistency). A single architecture ships in two sizes and three domain-tuned checkpoints, working on both static images and video at the frame level. - Frame-level — one model handles both images and videos (frame-level inference + aggregation). - Cross-domain — robust on both East-Asian (KoDF) and Western (Celeb-DF-v2, FaceForensics++) faces. - Two variants — Fast (b0) for…
ViViT model as introduced in the paper ViViT: A Video Vision Transformer by Arnab et al. and first released in this repository. Disclaimer: The team releasing ViViT did not write a model card for this model so this model card has been written by the Hugging Face team. ViViT is an extension of the Vision Transformer (ViT) to video. We refer to the paper for details. The model is mostly meant to intended to be fine-tuned on a downstream task, like video classification. See the model hub to look for fine-tuned versions on a task that interests you. For code examples, we refer to the documentation.
ViViT model as introduced in the paper ViViT: A Video Vision Transformer by Arnab et al. and first released in this repository. Disclaimer: The team releasing ViViT did not write a model card for this model so this model card has been written by the Hugging Face team. ViViT is an extension of the Vision Transformer (ViT) to video. We refer to the paper for details. The model is mostly meant to intended to be fine-tuned on a downstream task, like video classification. See the model hub to look for fine-tuned versions on a task that interests you. For code examples, we refer to the documentation.
TimeSformer model pre-trained on Kinetics-600. It was introduced in the paper TimeSformer: Is Space-Time Attention All You Need for Video Understanding? by Tong et al. and first released in this repository. Disclaimer: The team releasing TimeSformer did not write a model card for this model so this model card has been written by fcakyon. You can use the raw model for video classification into one of the 600 possible Kinetics-600 labels. Here is how to use this model to classify a video: For more code examples, we refer to the documentation.
TimeSformer model pre-trained on Kinetics-400. It was introduced in the paper TimeSformer: Is Space-Time Attention All You Need for Video Understanding? by Tong et al. and first released in this repository. Disclaimer: The team releasing TimeSformer did not write a model card for this model so this model card has been written by fcakyon. You can use the raw model for video classification into one of the 400 possible Kinetics-400 labels. Here is how to use this model to classify a video: For more code examples, we refer to the documentation.