VideoMAE model pre-trained for 2400 epochs in a self-supervised way and fine-tuned in a supervised way on Something-Something V2. It was introduced in the paper VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training by Tong et al. and first released in this repository. Disclaimer: The team releasing VideoMAE did not write a model card for this model so this model card has been written by the Hugging Face team. VideoMAE is an extension of Masked Autoencoders (MAE) to video. The architecture of the model is very similar to that of a standard Vision Transformer (ViT), with a decoder on top for predicting pixel values for masked patches. Videos are…
Open weights
cc-by-nc-4.0
transformers
TimeSformer model pre-trained on Something Something v2. It was introduced in the paper TimeSformer: Is Space-Time Attention All You Need for Video Understanding? by Tong et al. and first released in this repository. Disclaimer: The team releasing TimeSformer did not write a model card for this model so this model card has been written by fcakyon. You can use the raw model for video classification into one of the 174 possible Something Something v2 labels. Here is how to use this model to classify a video: For more code examples, we refer to the documentation.
Open weights
cc-by-nc-4.0
transformers
VideoMAE model pre-trained on Something-Something-v2 for 2400 epochs in a self-supervised way. It was introduced in the paper VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training by Tong et al. and first released in this repository. Disclaimer: The team releasing VideoMAE did not write a model card for this model so this model card has been written by the Hugging Face team. VideoMAE is an extension of Masked Autoencoders (MAE) to video. The architecture of the model is very similar to that of a standard Vision Transformer (ViT), with a decoder on top for predicting pixel values for masked patches. Videos are presented to the model as a sequence of…
Open weights
cc-by-nc-4.0
94M parameters
transformers
This model is a fine-tuned version of MCG-NJU/videomae-base on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 8 - evalbatchsize: 8 - lrschedulertype: linear - trainingsteps: 370 - Transformers 5.16.1 - Pytorch 2.14.0+cu126 - Datasets 5.0.1 - Tokenizers 0.23.2
Open weights
cc-by-nc-4.0
86M parameters
transformers
This model is a fine-tuned version of MCG-NJU/videomae-base on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 8 - evalbatchsize: 8 - lrschedulertype: linear - trainingsteps: 370 - Transformers 5.16.1 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.23.1
Open weights
cc-by-nc-4.0
86M parameters
transformers
VideoMAE model pre-trained for 1600 epochs in a self-supervised way and fine-tuned in a supervised way on Kinetics-400. It was introduced in the paper VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training by Tong et al. and first released in this repository. Disclaimer: The team releasing VideoMAE did not write a model card for this model so this model card has been written by the Hugging Face team. VideoMAE is an extension of Masked Autoencoders (MAE) to video. The architecture of the model is very similar to that of a standard Vision Transformer (ViT), with a decoder on top for predicting pixel values for masked patches. Videos are presented to…
Open weights
cc-by-nc-4.0
transformers
M
Model · Video classification
Mai
This model is a fine-tuned version of MCG-NJU/videomae-large on the Deepfake Detection Challenge dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 5e-05 - trainbatchsize: 8 - evalbatchsize: 8 - lrschedulertype: linear - lrschedulerwarmupratio: 0.1 - trainingsteps: 4470 - mixedprecisiontraining: Native AMP - Transformers 4.44.2 - Pytorch 2.5.0+cu121 - Datasets 3.1.0 - Tokenizers 0.19.1
Open weights
cc-by-nc-4.0
304M parameters
transformers
VideoMAEv2-giant model pre-trained for 1200 epochs in a self-supervised way on UnlabeldHybrid-1M dataset. It was introduced in the paper [[CVPR23]VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking](https://arxiv.org/abs/2203.12602) by Wang et al. and first released in GitHub. You can use the raw model for video feature extraction. Here is how to use this model to extract a video feature
Open weights
cc-by-nc-4.0
1B parameters
VideoMAE model pre-trained on Kinetics-400 for 800 epochs in a self-supervised way. It was introduced in the paper VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training by Tong et al. and first released in this repository. Disclaimer: The team releasing VideoMAE did not write a model card for this model so this model card has been written by the Hugging Face team. VideoMAE is an extension of Masked Autoencoders (MAE) to video. The architecture of the model is very similar to that of a standard Vision Transformer (ViT), with a decoder on top for predicting pixel values for masked patches. Videos are presented to the model as a sequence of fixed-size…
Open weights
cc-by-nc-4.0
94M parameters
transformers
Frozen V-JEPA 2.1 tokens on LED, tokenised as k frames (histogram) (16 time bins of 1.25 ms paired into 8 tubelet tokens per cell), scored per event: each event looks up the token at its (time bin / 2, patch) and is judged with its own E15 neighbourhood patch. Routes in this repo: Pixel-AC fuses the per-event scores into one decision per pixel and chunk (integrated/ /; result.json carries the Pixel-AC metrics with the AC-only metrics under ac). Run-name suffixes: nokv = reader bypassed, lb = scalar Local branch on, vitbase = ViT-B/384 tokens. All units: 6,000 training chunks, 4 epochs, official 713 chunks at a fixed threshold of 0.5, no selection on test. Reference on the same protocol: the…
Open weights
cc-by-nc-4.0
pytorch
The whole counsel of Scripture — to read, search and study, in the original Hebrew, Greek and Aramaic, on any device, offline. Free for life. Grab it first, then keep reading while it downloads. It's free — and it stays free. yes · requires payment or subscription · no. Only YahBible is fully offline — including its AI search and reasoning — and stays free with no trial or subscription. Other apps' names belong to their owners. The King James Bible with 120+ translations, verse for verse. Search however you remember a passage — "3:16 John", "the 23rd Psalm" — and switch versions without losing your place. Open any commandment in the centre: its Scripture, what it forbids, how to keep it…
Open weights
cc-by-nc-4.0
Four artist-style LoRAs that push YuE2-3B into modern militant roots reggae: dark raspy male patois vocals, steppers and one-drop grooves, deep sub bass, bubbling Hammond, nyabinghi drums, horn stabs, dub sirens and spring reverb. Conscious, apocalyptic, anthemic. Each file patches both halves of YuE2 in one go: the autoregressive planner (writes the score, decides the arrangement and the vocal lines) and the flow-matching decoder (the sound). Trigger word for all three: mltnt. All demos use the same original lyric, seed 7, 32 steps dpm2 / sgmuniform, no post-processing. MLTNT Frontline — baseline recipe, prompt prompts/steppersbaseline.txt, dense lyric (verses written at ~17 words per…
Open weights
cc-by-nc-4.0