SAVRN
Search Contact SAVRN

Organization

Mobile Perception Systems Lab

tue-mps

Models in Library2
Datasets in Library0
Models on Hugging Face59
Followers35

Models

EoMT (Encoder-only Mask Transformer) is a Vision Transformer (ViT) architecture designed for high-quality and efficient image segmentation. It was introduced in the CVPR 2025 highlight paper: by Tommie Kerssies, Niccolò Cavagnero, Alexander Hermans, Narges Norouzi, Giuseppe Averta, Bastian Leibe, Gijs Dubbelman, and Daan de Geus. The original implementation can be found in this repository. The HuggingFace model page is available at this link. Here is how to use this model for Panotpic Segmentation: If you find our work useful, please consider citing us as

Open weights mit 317M parameters transformers

This repository contains the Hugging Face Transformers conversion of the official VidEoMT checkpoint yt2019vitsmall52.8.pth from tue-mps/VidEoMT. The metrics above are the numbers reported by the authors in the official model zoo. Use processor.postprocessinstancesegmentation, processor.postprocesspanopticsegmentation, or processor.postprocesssemanticsegmentation depending on the target task.

Open weights 24M parameters transformers