SAVRN
Search Contact SAVRN

SAVRN Model Hub · Models by Task

Zero Shot Object Detection Models

3 open-weight zero shot object detection models in the SAVRN Model Hub, with IDEA-Research and Google publishing the most.

3Models
2Publishers
155M to 233MParameter range
1Licenses

Most Downloaded

ModelPublisherParametersLicenseMonthly downloadsCheapest GPUs at 16-bit
owlv2-base-patch16-ensemble Google 155M apache-2.0 1.7M 1x MI300X, $1.85/hr
grounding-dino-base IDEA-Research 233M apache-2.0 1.5M 1x MI300X, $1.85/hr
grounding-dino-tiny IDEA-Research 172M apache-2.0 761.2k 1x MI300X, $1.85/hr

Licenses

LicenseModelsCommercial use
apache-2.03Yes

Who Publishes Them

PublisherModels
IDEA-Research2
Google1

All 3 Models

Model · Zero shot object detection

owlv2-base-patch16-ensemble

Google

The OWLv2 model (short for Open-World Localization) was proposed in Scaling Open-Vocabulary Object Detection by Matthias Minderer, Alexey Gritsenko, Neil Houlsby. OWLv2, like OWL-ViT, is a zero-shot text-conditioned object detection model that can be used to query an image with one or multiple text queries. The model uses CLIP as its multi-modal backbone, with a ViT-like Transformer to get visual features and a causal language model to get the text features. To use CLIP for detection, OWL-ViT removes the final token pooling layer of the vision model and attaches a lightweight classification and box head to each transformer output token. Open-vocabulary classification is enabled by replacing…

Open weights apache-2.0 155M parameters transformers

Model · Zero shot object detection

grounding-dino-base

IDEA-Research

The Grounding DINO model was proposed in Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection by Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang. Grounding DINO extends a closed-set object detection model with a text encoder, enabling open-set object detection. The model achieves remarkable results, such as 52.5 AP on COCO zero-shot. alt="drawing" width="600"/> You can use the raw model for zero-shot object detection (the task of detecting things in an image out-of-the-box without labeled data). Here's how to use the model for zero-shot object detection

Open weights apache-2.0 233M parameters transformers

Model · Zero shot object detection

grounding-dino-tiny

IDEA-Research

The Grounding DINO model was proposed in Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection by Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang. Grounding DINO extends a closed-set object detection model with a text encoder, enabling open-set object detection. The model achieves remarkable results, such as 52.5 AP on COCO zero-shot. alt="drawing" width="600"/> You can use the raw model for zero-shot object detection (the task of detecting things in an image out-of-the-box without labeled data). Here's how to use the model for zero-shot object detection

Open weights apache-2.0 172M parameters transformers

Questions

Which Zero shot object detection models are most downloaded?

By monthly downloads reported by the Hugging Face Hub: owlv2-base-patch16-ensemble (1.7M); grounding-dino-base (1.5M); grounding-dino-tiny (761.2k).

Other Tasks

See all