SAVRN
Search Contact SAVRN

Open-weight model · Zero shot object detection

grounding-dino-base

by IDEA-Research IDEA-Research/grounding-dino-base

The Grounding DINO model was proposed in Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection by Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang.

Parameters233M
Context
Weights1.9 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.5M

Runs On

What it takes to serve grounding-dino-base (233M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.5 GB 0.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on grounding-dino-base

Ask a camera for the pallet jack near the loading door without training a pallet-jack detector, and you are describing what grounding-dino-base does. IDEA-Research attached a text encoder to a closed-set detector, so the classes come from the prompt instead of a fixed label list. It carries 233 million parameters stored in float32, yet at 16-bit it needs 0.6 GB, 0.3 GB at 8-bit and 0.1 GB at 4-bit. Against the cheapest Index listing, one MI300X with 192 GB at $1.85 an hour on-demand, that is a rounding error, so plan for many copies per card.

Apache 2.0 with commercial use permitted lets you ship it inside a product, modify it and redistribute it, provided the license and NOTICE files stay and significant changes are stated. No context length to size; the prompt is a short string. Released September 25, 2023, last updated May 12, 2024, method in arXiv:2303.05499.

Model Card

By IDEA-Research, published under apache-2.0, revision 12bdfa3120f3.

Grounding DINO model (base variant)

The Grounding DINO model was proposed in Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection by Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang. Grounding DINO extends a closed-set object detection model with a text encoder, enabling open-set object detection. The model achieves remarkable results, such as 52.5 AP on COCO zero-shot.

Grounding DINO overview. Taken from the original paper.

Intended uses & limitations

You can use the raw model for zero-shot object detection (the task of detecting things in an image out-of-the-box without labeled data).

How to use

Here's how to use the model for zero-shot object detection:

Read the full model card (259 words)

Configuration

Architecture
GroundingDinoForObjectDetection
Stored precision
float32
Model type
grounding-dino

Identity and Version

Repository
IDEA-Research/grounding-dino-base
Publisher
IDEA-Research
Task
Zero shot object detection
Modality
Other
Library
transformers
Parameters
233M parameters
Languages
Not stated by the source
Revision
12bdfa3120f3e7ec7b434d90674b3396eccf88eb
First published
2023-09-25
Last updated
2024-05-12

Files and Weights

10 files, 1.9 GB in total. The weights are 2 files totalling 1.9 GB in bin, safetensors.

Weights2 files · 1.9 GB
Configuration3 files · 2.3 KB
Tokenizer3 files · 944.1 KB
Documentation1 file · 2.6 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights933.4 MB 5548f844c928
pytorch_model.binWeights936.0 MB 4b51daf29695
config.jsonConfiguration1.7 KB
preprocessor_config.jsonConfiguration457 B
special_tokens_map.jsonConfiguration125 B
README.mdDocumentation2.6 KB
.gitattributesRepository1.5 KB
tokenizer.jsonTokenizer711.4 KB
tokenizer_config.jsonTokenizer1.2 KB
vocab.txtTokenizer231.5 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.9 GB
Download from IDEA-Research

Released by IDEA-Research through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published1.9 GB
16-bit0.5 GB
8-bit0.2 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About grounding-dino-base

How much GPU memory does grounding-dino-base need?

About 0.6 GB at 16-bit and 0.1 GB at 4-bit: the weights (233M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run grounding-dino-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use grounding-dino-base commercially?

Yes. grounding-dino-base is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Zero shot object detection

grounding-dino-tiny

IDEA-Research

The Grounding DINO model was proposed in Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection by Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang. Grounding DINO extends a closed-set object detection model with a text encoder, enabling open-set object detection. The model achieves remarkable results, such as 52.5 AP on COCO zero-shot. alt="drawing" width="600"/> You can use the raw model for zero-shot object detection (the task of detecting things in an image out-of-the-box without labeled data). Here's how to use the model for zero-shot object detection

Open weights apache-2.0 172M parameters transformers

Model · Zero shot object detection

owlv2-base-patch16-ensemble

Google

The OWLv2 model (short for Open-World Localization) was proposed in Scaling Open-Vocabulary Object Detection by Matthias Minderer, Alexey Gritsenko, Neil Houlsby. OWLv2, like OWL-ViT, is a zero-shot text-conditioned object detection model that can be used to query an image with one or multiple text queries. The model uses CLIP as its multi-modal backbone, with a ViT-like Transformer to get visual features and a causal language model to get the text features. To use CLIP for detection, OWL-ViT removes the final token pooling layer of the vision model and attaches a lightweight classification and box head to each transformer output token. Open-vocabulary classification is enabled by replacing…

Open weights apache-2.0 155M parameters transformers