SAVRN
Search Contact SAVRN

Open-weight model · Image feature extraction

dinov2-base

by AI at Meta facebook/dinov2-base

Vision Transformer (ViT) model trained using the DINOv2 method. It was introduced in the paper DINOv2: Learning Robust Visual Features without Supervision by Oquab et al. and first released in this repository.

Parameters87M
Context
Weights692.7 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads3.4M

Runs On

What it takes to serve dinov2-base (87M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on dinov2-base

Nothing about this one strains a card. The 16-bit weights load in 0.2 GB, and the cheapest option on the page is one MI300X with 192 GB at $1.85 an hour, so share the card; the image batch sets the bill, not the encoder. What comes out is a feature vector, not an answer, which puts its 87M parameters and 768-wide hidden state ahead of your own classifier or retrieval index.

Nothing in the license slows a deployment: Apache 2.0 covers commercial use, modification and redistribution as long as the notices stay attached, so a fine-tuned head ships with the encoder in a product. Two things to confirm: the page lists no reported evaluations, so your own held-out images are the benchmark, and the publisher's team did not write the model card, so the paper it cites, arXiv:2304.07193, is the record of training. Weights last updated January 17, 2024.

Model Card

By AI at Meta, published under apache-2.0, revision f9e44c814b77.

Vision Transformer (base-sized model) trained using DINOv2

Vision Transformer (ViT) model trained using the DINOv2 method. It was introduced in the paper DINOv2: Learning Robust Visual Features without Supervision by Oquab et al. and first released in this repository.

Disclaimer: The team releasing DINOv2 did not write a model card for this model so this model card has been written by the Hugging Face team.

Model description

The Vision Transformer (ViT) is a transformer encoder model (BERT-like) pretrained on a large collection of images in a self-supervised fashion.

Images are presented to the model as a sequence of fixed-size patches, which are linearly embedded. One also adds a [CLS] token to the beginning of a sequence to use it for classification tasks. One also adds absolute position embeddings before feeding the sequence to the layers of the Transformer encoder.

Note that this model does not include any fine-tuned heads.

Read the full model card (397 words)

Configuration

Architecture
Dinov2Model
Layers
12
Hidden size
768
Attention heads
12
Stored precision
float32
Model type
dinov2

Identity and Version

Repository
facebook/dinov2-base
Publisher
AI at Meta
Task
Image feature extraction
Modality
Other
Library
transformers
Parameters
87M parameters
Languages
Not stated by the source
Revision
f9e44c814b77203eaa57a6bdbbd535f21ede1415
First published
2023-07-17
Last updated
2024-01-17

Files and Weights

6 files, 692.7 MB in total. The weights are 2 files totalling 692.7 MB in bin, safetensors.

Weights2 files · 692.7 MB
Configuration2 files · 984 B
Documentation1 file · 3.0 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights346.3 MB d73036b56966
pytorch_model.binWeights346.4 MB 014965d9e330
config.jsonConfiguration548 B
preprocessor_config.jsonConfiguration436 B
README.mdDocumentation3.0 KB
.gitattributesRepository1.5 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
692.7 MB
Download from AI at Meta

Released by AI at Meta through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published692.7 MB
16-bit0.2 GB
8-bit0.1 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare dinov2-base

Questions About dinov2-base

How much GPU memory does dinov2-base need?

About 0.2 GB at 16-bit and 0.1 GB at 4-bit: the weights (87M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run dinov2-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use dinov2-base commercially?

Yes. dinov2-base is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image feature extraction

vit-base-patch16-224-in21k

Google

Vision Transformer (ViT) model pre-trained on ImageNet-21k (14 million images, 21,843 classes) at resolution 224x224. It was introduced in the paper An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale by Dosovitskiy et al. and first released in this repository. However, the weights were converted from the timm repository by Ross Wightman, who already converted the weights from JAX to PyTorch. Credits go to him. Disclaimer: The team releasing ViT did not write a model card for this model so this model card has been written by the Hugging Face team. The Vision Transformer (ViT) is a transformer encoder model (BERT-like) pretrained on a large collection of images in a…

Open weights apache-2.0 86M parameters transformers

Model · Image feature extraction

dinov2-small

AI at Meta

Vision Transformer (ViT) model trained using the DINOv2 method. It was introduced in the paper DINOv2: Learning Robust Visual Features without Supervision by Oquab et al. and first released in this repository. Disclaimer: The team releasing DINOv2 did not write a model card for this model so this model card has been written by the Hugging Face team. The Vision Transformer (ViT) is a transformer encoder model (BERT-like) pretrained on a large collection of images in a self-supervised fashion. Images are presented to the model as a sequence of fixed-size patches, which are linearly embedded. One also adds a [CLS] token to the beginning of a sequence to use it for classification tasks. One…

Open weights apache-2.0 22M parameters transformers

Model · Image feature extraction

vit_small_patch14_dinov2.lvd142m

PyTorch Image Models

A Vision Transformer (ViT) image feature model. Pretrained on LVD-142M with self-supervised DINOv2 method. - An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: https://arxiv.org/abs/2010.11929v2 Explore the dataset and runtime metrics of this model in timm model results.

Open weights apache-2.0 22M parameters timm

Model · Image feature extraction

dinov3-vitl16-pretrain-lvd1689m

Camenduru

DINOv3 is a family of versatile vision foundation models that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models. These are Vision Transformer and ConvNeXt models trained following the method described in the DINOv3 paper. 12 models are provided: - 10 models pretrained on web data (LVD-1689M dataset) - 1 ViT-7B trained from scratch, - 5 ViT-S/S+/B/L/H+ models distilled from the ViT-7B, - 4 ConvNeXt-{T/S/B/L} models distilled from the ViT-7B, - 2 models…

Open weights other 303M parameters transformers

Model · Image feature extraction

dinov2-large

AI at Meta

Vision Transformer (ViT) model trained using the DINOv2 method. It was introduced in the paper DINOv2: Learning Robust Visual Features without Supervision by Oquab et al. and first released in this repository. Disclaimer: The team releasing DINOv2 did not write a model card for this model so this model card has been written by the Hugging Face team. The Vision Transformer (ViT) is a transformer encoder model (BERT-like) pretrained on a large collection of images in a self-supervised fashion. Images are presented to the model as a sequence of fixed-size patches, which are linearly embedded. One also adds a [CLS] token to the beginning of a sequence to use it for classification tasks. One…

Open weights apache-2.0 304M parameters transformers