SAVRN
Search Contact SAVRN

Open-weight model · Feature extraction

vit_giant_patch14_224.dinobloom

by Laureηt Fainsin 1aurent/vit_giant_patch14_224.dinobloom

vit_giant_patch14_224.dinobloom is an open-weight model for feature extraction from Laureηt Fainsin, released under Apache License 2.0. It has 1.1B parameters. At 16-bit it needs about 2.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 56 downloads a month.

Model Type: Feature backbone; Params: 1136M (giant); Image size: 224 x 224 x 3; Patch size: 14 x 14 x 3; Repository: github.com:marrlab/DinoBloom; Original Weights: Zenodo.

Parameters1.1B
Context—
Weights4.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads56

Runs On

What it takes to serve vit_giant_patch14_224.dinobloom (1.1B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 2.3 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.1 GB 1.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.6 GB 0.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

vit_giant_patch14_224.dinobloom on every accelerator the SAVRN Index prices, at every precision

Model Card

By Laureηt Fainsin, published under apache-2.0, revision a12c5d7531bb.

Model Type: Feature backbone; Params: 1136M (giant); Image size: 224 x 224 x 3; Patch size: 14 x 14 x 3; Repository: github.com:marrlab/DinoBloom; Original Weights: Zenodo.

Read Laureηt Fainsin's full model card

[!WARNING]
Please use https://huggingface.co/MarrLab/DinoBloom instead

Model card for vit_giant_patch14_224.dinobloom

Model Details

Model Usage

Image Embeddings

from urllib.request import urlopen
from PIL import Image
import timm

# get example histology image
img = Image.open(
  urlopen(
    "https://raw.githubusercontent.com/zxaoyou/segmentation_WBC/master/Dataset%201/001.bmp"
  )
)

# load model from the hub
model = timm.create_model(
  model_name="hf-hub:1aurent/vit_giant_patch14_224.dinobloom",
  pretrained=True,
).eval()

# get model specific transforms (normalization, resize)
data_config = timm.data.resolve_model_data_config(model)
transforms = timm.data.create_transform(**data_config, is_training=False)

data = transforms(img).unsqueeze(0) # input is a (batch_size, num_channels, img_size, img_size) shaped tensor
output = model(data)  # output is a (batch_size, num_features) shaped tensor

Citation

@misc{koch2024dinobloom,
  title         = {DinoBloom: A Foundation Model for Generalizable Cell Embeddings in Hematology}, 
  author        = {Valentin Koch and Sophia J. Wagner and Salome Kazeminia and Ece Sancar and Matthias Hehr and Julia Schnabel and Tingying Peng and Carsten Marr},
  year          = {2024},
  eprint        = {2404.05022},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV}
}

Identity and Version

Repository
1aurent/vit_giant_patch14_224.dinobloom
Publisher
Laureηt Fainsin
Task
Feature extraction
Modality
Text
Library
timm
Parameters
1.1B parameters
Languages
Not stated by the source
Revision
a12c5d7531bbc2547f216e71b334acf0222c0643
First published
2024-05-20
Last updated
2026-10-05

Files and Weights

4 files, 4.5 GB in total. The weights are 1 file totalling 4.5 GB in safetensors.

Weights1 file · 4.5 GB
Configuration1 file · 614 B
Documentation1 file · 2.1 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights4.5 GB 6cb92046492f
config.jsonConfiguration614 B —
README.mdDocumentation2.1 KB —
.gitattributesRepository1.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
4.5 GB
Download from Laureηt Fainsin

Released by Laureηt Fainsin through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2404.05022

Memory Requirements

PrecisionWeights in memory
As published4.5 GB
16-bit2.3 GB
8-bit1.1 GB
4-bit0.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About vit_giant_patch14_224.dinobloom

How much GPU memory does vit_giant_patch14_224.dinobloom need?

About 2.7 GB at 16-bit and 0.7 GB at 4-bit: the weights (1.1B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run vit_giant_patch14_224.dinobloom on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use vit_giant_patch14_224.dinobloom commercially?

Yes. vit_giant_patch14_224.dinobloom is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Feature extraction

Vela-1.0-Omni-Mini

vLLM Semantic Router

Text, images, speech and environmental sounds in one embedding space. Vela Omni Mini supports multimodal search, routing and clustering with normalized vectors that can be compared directly. Scores are 0–100; higher is better. The comparison uses the same held-out evaluation examples, complete retrieval pools and 128-token text cap for both models. Bold marks improvement over the original large model. These known test pools are reused across releases. The common protocol uses labeled TRAIN prototypes for text classification and all matching positives for retrieval; it is separate from official MTEB classification. All 14 metrics, Macro-F1, exact counts and uncertainty. The primary metric…

Open weights apache-2.0 1.4B parameters pytorch

Model · Feature extraction

voxtral-mini-audio-extractor

Vxltxr Llc

A lightweight standalone audio feature extractor derived from Voxtral-Mini-3B-2507. This repository isolates the Whisper-based audio encoder and multi-modal projector from the original 3B language model. The extracted module is intended for offline audio preprocessing, dataset preparation, feature caching, and downstream multimodal training pipelines. The language-model decoder is not included. It does not contain the 3B LLaMA language-model decoder. For offline preprocessing, loading the complete multimodal language model is unnecessary when the only required output is the projected audio representation. This standalone checkpoint can therefore be used as a dedicated audio feature…

Open weights apache-2.0 662M parameters

Model · Feature extraction

Qwen3-Embedding-0.6B

Qwen

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining. Exceptional Versatility: The…

Open weights apache-2.0 596M parameters 32,768 tokens sentence-transformers

Model · Feature extraction

w2v-bert-2.0

AI at Meta

We are open-sourcing our Conformer-based W2v-BERT 2.0 speech encoder as described in Section 3.2.1 of the paper, which is at the core of our Seamless models. This model was pre-trained on 4.5M hours of unlabeled audio data covering more than 143 languages. It requires finetuning to be used for downstream tasks such as Automatic Speech Recognition (ASR), or Audio Classification. This model and its training are supported by Transformers, more on it in the docs. This is a bare checkpoint without any modeling head, and thus requires finetuning to be used for downstream tasks such as ASR. You can however use it to extract audio embeddings from the top layer with this code snippet: To learn more…

Open weights mit 580M parameters transformers

Model · Feature extraction

jina-embeddings-v3

Jina AI

jina-embeddings-v3 is a multilingual multi-task text embedding model designed for a variety of NLP applications. Based on the Jina-XLM-RoBERTa architecture, this model supports Rotary Position Embeddings to handle long input sequences up to 8192 tokens. Additionally, it features 5 LoRA adapters to generate task-specific embeddings efficiently. - retrieval.query: Used for query embeddings in asymmetric retrieval tasks - retrieval.passage: Used for passage embeddings in asymmetric retrieval tasks - separation: Used for embeddings in clustering and re-ranking applications - classification: Used for embeddings in classification tasks - text-matching: Used for embeddings in tasks that quantify…

Open weights cc-by-nc-4.0 572M parameters 8,194 tokens transformers

We have updated the new reranker, supporting larger lengths, more languages, and achieving better performance. More details please refer to our Github: FlagEmbedding. FlagEmbedding focuses on retrieval-augmented LLMs, consisting of the following projects currently: - 3/18/2024: Release new rerankers, built upon powerful M3 and LLM (GEMMA and MiniCPM, not so large actually) backbones, supporitng multi-lingual processing and larger inputs, massive improvements of ranking performances on BEIR, C-MTEB/Retrieval, MIRACL, LlamaIndex Evaluation. - 3/18/2024: Release Visualized-BGE, equipping BGE with visual capabilities. Visualized-BGE can be utilized to generate embeddings for hybrid image-text…

Open weights mit 560M parameters 514 tokens transformers