SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

vllm-translategemma-4b-it

by Infomaniak Network SA Infomaniak-AI/vllm-translategemma-4b-it

This is a modified version of google/translategemma-4b-it optimized for deployment with vLLM.

Parameters5B
Context131,072
Weights8.6 GB
Licensegemma
AccessOpen weights
Monthly Downloads743.2k

Runs On

What it takes to serve vllm-translategemma-4b-it (5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 9.9 GB 11.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 5.0 GB 6.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.5 GB 3.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on vllm-translategemma-4b-it

Infomaniak Network SA published this, not Google. It is google/translategemma-4b-it with the chat template rewritten so vLLM can take the source and target language codes inline in the message, marked by double arrows, rather than in separate fields. Image and text in, text out, 5 billion parameters, a 131,072-token window with a 1,024-token sliding window. The 16-bit weights are 9.9 GB and need 11.9 GB; 8-bit needs 6.0 GB and 4-bit needs 3.0 GB. The cheapest fit we list is one MI300X with 192 GB at $1.85 an hour.

The Gemma Terms of Use govern it whoever re-hosts the weights: commercial use is allowed provided the Prohibited Use Policy travels with every copy you pass on, so redistribution carries paperwork that an internal deployment does not. Confirm your serving stack honors the modified template, and know that our file holds no evaluations for it and no host prices.

Model Card

This is a modified version of google/translategemma-4b-it optimized for deployment with vLLM. The original TranslateGemma model requires a structured payload with dedicated sourcelangcode and targetlangcode fields: However, vLLM does not support these custom content parameters. To maintain compatibility, the chat template has been modified to encode language codes directly in the message content using a delimiter-based format: Format: >>{sourcelang} >>{targetlang} >>{texttotranslate} If you need to provide a custom prompt input The original model uses the new Transformers RoPE configuration format with separate attention type settings: This has been simplified for vLLM compatibility: The…

Excerpt from the card by Infomaniak Network SA, licensed gemma.

Configuration

Architecture
Gemma3ForConditionalGeneration
Context length (tokens)
131,072
Layers
34
Hidden size
2,560
Feed-forward size
10,240
Attention heads
8
Key/value heads
4
Head dimension
256
Vocabulary size
262,208
Sliding window (tokens)
1,024
RoPE base
1,000,000
Model type
gemma3

Identity and Version

Repository
Infomaniak-AI/vllm-translategemma-4b-it
Publisher
Infomaniak Network SA
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
5B parameters
Languages
Not stated by the source
Revision
cb3e0b2504f0db832982cca7fafd3ea4aaf777a0
First published
2026-01-26
Last updated
2026-01-27

Files and Weights

16 files, 8.6 GB in total. The weights are 2 files totalling 8.6 GB in safetensors.

Weights2 files · 8.6 GB
Configuration8 files · 104.8 KB
Tokenizer3 files · 39.2 MB
Documentation1 file · 19.3 KB
Other1 file · 17.3 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights5.0 GB 6187566ec8d0
model-00002-of-00002.safetensorsWeights3.6 GB 40c6a896b227
added_tokens.jsonConfiguration35 B
config.jsonConfiguration2.5 KB
generation_config.jsonConfiguration229 B
model.safetensors.index.jsonConfiguration90.6 KB
preprocessor_config.jsonConfiguration570 B
processor_config.jsonConfiguration70 B
special_tokens_map.jsonConfiguration662 B
test_chat_template.pyConfiguration10.2 KB
README.mdDocumentation19.3 KB
chat_template.jinjaOther17.3 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer33.4 MB 7d4046bf0505
tokenizer.modelTokenizer4.7 MB 1299c11d7cf6
tokenizer_config.jsonTokenizer1.2 MB

License and Download

License
gemma
Access
Open weights, no gate
Download size
8.6 GB
Download from Infomaniak Network SA

Released by Infomaniak Network SA through its official repository on Hugging Face.

Built From

  • Described by arXiv:2503.19786
  • Described by arXiv:2601.09012

Memory Requirements

PrecisionWeights in memory
As published8.6 GB
16-bit9.9 GB
8-bit5.0 GB
4-bit2.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare vllm-translategemma-4b-it

Questions About vllm-translategemma-4b-it

How much GPU memory does vllm-translategemma-4b-it need?

About 11.9 GB at 16-bit and 3 GB at 4-bit: the weights (5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run vllm-translategemma-4b-it on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use vllm-translategemma-4b-it commercially?

Yes, with conditions. vllm-translategemma-4b-it is released under Gemma Terms of Use. Gemma models are released under Google's Gemma Terms of Use, which permit commercial use and redistribution subject to the Gemma Prohibited Use Policy, whose restrictions must be passed on to anyone the model is distributed to.

What is vllm-translategemma-4b-it's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.5-4B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 4.7B parameters 262,144 tokens transformers

Model · Image and text to text

chandra-ocr-2

Datalab

Chandra 2 is a state of the art OCR model from Datalab that outputs markdown, HTML, and JSON. It is highly accurate at extracting text from images and PDFs, while preserving layout information. Try Chandra in the free playground, or use the hosted API for higher accuracy and speed. - 85.8% olmocr bench score (sota), 77.8% multilingual bench score (12% improvement over Chandra 1) - Significant improvements to math, tables, complex layouts - 90+ language support with major accuracy gains - Convert documents to markdown, HTML, or JSON with detailed layout information - Reconstructs forms accurately, including checkboxes - Strong performance with tables, math, and complex layouts - Extracts…

Open weights openrail 5.3B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next-4bit-paged

GREENBITAI

Expert-paged build of Vontra/Qwen3.8-Flash-Next-MLX-4bit. The weights that are read a fraction at a time live in their own containers, so a machine loads what it needs rather than all Total 105.46 GiB. Of that, 103.94 GiB is the source build, whose bytes moved into containers rather than being copied, and 1.52 GiB is the draft head, which no published build of this model carries. Where the weights fit they are filled from experts.bin and the model runs the stock path at stock speed; where they do not, they stream from disk. Reading the machine decides that, not a flag. To override that: GBXPAGING=off holds the experts resident, GBXPLE=off holds the n-gram table resident. Checked at build…

Open weights other 5.4B parameters 262,144 tokens mlx

Model · Image and text to text

Omni-Edu-4B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3-VL-4B-Instruct

Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension. 1. Interleaved-MRoPE: Full‑frequency allocation over time, width, and height…

Open weights apache-2.0 4.4B parameters 262,144 tokens transformers

Model · Image and text to text

gemma-3-4b-it

Google

[Gemma 3 Technical Report][g3-tech-report] [Responsible Generative AI Toolkit][rai-toolkit] [Gemma on Kaggle][kaggle-gemma] [Gemma on Vertex Model Garden][vertex-mg-gemma3] Summary description and brief definition of inputs and outputs. Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned variants. Gemma 3 has a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous…

Access requested at publisher gemma 4.3B parameters transformers