SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

gemma-3-27b-it-GPTQ-4b-128g

by IST Austria Distributed Algorithms and Systems Lab ISTA-DASLab/gemma-3-27b-it-GPTQ-4b-128g

This model was obtained by quantizing the weights of gemma-3-27b-it to INT4 data type. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%.

Parameters27.6B
Context131,072
Weights16.9 GB
Licensegemma
AccessOpen weights
Monthly Downloads715.4k

Runs On

What it takes to serve gemma-3-27b-it-GPTQ-4b-128g (27.6B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 55.3 GB 66.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.6 GB 33.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.8 GB 16.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

This model was obtained by quantizing the weights of gemma-3-27b-it to INT4 data type. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%. Only the weights of the linear operators within languagemodel transformers blocks are quantized. Vision model and multimodal projection are kept in original precision. Weights are quantized using a symmetric per-group scheme, with group size 128. The GPTQ algorithm is applied for quantization. Model checkpoint is saved in compressedtensors format. This model was evaluated on the OpenLLM v1 benchmarks. Model outputs were generated with the vLLM engine. The…

Excerpt from the card by IST Austria Distributed Algorithms and Systems Lab, licensed gemma.

Configuration

Architecture
Gemma3ForConditionalGeneration
Context length (tokens)
131,072
Layers
62
Hidden size
5,376
Feed-forward size
21,504
Attention heads
32
Key/value heads
16
Head dimension
128
Vocabulary size
262,208
Sliding window (tokens)
1,024
RoPE base
1e+06
Stored precision
bfloat16
Model type
gemma3
Quantization
compressed-tensors

Identity and Version

Repository
ISTA-DASLab/gemma-3-27b-it-GPTQ-4b-128g
Publisher
IST Austria Distributed Algorithms and Systems Lab
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
27.6B parameters
Languages
Not stated by the source
Revision
9c5f61a23e1ccf3921173c2b32f6ff0af1bffd1b
First published
2025-03-14
Last updated
2025-03-20

Files and Weights

17 files, 16.9 GB in total. The weights are 5 files totalling 16.9 GB in safetensors.

Weights5 files · 16.9 GB
Configuration8 files · 215.5 KB
Tokenizer2 files · 5.8 MB
Documentation1 file · 3.7 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00005.safetensorsWeights3.4 GB 3ee3056de4d0
model-00002-of-00005.safetensorsWeights3.4 GB fc7d43f12cdd
model-00003-of-00005.safetensorsWeights3.4 GB cffdaddf3828
model-00004-of-00005.safetensorsWeights3.0 GB 8abfbbdc62e6
model-00005-of-00005.safetensorsWeights3.7 GB 67dd90abb5fb
added_tokens.jsonConfiguration35 B
chat_template.jsonConfiguration1.6 KB
config.jsonConfiguration2.4 KB
generation_config.jsonConfiguration215 B
model.safetensors.index.jsonConfiguration210.0 KB
preprocessor_config.jsonConfiguration570 B
processor_config.jsonConfiguration70 B
special_tokens_map.jsonConfiguration662 B
README.mdDocumentation3.7 KB
.gitattributesRepository1.5 KB
tokenizer.modelTokenizer4.7 MB 1299c11d7cf6
tokenizer_config.jsonTokenizer1.2 MB

License and Download

License
gemma
Access
Open weights, no gate
Download size
16.9 GB
Download from IST Austria Distributed Algorithms and Systems Lab

Released by IST Austria Distributed Algorithms and Systems Lab through its official repository on Hugging Face.

Built From

  • Derived from google/gemma-3-27b-it
  • Quantized from google/gemma-3-27b-it

Memory Requirements

PrecisionWeights in memory
As published16.9 GB
16-bit55.3 GB
8-bit27.6 GB
4-bit13.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About gemma-3-27b-it-GPTQ-4b-128g

How much GPU memory does gemma-3-27b-it-GPTQ-4b-128g need?

About 66.3 GB at 16-bit and 16.6 GB at 4-bit: the weights (27.6B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run gemma-3-27b-it-GPTQ-4b-128g on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use gemma-3-27b-it-GPTQ-4b-128g commercially?

Yes, with conditions. gemma-3-27b-it-GPTQ-4b-128g is released under Gemma Terms of Use. Gemma models are released under Google's Gemma Terms of Use, which permit commercial use and redistribution subject to the Gemma Prohibited Use Policy, whose restrictions must be passed on to anyone the model is distributed to.

What is gemma-3-27b-it-GPTQ-4b-128g's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.8-27B

Qwen

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-FP8

Qwen

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.6-27B

Qwen

Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. This release delivers substantial upgrades, particularly in For more details, please refer to our blog post Qwen3.6-27B. Empty cells (--) indicate scores not yet available or not applicable. For streamlined integration, we recommend using Qwen3.6 via APIs. Below is a guide to use Qwen3.6 via OpenAI-compatible API. Qwen3.6 can be served via APIs with popular inference frameworks. In…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-27B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-AWQ-INT4

Cyankiwi

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B

Kyle Thomas

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Qwen3.8 can be…

Open weights apache-2.0 27.8B parameters 262,144 tokens transformers