SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

gemma-3-27b-it-int4-awq

by Thien Tran gaunernst/gemma-3-27b-it-int4-awq

This is the QAT INT4 Flax checkpoint (from Kaggle) converted to HF+AWQ format for ease of use. AWQ was NOT used for quantization. You can find the conversion script convertflax.py in this model repo.

Parameters27.4B
Context
Weights18.5 GB
Licensegemma
AccessOpen weights
Monthly Downloads1.3M

Runs On

What it takes to serve gemma-3-27b-it-int4-awq (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 54.9 GB 65.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 27.4 GB 32.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 13.7 GB 16.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

This is the QAT INT4 Flax checkpoint (from Kaggle) converted to HF+AWQ format for ease of use. AWQ was NOT used for quantization. You can find the conversion script convertflax.py in this model repo. NOTE: this is NOT the same as the official QAT INT4 GGUFs released here https://huggingface.co/collections/google/gemma-3-qat-67ee61ccacbf2be4195c265b Below is the original Model card from https://huggingface.co/google/gemma-3-27b-it [Gemma 3 Technical Report][g3-tech-report] [Responsible Generative AI Toolkit][rai-toolkit] [Gemma on Kaggle][kaggle-gemma] [Gemma on Vertex Model Garden][vertex-mg-gemma3] Summary description and brief definition of inputs and outputs. Gemma is a family of…

Excerpt from the card by Thien Tran, licensed gemma.

Configuration

Architecture
Gemma3ForConditionalGeneration
Layers
62
Hidden size
5,376
Feed-forward size
21,504
Attention heads
32
Key/value heads
16
Head dimension
128
Sliding window (tokens)
1,024
Stored precision
bfloat16
Model type
gemma3
Quantization
awq

Identity and Version

Repository
gaunernst/gemma-3-27b-it-int4-awq
Publisher
Thien Tran
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
27.4B parameters
Languages
awq
Revision
7cf8bdc81343c635390dd1e1bf590ab22dd6f366
First published
2025-03-21
Last updated
2025-04-06

Files and Weights

18 files, 18.5 GB in total. The weights are 4 files totalling 18.5 GB in safetensors.

Weights4 files · 18.5 GB
Configuration9 files · 234.4 KB
Tokenizer3 files · 39.2 MB
Documentation1 file · 25.6 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00004.safetensorsWeights5.0 GB dfe46693f149
model-00002-of-00004.safetensorsWeights4.8 GB b7e265b15839
model-00003-of-00004.safetensorsWeights4.8 GB fc7283de8174
model-00004-of-00004.safetensorsWeights3.9 GB 0cff33cb628f
added_tokens.jsonConfiguration35 B
chat_template.jsonConfiguration1.6 KB
config.jsonConfiguration1.2 KB
convert_flax.pyConfiguration10.9 KB
generation_config.jsonConfiguration215 B
model.safetensors.index.jsonConfiguration219.1 KB
preprocessor_config.jsonConfiguration570 B
processor_config.jsonConfiguration70 B
special_tokens_map.jsonConfiguration662 B
README.mdDocumentation25.6 KB
.gitattributesRepository1.6 KB
tokenizer.jsonTokenizer33.4 MB 4667f2089529
tokenizer.modelTokenizer4.7 MB 1299c11d7cf6
tokenizer_config.jsonTokenizer1.2 MB

License and Download

License
gemma
Access
Open weights, no gate
Download size
18.5 GB
Download from Thien Tran

Released by Thien Tran through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published18.5 GB
16-bit54.9 GB
8-bit27.4 GB
4-bit13.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About gemma-3-27b-it-int4-awq

How much GPU memory does gemma-3-27b-it-int4-awq need?

About 65.8 GB at 16-bit and 16.5 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run gemma-3-27b-it-int4-awq on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use gemma-3-27b-it-int4-awq commercially?

Yes, with conditions. gemma-3-27b-it-int4-awq is released under Gemma Terms of Use. Gemma models are released under Google's Gemma Terms of Use, which permit commercial use and redistribution subject to the Gemma Prohibited Use Policy, whose restrictions must be passed on to anyone the model is distributed to.

Similar Models

An INT8 W8A8 quantization of (the abliterated BF16 build of a merged LoRA finetune of Qwen/Qwen3.8-27B), for fast serving on GPUs where native FP8 is unavailable or undesirable. int-quantized, native CUTLASS INT8 tensor-core path. embedtokens, the vision tower, GatedDeltaNet recurrent gates, all norms. Note these are FP16, not the source's BF16 — the 16-bit residual precision of this build is float16 throughout (config.json: "dtype": "float16"). (llm-compressor), with the MTP drafter ablated in the rotated basis. The MTP head is unquantized but is not byte-identical to the BF16 build's: it carries the same residual-stream rotation and RMSNorm fold as the rest of the model, so the two heads…

Open weights other 27.4B parameters 262,144 tokens vllm

Model · Image and text to text

Qwen3.8-27B-MLX-4bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-8bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-6bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-MLX-5bit

LM Studio Community

LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-27B-Continuum-mxfp4-mlx

Gheorghe Chesler

(He slides three glasses across the bar—wine for Shakespeare, water for Data, and a mysterious blue liquid for Spock) This model is a merge of: - nightmedia/Qwen3.8-27B-Brainwaves - migtissera/Synthia-4-27B Brainwaves nightmedia/Qwen3.8-27B-Brainwaves migtissera/Synthia-4-27B Late-Night Architectural Takeaways The Quantization Shield (mxfp8 Master Pass): Hitting 0.735 ARC-C on the 8-bit layout proves that your Cold-Fusion flagship anchors and Migel Tissera's Synthia agent paths reached absolute geometric equilibrium. Instead of losing performance to the 0.596 Heretic collapse, the curved hypersphere calculation completely shielded the model's core intelligence. The Perplexity Sweet Spot…

Open weights apache-2.0 27.4B parameters 262,144 tokens transformers