An INT8 W8A8 quantization of (the abliterated BF16 build of a merged LoRA finetune of Qwen/Qwen3.8-27B), for fast serving on GPUs where native FP8 is unavailable or undesirable. int-quantized, native CUTLASS INT8 tensor-core path. embedtokens, the vision tower, GatedDeltaNet recurrent gates, all norms. Note these are FP16, not the source's BF16 — the 16-bit residual precision of this build is float16 throughout (config.json: "dtype": "float16"). (llm-compressor), with the MTP drafter ablated in the rotated basis. The MTP head is unquantized but is not byte-identical to the BF16 build's: it carries the same residual-stream rotation and RMSNorm fold as the rest of the model, so the two heads…
Open-weight model · Image and text to text
gemma-3-27b-it-int4-awq
by Thien Tran gaunernst/gemma-3-27b-it-int4-awq
This is the QAT INT4 Flax checkpoint (from Kaggle) converted to HF+AWQ format for ease of use. AWQ was NOT used for quantization. You can find the conversion script convertflax.py in this model repo.
Runs On
What it takes to serve gemma-3-27b-it-int4-awq (27.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 54.9 GB | 65.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 27.4 GB | 32.9 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 13.7 GB | 16.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
This is the QAT INT4 Flax checkpoint (from Kaggle) converted to HF+AWQ format for ease of use. AWQ was NOT used for quantization. You can find the conversion script convertflax.py in this model repo. NOTE: this is NOT the same as the official QAT INT4 GGUFs released here https://huggingface.co/collections/google/gemma-3-qat-67ee61ccacbf2be4195c265b Below is the original Model card from https://huggingface.co/google/gemma-3-27b-it [Gemma 3 Technical Report][g3-tech-report] [Responsible Generative AI Toolkit][rai-toolkit] [Gemma on Kaggle][kaggle-gemma] [Gemma on Vertex Model Garden][vertex-mg-gemma3] Summary description and brief definition of inputs and outputs. Gemma is a family of…
Excerpt from the card by Thien Tran, licensed gemma.
Configuration
- Architecture
- Gemma3ForConditionalGeneration
- Layers
- 62
- Hidden size
- 5,376
- Feed-forward size
- 21,504
- Attention heads
- 32
- Key/value heads
- 16
- Head dimension
- 128
- Sliding window (tokens)
- 1,024
- Stored precision
- bfloat16
- Model type
- gemma3
- Quantization
- awq
Identity and Version
- Repository
- gaunernst/gemma-3-27b-it-int4-awq
- Publisher
- Thien Tran
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 27.4B parameters
- Languages
- awq
- Revision
- 7cf8bdc81343c635390dd1e1bf590ab22dd6f366
- First published
- 2025-03-21
- Last updated
- 2025-04-06
Files and Weights
18 files, 18.5 GB in total. The weights are 4 files totalling 18.5 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00004.safetensors | Weights | 5.0 GB | dfe46693f149 |
| model-00002-of-00004.safetensors | Weights | 4.8 GB | b7e265b15839 |
| model-00003-of-00004.safetensors | Weights | 4.8 GB | fc7283de8174 |
| model-00004-of-00004.safetensors | Weights | 3.9 GB | 0cff33cb628f |
| added_tokens.json | Configuration | 35 B | — |
| chat_template.json | Configuration | 1.6 KB | — |
| config.json | Configuration | 1.2 KB | — |
| convert_flax.py | Configuration | 10.9 KB | — |
| generation_config.json | Configuration | 215 B | — |
| model.safetensors.index.json | Configuration | 219.1 KB | — |
| preprocessor_config.json | Configuration | 570 B | — |
| processor_config.json | Configuration | 70 B | — |
| special_tokens_map.json | Configuration | 662 B | — |
| README.md | Documentation | 25.6 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 33.4 MB | 4667f2089529 |
| tokenizer.model | Tokenizer | 4.7 MB | 1299c11d7cf6 |
| tokenizer_config.json | Tokenizer | 1.2 MB | — |
License and Download
- License
- gemma
- Access
- Open weights, no gate
- Download size
- 18.5 GB
Released by Thien Tran through its official repository on Hugging Face.
Built From
- Derived from google/gemma-3-27b-it
- Described by arXiv:1705.03551
- Described by arXiv:1810.12440
- Described by arXiv:1903.00161
- Described by arXiv:1904.09728
- Described by arXiv:1905.07830
- Described by arXiv:1905.10044
- Described by arXiv:1907.10641
- Described by arXiv:1908.02660
- Described by arXiv:1910.11856
- Described by arXiv:1911.01547
- Described by arXiv:1911.11641
- Described by arXiv:2009.03300
- Described by arXiv:2103.03874
- Described by arXiv:2104.12756
- Described by arXiv:2106.03193
- Described by arXiv:2107.03374
- Described by arXiv:2108.07732
- Described by arXiv:2110.14168
- Described by arXiv:2203.10244
- Described by arXiv:2210.03057
- Described by arXiv:2304.06364
- Described by arXiv:2311.12022
- Described by arXiv:2311.16502
- Described by arXiv:2312.11805
- Described by arXiv:2404.12390
- Described by arXiv:2404.16816
- Described by arXiv:2502.12404
- Described by arXiv:2502.21228
- Quantized from google/gemma-3-27b-it
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 18.5 GB |
| 16-bit | 54.9 GB |
| 8-bit | 27.4 GB |
| 4-bit | 13.7 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About gemma-3-27b-it-int4-awq
How much GPU memory does gemma-3-27b-it-int4-awq need?
About 65.8 GB at 16-bit and 16.5 GB at 4-bit: the weights (27.4B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run gemma-3-27b-it-int4-awq on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use gemma-3-27b-it-int4-awq commercially?
Yes, with conditions. gemma-3-27b-it-int4-awq is released under Gemma Terms of Use. Gemma models are released under Google's Gemma Terms of Use, which permit commercial use and redistribution subject to the Gemma Prohibited Use Policy, whose restrictions must be passed on to anyone the model is distributed to.
Similar Models
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 4-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 8-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 6-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
LM Studio Community models highlights program. Highlighting new & noteworthy models by the community. Join the conversation on Discord. 5-bit quantized version of Qwen3.8-27B using MLX, optimized for Apple Silicon. Special thanks to the Apple Machine Learning Research team for creating MLX. LM Studio is not the creator, originator, or owner of any Model featured in the Community Model Program. Each Community Model is created and provided by third parties. LM Studio does not endorse, support, represent or guarantee the completeness, truthfulness, accuracy, or reliability of any Community Model. You understand that Community Models can produce content that might be offensive, harmful…
(He slides three glasses across the bar—wine for Shakespeare, water for Data, and a mysterious blue liquid for Spock) This model is a merge of: - nightmedia/Qwen3.8-27B-Brainwaves - migtissera/Synthia-4-27B Brainwaves nightmedia/Qwen3.8-27B-Brainwaves migtissera/Synthia-4-27B Late-Night Architectural Takeaways The Quantization Shield (mxfp8 Master Pass): Hitting 0.735 ARC-C on the 8-bit layout proves that your Cold-Fusion flagship anchors and Migel Tissera's Synthia agent paths reached absolute geometric equilibrium. Instead of losing performance to the 0.596 Heretic collapse, the curved hypersphere calculation completely shielded the model's core intelligence. The Perplexity Sweet Spot…