SAVRN
Search Contact SAVRN

Open-weight model · Any to any

gemma-4-12B-it-GGUF

by GGML Org ggml-org/gemma-4-12B-it-GGUF

Run with https://llama.app - https://huggingface.co/google/gemma-4-12B-it - https://huggingface.co/google/gemma-4-12B-it-assistant - https://huggingface.co/google/gemma-4-12B-it-qat-q40-unquantized-assistant …

Parameters
Context
Weights45.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads49.1k

Model Card

By GGML Org, published under apache-2.0, revision e3e681731089.

Run with https://llama.app - https://huggingface.co/google/gemma-4-12B-it - https://huggingface.co/google/gemma-4-12B-it-assistant - https://huggingface.co/google/gemma-4-12B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-12B-it-qat-q40-unquantized - add info - add dflash

Read GGML Org's full model card

gemma-4-12B-it

Run with https://llama.app

llama serve -hf ggml-org/gemma-4-12B-it-GGUF

Source models

  • https://huggingface.co/google/gemma-4-12B-it
  • https://huggingface.co/google/gemma-4-12B-it-assistant
  • https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-unquantized-assistant
  • https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-unquantized

TODOs

  • add info
  • add dflash

[!IMPORTANT] This model is automatically converted using https://github.com/ggml-org/convert

Identity and Version

Repository
ggml-org/gemma-4-12B-it-GGUF
Publisher
GGML Org
Task
Any to any
Modality
Multimodal
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
e3e681731089efaa3f0917336944ac64752db8ba
First published
2026-06-03
Last updated
2026-08-23

Files and Weights

12 files, 45.6 GB in total. The weights are 8 files totalling 45.6 GB in gguf.

Weights8 files · 45.6 GB
Documentation1 file · 621 B
Other1 file · 439.2 KB
Repository2 files · 2.4 KB
Every file
FileTypeSizeSHA-256
gemma-4-12B-it-BF16.ggufWeights23.8 GB 2851bc5b1a24
gemma-4-12B-it-Q4_0.ggufWeights7.2 GB 3712b9bd32ca
gemma-4-12B-it-Q8_0.ggufWeights12.7 GB abfc3044b937
mmproj-gemma-4-12B-it-BF16.ggufWeights175.1 MB 9b1edfa05b63
mmproj-gemma-4-12B-it-Q8_0.ggufWeights159.0 MB 59e62255435d
mtp-gemma-4-12B-it-BF16.ggufWeights861.5 MB 9b67d8e3650f
mtp-gemma-4-12B-it-Q4_0.ggufWeights253.7 MB b894e614824d
mtp-gemma-4-12B-it-Q8_0.ggufWeights465.1 MB 16c90eb9f2b2
README.mdDocumentation621 B
convert.logOther439.2 KB
.gitattributesRepository2.2 KB
.src_shaRepository212 B

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
45.6 GB
Download from GGML Org

Released by GGML Org through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published45.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About gemma-4-12B-it-GGUF

Can I use gemma-4-12B-it-GGUF commercially?

Yes. gemma-4-12B-it-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Any to any

gemma-4-E4B-it-GGUF

GGML Org

Run with https://llama.app - https://huggingface.co/google/gemma-4-E4B-it - https://huggingface.co/google/gemma-4-E4B-it-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized - add info - add dflash

Open weights apache-2.0

Model · Any to any

gemma-4-12B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-12B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-12B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-E4B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0

Model · Any to any

gemma-4-E4B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E4B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers

Model · Any to any

gemma-4-E2B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E2B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers