SAVRN
Search Contact SAVRN

Open-weight model · Any to any

gemma-4-E4B-it-GGUF

by GGML Org ggml-org/gemma-4-E4B-it-GGUF

Run with https://llama.app - https://huggingface.co/google/gemma-4-E4B-it - https://huggingface.co/google/gemma-4-E4B-it-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized-assistant …

Parameters
Context
Weights29.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.5M

Model Card

By GGML Org, published under apache-2.0, revision b8093469224f.

Run with https://llama.app - https://huggingface.co/google/gemma-4-E4B-it - https://huggingface.co/google/gemma-4-E4B-it-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized - add info - add dflash

Read GGML Org's full model card

gemma-4-E4B-it

Run with https://llama.app

llama serve -hf ggml-org/gemma-4-E4B-it-GGUF

Source models

  • https://huggingface.co/google/gemma-4-E4B-it
  • https://huggingface.co/google/gemma-4-E4B-it-assistant
  • https://huggingface.co/google/gemma-4-E4B-it-qat-q4_0-unquantized-assistant
  • https://huggingface.co/google/gemma-4-E4B-it-qat-q4_0-unquantized

TODOs

  • add info
  • add dflash

[!IMPORTANT] This model is automatically converted using https://github.com/ggml-org/convert

Identity and Version

Repository
ggml-org/gemma-4-E4B-it-GGUF
Publisher
GGML Org
Task
Any to any
Modality
Multimodal
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
b8093469224f83f5c38f691eb906c380e9e63114
First published
2026-04-01
Last updated
2026-07-26

Files and Weights

12 files, 29.6 GB in total. The weights are 8 files totalling 29.6 GB in gguf.

Weights8 files · 29.6 GB
Documentation1 file · 621 B
Other1 file · 708.1 KB
Repository2 files · 2.8 KB
Every file
FileTypeSizeSHA-256
gemma-4-E4B-it-BF16.ggufWeights15.1 GB f0669d20ec5b
gemma-4-E4B-it-Q4_0.ggufWeights4.6 GB a555b900214b
gemma-4-E4B-it-Q8_0.ggufWeights8.0 GB 34be82b17b49
mmproj-gemma-4-E4B-it-BF16.ggufWeights991.6 MB f77995e4b6a5
mmproj-gemma-4-E4B-it-Q8_0.ggufWeights559.9 MB 197f49a93027
mtp-gemma-4-E4B-it-BF16.ggufWeights171.8 MB 5fe3a6c5ebd2
mtp-gemma-4-E4B-it-Q4_0.ggufWeights59.7 MB d19a11e4fd50
mtp-gemma-4-E4B-it-Q8_0.ggufWeights98.7 MB f38ae6296265
README.mdDocumentation621 B
convert.logOther708.1 KB
.gitattributesRepository2.6 KB
.src_shaRepository212 B

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
29.6 GB
Download from GGML Org

Released by GGML Org through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published29.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About gemma-4-E4B-it-GGUF

Can I use gemma-4-E4B-it-GGUF commercially?

Yes. gemma-4-E4B-it-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Any to any

gemma-4-12B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-12B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-12B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-E4B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0

Model · Any to any

gemma-4-E4B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E4B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers

Model · Any to any

gemma-4-E2B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E2B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers

Model · Any to any

gemma-4-E2B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 transformers