SAVRN
Search Contact SAVRN

Open-weight model · Any to any

gemma-4-E2B-it-GGUF

by GGML Org ggml-org/gemma-4-E2B-it-GGUF

Run with https://llama.app - https://huggingface.co/google/gemma-4-E2B-it - https://huggingface.co/google/gemma-4-E2B-it-assistant - https://huggingface.co/google/gemma-4-E2B-it-qat-q40-unquantized-assistant …

Parameters
Context
Weights19.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads98.6k

Model Card

By GGML Org, published under apache-2.0, revision b4243c156154.

Run with https://llama.app - https://huggingface.co/google/gemma-4-E2B-it - https://huggingface.co/google/gemma-4-E2B-it-assistant - https://huggingface.co/google/gemma-4-E2B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-E2B-it-qat-q40-unquantized - add info - add dflash

Read GGML Org's full model card

gemma-4-E2B-it

Run with https://llama.app

llama serve -hf ggml-org/gemma-4-E2B-it-GGUF

Source models

  • https://huggingface.co/google/gemma-4-E2B-it
  • https://huggingface.co/google/gemma-4-E2B-it-assistant
  • https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-unquantized-assistant
  • https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-unquantized

TODOs

  • add info
  • add dflash

[!IMPORTANT] This model is automatically converted using https://github.com/ggml-org/convert

Identity and Version

Repository
ggml-org/gemma-4-E2B-it-GGUF
Publisher
GGML Org
Task
Any to any
Modality
Multimodal
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
b4243c156154b6dca9324415f8c7ccc098b4aed1
First published
2026-04-01
Last updated
2026-07-26

Files and Weights

12 files, 19.0 GB in total. The weights are 8 files totalling 19.0 GB in gguf.

Weights8 files · 19.0 GB
Documentation1 file · 621 B
Other1 file · 650.3 KB
Repository2 files · 2.7 KB
Every file
FileTypeSizeSHA-256
gemma-4-E2B-it-BF16.ggufWeights9.3 GB 246c5fb19e64
gemma-4-E2B-it-Q4_0.ggufWeights2.8 GB 8e30dff3ac4c
gemma-4-E2B-it-Q8_0.ggufWeights5.0 GB 996d08777aad
mmproj-gemma-4-E2B-it-BF16.ggufWeights986.8 MB 711e1e8f43fa
mmproj-gemma-4-E2B-it-Q8_0.ggufWeights557.4 MB 9406f99c16d6
mtp-gemma-4-E2B-it-BF16.ggufWeights170.2 MB 48a801d2ea0f
mtp-gemma-4-E2B-it-Q4_0.ggufWeights59.2 MB 718d3a440579
mtp-gemma-4-E2B-it-Q8_0.ggufWeights97.8 MB c4fba8d43b40
README.mdDocumentation621 B
convert.logOther650.3 KB
.gitattributesRepository2.5 KB
.src_shaRepository212 B

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
19.0 GB
Download from GGML Org

Released by GGML Org through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published19.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About gemma-4-E2B-it-GGUF

Can I use gemma-4-E2B-it-GGUF commercially?

Yes. gemma-4-E2B-it-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Any to any

gemma-4-E4B-it-GGUF

GGML Org

Run with https://llama.app - https://huggingface.co/google/gemma-4-E4B-it - https://huggingface.co/google/gemma-4-E4B-it-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized - add info - add dflash

Open weights apache-2.0

Model · Any to any

gemma-4-12B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-12B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-12B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-E4B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0

Model · Any to any

gemma-4-E4B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E4B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers

Model · Any to any

gemma-4-E2B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E2B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers