# gemma-4-26B-A4B-NVFP4-lmhead by Tenhkspark: Open Model
Source: https://savrn.com/models/gemma-4-26b-a4b-nvfp4-lmhead
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve gemma-4-26B-A4B-NVFP4-lmhead (14.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 28.8 GB | 34.5 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 14.4 GB | 17.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 7.2 GB | 8.6 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[gemma-4-26B-A4B-NVFP4-lmhead on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/gemma-4-26b-a4b-nvfp4-lmhead/gpus)

## Model Card

By Tenhkspark, published under apache-2.0, revision 365b3bc730a0.

### Gemma 4 26B A4B — NVFP4 with untied lm_head, v2

English | 日本語 | 한국어 | 中文

NVFP4 derivative checkpoint built on nvidia/Gemma-4-26B-A4B-NVFP4 (NVIDIA's ModelOpt NVFP4 quantization of Google's gemma-4-26B-A4B-it), re-saved with tie_word_embeddings=false and a separately NVFP4-quantized lm_head.weight. Weights: 19.2 GB, on-GPU footprint 17.08 GiB. The weights are unchanged in v2; v2 is a new serving setup and image ([tenhkspark/gemma-4-v2:v2](https://hub.docker.com/r/tenhkspark/gemma-4-v2)). Serving setup: https://github.com/tenhkspark/gemma4-spark

### v1 to v2

[Read the full model card (565 words)](https://savrn.com/models/gemma-4-26b-a4b-nvfp4-lmhead/card)

## Configuration

Architecture

Gemma4ForConditionalGeneration

Context length (tokens)

262,144

Layers

30

Hidden size

2,816

Feed-forward size

2,112

Attention heads

16

Key/value heads

8

Head dimension

256

Vocabulary size

262,144

Experts

128

Sliding window (tokens)

1,024

Model type

gemma4

Quantization

modelopt

## Identity and Version

Repository

tenhkspark/gemma-4-26B-A4B-NVFP4-lmhead

Publisher

Tenhkspark

Task

Not stated by the source

Modality

Other

Library

Not stated by the source

Parameters

14.4B parameters

Languages

ja

Revision

365b3bc730a02a63e6796057269fc22257fc1efa

First published

2026-09-20

Last updated

2026-10-01

## Files and Weights

30 files, 19.2 GB in total. The weights are 3 files totalling 19.2 GB in safetensors.

Weights3 files · 19.2 GB

Configuration8 files · 4.8 MB

Tokenizer2 files · 32.2 MB

Documentation6 files · 30.2 KB

Other10 files · 31.6 KB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| lm_head_nvfp4.safetensors | Weights | 415.2 MB | de5b97961d85 |
| model-00001-of-00002.safetensors | Weights | 10.0 GB | b5df31122600 |
| model-00002-of-00002.safetensors | Weights | 8.8 GB | ff11061ebf57 |
| bench-cell.py | Configuration | 3.9 KB | — |
| config.json | Configuration | 8.9 KB | — |
| generation_config.json | Configuration | 208 B | — |
| hf_quant_config.json | Configuration | 4.6 KB | — |
| model.safetensors.index.json | Configuration | 4.8 MB | — |
| processor_config.json | Configuration | 1.7 KB | — |
| tools/router.py | Configuration | 12.2 KB | — |
| untie-lmhead-fp8.py | Configuration | 10.3 KB | — |
| LICENSE | Documentation | 11.3 KB | — |
| NOTICE | Documentation | 527 B | — |
| README.ja.md | Documentation | 5.1 KB | — |
| README.ko.md | Documentation | 4.7 KB | — |
| README.md | Documentation | 4.6 KB | — |
| README.zh.md | Documentation | 4.0 KB | — |
| MD5SUMS | Other | 1.5 KB | — |
| chat_template.jinja | Other | 18.7 KB | — |
| files.tsv | Other | 695 B | — |
| gemma4-v2-balanced-mtp8.env | Other | 200 B | — |
| gemma4-v2-balanced.env | Other | 200 B | — |
| gemma4-v2-prefill-first.env | Other | 299 B | — |
| gemma4-v2-serve.sh | Other | 8.1 KB | — |
| gemma4-v2.env | Other | 902 B | — |
| gemma4.small.env | Other | 856 B | — |
| tools/router-v2.tsv | Other | 170 B | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 32.2 MB | cc8d3a0ce364 |
| tokenizer_config.json | Tokenizer | 2.1 KB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

19.2 GB

[Download from Tenhkspark](https://huggingface.co/tenhkspark/gemma-4-26B-A4B-NVFP4-lmhead)

Released by Tenhkspark through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from [google/gemma-4-26B-A4B-it](https://savrn.com/models/gemma-4-26b-a4b-it)
- Quantized from [google/gemma-4-26B-A4B-it](https://savrn.com/models/gemma-4-26b-a4b-it)

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 19.2 GB |
| 16-bit | 28.8 GB |
| 8-bit | 14.4 GB |
| 4-bit | 7.2 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About gemma-4-26B-A4B-NVFP4-lmhead

### How much GPU memory does gemma-4-26B-A4B-NVFP4-lmhead need?

About 34.5 GB at 16-bit and 8.6 GB at 4-bit: the weights (14.4B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run gemma-4-26B-A4B-NVFP4-lmhead on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use gemma-4-26B-A4B-NVFP4-lmhead commercially?

Yes. gemma-4-26B-A4B-NVFP4-lmhead is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

### What is gemma-4-26B-A4B-NVFP4-lmhead's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Tenhkspark

[All models and datasets](https://savrn.com/model-publishers/tenhkspark)

## Versions

- [365b3bc730a0](https://savrn.com/models/gemma-4-26b-a4b-nvfp4-lmhead/versions/365b3bc730a0) · current 2026-10-01

## Explore More

- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-01.
- [Hugging Face record](https://huggingface.co/tenhkspark/gemma-4-26B-A4B-NVFP4-lmhead)
- [How the hub is built](https://savrn.com/model-hub/methodology)
