SAVRN
Search Contact SAVRN

Open-weight model

gemma-4-26B-A4B-NVFP4-lmhead

by Tenhkspark tenhkspark/gemma-4-26B-A4B-NVFP4-lmhead

gemma-4-26B-A4B-NVFP4-lmhead is an open-weight model from Tenhkspark, released under Apache License 2.0. It has 14.4B parameters and a 262,144-token context. At 16-bit it needs about 34.5 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 244 downloads a month.

NVFP4 derivative checkpoint built on nvidia/Gemma-4-26B-A4B-NVFP4 (NVIDIA's ModelOpt NVFP4 quantization of Google's gemma-4-26B-A4B-it), re-saved with tiewordembeddings=false and a separately NVFP4-quantized lmhead.weight.

Parameters14.4B
Context262,144
Weights19.2 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads244

Runs On

What it takes to serve gemma-4-26B-A4B-NVFP4-lmhead (14.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 28.8 GB 34.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 14.4 GB 17.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 7.2 GB 8.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

gemma-4-26B-A4B-NVFP4-lmhead on every accelerator the SAVRN Index prices, at every precision

Model Card

By Tenhkspark, published under apache-2.0, revision 365b3bc730a0.

Gemma 4 26B A4B — NVFP4 with untied lm_head, v2

English | 日本語 | 한국어 | 中文

NVFP4 derivative checkpoint built on nvidia/Gemma-4-26B-A4B-NVFP4 (NVIDIA's ModelOpt NVFP4 quantization of Google's gemma-4-26B-A4B-it), re-saved with tie_word_embeddings=false and a separately NVFP4-quantized lm_head.weight. Weights: 19.2 GB, on-GPU footprint 17.08 GiB. The weights are unchanged in v2; v2 is a new serving setup and image (tenhkspark/gemma-4-v2:v2). Serving setup: https://github.com/tenhkspark/gemma4-spark

v1 to v2

Read the full model card (565 words)

Configuration

Architecture
Gemma4ForConditionalGeneration
Context length (tokens)
262,144
Layers
30
Hidden size
2,816
Feed-forward size
2,112
Attention heads
16
Key/value heads
8
Head dimension
256
Vocabulary size
262,144
Experts
128
Sliding window (tokens)
1,024
Model type
gemma4
Quantization
modelopt

Identity and Version

Repository
tenhkspark/gemma-4-26B-A4B-NVFP4-lmhead
Publisher
Tenhkspark
Task
Not stated by the source
Modality
Other
Library
Not stated by the source
Parameters
14.4B parameters
Languages
ja
Revision
365b3bc730a02a63e6796057269fc22257fc1efa
First published
2026-09-20
Last updated
2026-10-01

Files and Weights

30 files, 19.2 GB in total. The weights are 3 files totalling 19.2 GB in safetensors.

Weights3 files · 19.2 GB
Configuration8 files · 4.8 MB
Tokenizer2 files · 32.2 MB
Documentation6 files · 30.2 KB
Other10 files · 31.6 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
lm_head_nvfp4.safetensorsWeights415.2 MB de5b97961d85
model-00001-of-00002.safetensorsWeights10.0 GB b5df31122600
model-00002-of-00002.safetensorsWeights8.8 GB ff11061ebf57
bench-cell.pyConfiguration3.9 KB —
config.jsonConfiguration8.9 KB —
generation_config.jsonConfiguration208 B —
hf_quant_config.jsonConfiguration4.6 KB —
model.safetensors.index.jsonConfiguration4.8 MB —
processor_config.jsonConfiguration1.7 KB —
tools/router.pyConfiguration12.2 KB —
untie-lmhead-fp8.pyConfiguration10.3 KB —
LICENSEDocumentation11.3 KB —
NOTICEDocumentation527 B —
README.ja.mdDocumentation5.1 KB —
README.ko.mdDocumentation4.7 KB —
README.mdDocumentation4.6 KB —
README.zh.mdDocumentation4.0 KB —
MD5SUMSOther1.5 KB —
chat_template.jinjaOther18.7 KB —
files.tsvOther695 B —
gemma4-v2-balanced-mtp8.envOther200 B —
gemma4-v2-balanced.envOther200 B —
gemma4-v2-prefill-first.envOther299 B —
gemma4-v2-serve.shOther8.1 KB —
gemma4-v2.envOther902 B —
gemma4.small.envOther856 B —
tools/router-v2.tsvOther170 B —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer32.2 MB cc8d3a0ce364
tokenizer_config.jsonTokenizer2.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
19.2 GB
Download from Tenhkspark

Released by Tenhkspark through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published19.2 GB
16-bit28.8 GB
8-bit14.4 GB
4-bit7.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About gemma-4-26B-A4B-NVFP4-lmhead

How much GPU memory does gemma-4-26B-A4B-NVFP4-lmhead need?

About 34.5 GB at 16-bit and 8.6 GB at 4-bit: the weights (14.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run gemma-4-26B-A4B-NVFP4-lmhead on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use gemma-4-26B-A4B-NVFP4-lmhead commercially?

Yes. gemma-4-26B-A4B-NVFP4-lmhead is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is gemma-4-26B-A4B-NVFP4-lmhead's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.