This repository contains an MXFP8-weight checkpoint derived from exported in compressed-tensors format. The checkpoint retains the source multimodal components, but the evaluation reported here covers text tasks only. The exported checkpoint contains 11,635 F8E4M3 weight tensors, 838 BF16 weight tensors, 11,635 U8 block-scale tensors, and 171 FP32 scale/metadata tensors. Routed-expert weights and source-present text self-attention projections are MXFP8; the router, shared MLP, vision tower, and embeddings remain BF16. Measured on 2026-09-30 with lm-eval 0.4.13 and vLLM 0.29.0. The tested settings were TRITONATTN, tensor parallelism 2, pipeline parallelism 1, batch size 64, maxnumseqs=64…
Open-weight model · Text generation
gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
by INC Optimized Models 4 INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound is an open-weight model for text generation from INC Optimized Models 4, released under Apache License 2.0. It has 25.8B parameters and a 262,144-token context. At 16-bit it needs about 61.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
This checkpoint is an AutoRound model-free MXFP8 RTN quantization of exported in llmcompressor / compressed-tensors format.
Runs On
What it takes to serve gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound (25.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 51.6 GB | 61.9 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 25.8 GB | 31.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 12.9 GB | 15.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
Model Card
By INC Optimized Models 4, published under apache-2.0, revision dc8c7c2ba6d6.
This checkpoint is an AutoRound model-free MXFP8 RTN quantization of exported in llmcompressor / compressed-tensors format. Routed experts and the self-attention projections present in the source are stored as F8E4M3; sensitive/shared and multimodal weights remain BF16. Static FP8 KV scales were calibrated with AutoRound using the text dataset NeelNanda/pile-10k; the vision tower was not quantized. On the paired repository lmeval protocol, the four primary metrics were non-decreasing relative to the BF16 baseline using the same vLLM FP8 KV cache. The AQA gate was GO. This is a result for those tasks and settings only, not a claim of lossless quantization or general quality improvement. The…
Read INC Optimized Models 4's full model card
This checkpoint is an AutoRound model-free MXFP8 RTN quantization of
google/gemma-4-26B-A4B,
exported in llm_compressor / compressed-tensors format. Routed experts and
the self-attention projections present in the source are stored as
F8_E4M3; sensitive/shared and multimodal weights remain BF16. Static FP8 KV
scales were calibrated with AutoRound using the text dataset
NeelNanda/pile-10k; the vision tower was not quantized.
On the paired repository lm_eval protocol, the four primary metrics were
non-decreasing relative to the BF16 baseline using the same vLLM FP8 KV cache.
The AQA gate was GO. This is a result for those tasks and settings only,
not a claim of lossless quantization or general quality improvement.
Read this first:
max_model_len=131072was configured, but the benchmark prompts were not a 128K long-context stress test. This run evaluated text tasks only; vision quality, multimodal prompts, thinking-on, throughput and latency were not measured. The tokenizer configuration used by the run had nochat_template, so GSM8K used the repository's bare-prompt path.
1. Model summary
| Base model | This checkpoint | |
|---|---|---|
| Name / architecture / parameters | Gemma 4 26B-A4B; Gemma4ForConditionalGeneration; nominal count 26,544,131,376 (26.544B; tied-weight counting convention) |
gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound; 25,805,936,206 unique serialized weight elements |
| Layer structure | 30 text layers; 128 routed experts/layer, top-8; 25 sliding-attention + 5 full-attention layers; vision tower | Routed-expert and source-present text self-attention projection weights MXFP8; protected modules BF16 |
| Weight tensors | 1,013 source index entries, BF16 | 24,222 exported tensor entries: 11,635 F8_E4M3 weights, 11,635 U8 block scales, 838 BF16 weights and 114 F32 static KV scales |
| Size on disk | 51,611,872,412 bytes, source safetensors index total_size |
28,412,087,908 bytes, exported safetensors index total_size |
| Compression / effective bits per parameter | BF16 source | 1.8165× smaller; 8.8079 effective bits per unique serialized weight element (8 × export bytes / 25,805,936,206) |
| Quantization / format / context / license | Gemma 4; Apache 2.0 | MXFP8 weights + static FP8 KV; compressed-tensors; evaluated with a 131,072-token maximum length; Apache 2.0 terms apply |
The model configuration sets tie_word_embeddings: true. The nominal
26.544B model count and the index-derived unique serialized element count are
different counting bases; the size and dtype accounting below uses the latter.
The A4B designation is the upstream model's active-parameter naming, not a
measurement made by this quantization run.
Serving was validated through the vLLM 0.29.0 backend on three RTX 5090 GPUs with TP=1 / PP=3. That is the tested topology, not a claim that other hardware or parallel layouts are equivalent.
2. Precision plan
Percentages below are relative to the 25,805,936,206 unique stored model-weight elements. Scale tensors are metadata and are not included in this denominator.
| Bit width / stored dtype | Module (layers) | Weight tensors | Weight elements | Share |
|---|---|---|---|---|
8-bit float (F8_E4M3) |
Routed expert gate/up/down weights (30 × 128 experts) | 11,520 | 22,837,985,280 | 88.4990% |
8-bit float (F8_E4M3) |
Existing text self-attention projections (q/k/o in 30 layers; v in 25) | 115 | 1,110,179,840 | 4.3020% |
16-bit (BF16) |
Shared MLP branch | 90 | 535,265,280 | 2.0742% |
16-bit (BF16) |
Router | 90 | 10,901,760 | 0.0422% |
16-bit (BF16) |
Vision tower | 355 | 569,550,384 | 2.2071% |
16-bit (BF16) |
Vision embedding projection | 1 | 3,244,032 | 0.0126% |
16-bit (BF16) |
Token embedding | 1 | 738,197,504 | 2.8606% |
16-bit (BF16) |
Other retained tensors, including norms and patch/position tensors | 301 | 612,126 | 0.0024% |
| Total model weights | F8_E4M3 + BF16 | 12,473 | 25,805,936,206 | 100.0000% |
Shares are rounded independently; the unrounded category shares sum to 100%.
The exported weight configuration is Linear, 8-bit float, symmetric,
group-wise, group size 32, static weights, with U8 scales and
memoryless_minmax observer. The config also records dynamic 8-bit float
input-activation quantization metadata. Static KV is 8-bit float, tensor-wise
and dynamic=false.
The exported checkpoint also has 11,635 U8 block-scale tensors and 114 F32 static KV scales. Of the latter, the 60 text K/V scales (30 layers × K/V) are finite and nonzero; the 54 vision scales are zero because calibration and evaluation were text-only.
Ignore groups
| Group | Reason / evidence |
|---|---|
router |
Preserve routing decisions; retained as BF16 in the header audit. |
mlp |
This is the always-on shared branch added to routed-expert output; retained as BF16. |
vision_tower, embed_vision |
Preserve the multimodal tower and projection; not part of this text-only evaluation. |
embed_tokens |
Embedding is not a supported AutoRound Linear target; source/export key is present and BF16. The logged unmatched-module warning is benign, not a missing checkpoint key. |
Norms and other non-target tensors remain BF16 by the model-free defaults. The
checkpoint contains no multi_modal_projector key, so it is not included in
the ignore list.
2.1 Quantization plan and artifact deviations
| Audited plan | Export observation | Difference / consequence |
|---|---|---|
| Quantize all routed experts | 11,520 F8_E4M3 tensors across 30 layers × 128 experts | Matches plan; 60 fused source expert keys were expanded into per-expert tensors. |
| Quantize source-present text self-attention projections | 115 F8_E4M3 projection tensors | Matches source index: q/k/o exist in all 30 layers, while v exists only in 25 sliding-attention layers. The five full-attention source layers have no v_proj. |
| Keep shared MLP, router, vision, embeddings and other non-target tensors BF16 | 838 BF16 tensors, 1,857,771,086 elements | Matches plan; no non-fused source keys were absent from the exported index. |
| Add static FP8 KV scales | 114 F32 scales; all finite; 60 text scales nonzero | Text scales cover the evaluated workload. Vision scales are zero after text-only calibration and were not exercised because language_model_only=true. |
There was no unplanned weight-precision change. The embed_tokens warning
reflects that embeddings are not Linear quantization targets; the key remains
present in BF16.
2.2 Size accounting
| Artifact | Header/index bytes | Comparison |
|---|---|---|
| BF16 source checkpoint | 51,611,872,412 | Source index total_size; 1,013 BF16 entries |
| MXFP8 + BF16 export | 28,412,087,908 | Export index total_size; 24,222 weight/scale entries |
The ratio is 51,611,872,412 / 28,412,087,908 = 1.8165×. Effective bits per
unique serialized model-weight element are 28,412,087,908 × 8 /
25,805,936,206 = 8.8079; this includes block/KV scale storage represented in
the exported index and is not the nominal MXFP8 weight bit width.
3. Evaluation results
Measured 2026-09-28 with lm_eval 0.4.13 using its vLLM backend (vLLM 0.29.0),
TP=1 / PP=3, seed 42, BF16 compute dtype and FP8 KV cache for both baseline and
candidate. Values are means ± lm_eval standard error (SE); N is the evaluated
sample count. The final column is the score difference divided by
sqrt(SE_baseline² + SE_MXFP8²); it is an approximate independent-SE scale,
not a paired significance test.
| Task / primary metric | N | BF16 baseline | MXFP8 + FP8 KV | Δ (pp) | Δ / combined SE |
|---|---|---|---|---|---|
PIQA (acc) |
1,838 | 0.8226 ± 0.0089 | 0.8286 ± 0.0088 | +0.60 | +0.48 |
MMLU (acc) |
14,042 | 0.7434 ± 0.0034 | 0.7441 ± 0.0034 | +0.07 | +0.15 |
HellaSwag (acc) |
10,042 | 0.6325 ± 0.0048 | 0.6347 ± 0.0048 | +0.22 | +0.32 |
GSM8K (exact_match,strict-match) |
1,319 | 0.7180 ± 0.0124 | 0.7324 ± 0.0122 | +1.44 | +0.83 |
The arithmetic mean of the four primary scores is 0.7291 (baseline) vs 0.7350
(MXFP8), a +0.58 pp difference; this mean is descriptive and has no combined
SE. PIQA acc_norm was 0.8395 ± 0.0086 vs 0.8341 ± 0.0087 (−0.54 pp);
HellaSwag acc_norm was 0.8305 ± 0.0037 vs 0.8317 ± 0.0037 (+0.12 pp).
The AQA gate uses the primary metrics above and reports GO
(max_accuracy_drop=0.02). The benchmark differences are within one
approximate combined SE; this run does not establish statistical superiority.
3.2 What was held fixed
Both measured runs used max_model_len=131072, kv_cache_dtype=fp8,
tensor_parallel_size=1, pipeline_parallel_size=3,
gpu_memory_utilization=0.85, dtype=bfloat16,
attention_backend=FLASHINFER, language_model_only=true, max_num_seqs=1,
trust_remote_code=true, add_bos_token=true, prefix caching disabled,
thinking disabled and seed 42. PIQA/MMLU/HellaSwag used the repository task
group with batch_size=auto; GSM8K used 5-shot and batch_size=64, with a
2,048-token generation budget. Inference-engine internal scheduling details
not exposed by lm_eval are not claimed as controlled measurements.
3.3 Versus baseline
All four AQA primary scores are non-decreasing in this paired run. The differences are small relative to their reported task-level uncertainty, so the result supports the stated gate decision but does not prove that MXFP8 improves accuracy. An earlier failed run had a different repeated BF16 GSM8K strict-match score (0.7286 vs 0.7180); its cause is unknown and it is not mixed into this gate comparison.
3.4 Functional evidence and provenance
Direct evidence: the exported checkpoint loaded in vLLM 0.29.0 with FP8 KV
cache and completed all four listed lm_eval tasks. Its
quantization_config.json records compressed-tensors MXFP8 weights and a
static 8-bit float KV scheme. The validation run ID is
gemma-4-26B-A4B_MXFP8_20260927-220842_a0c394.
Not measured: full 128K prompts, multimodal/vision quality, thinking-on, agentic tasks, throughput or latency. Environment support probes and sibling model runs are not substituted for this checkpoint's measured evidence.
4. Usage
4.1 vLLM through the validated lm_eval path
The run validated vLLM loading/evaluation via lm_eval. After the repository is available, the following uses its Hub ID in place of the original local checkpoint path:
MODEL_ID=INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
MODEL_ARGS="pretrained=${MODEL_ID},tensor_parallel_size=1,pipeline_parallel_size=3,max_model_len=131072,gpu_memory_utilization=0.85,dtype=bfloat16,trust_remote_code=True,add_bos_token=True,enable_prefix_caching=False,max_gen_toks=2048,attention_backend=FLASHINFER,language_model_only=True,kv_cache_dtype=fp8,max_num_seqs=1,enable_thinking=False"
# Select three idle GPUs before running; this example assumes physical GPUs 1-3 are idle.
export CUDA_VISIBLE_DEVICES=1,2,3
lm_eval --model vllm --model_args "$MODEL_ARGS" \
--tasks piqa,mmlu,hellaswag --batch_size auto --log_samples --seed 42
lm_eval --model vllm --model_args "$MODEL_ARGS" \
--tasks gsm8k --batch_size 64 --log_samples --seed 42 --num_fewshot 5
Hard constraints and evidence:
| Setting | Requirement / evidence |
|---|---|
| KV cache | Keep kv_cache_dtype=fp8 to match the evaluated quantized model protocol; confirmed in both baseline/candidate lm_eval metadata and the export KV scheme. |
| Context | max_model_len=131072 is the measured configuration. This is not evidence of 128K prompt throughput or quality. |
| Parallel layout | TP=1 / PP=3 on three GPUs is the tested layout. PP=2 failed CUDA graph warmup with an additional 636 MiB allocation in logs/aqa-run-retry-01.log; other layouts are unverified. |
| Backend | vLLM 0.29.0 with attention_backend=FLASHINFER was exercised by the four evaluation tasks. For head dimension 512, the XQA decode path logged a fallback to native FlashInfer decode; KV remained FP8 and evaluation completed. |
| Prompt formatting | The run's local/upstream tokenizer configuration had no chat_template. GSM8K therefore used apply_chat_template=false and fewshot_as_multiturn=false; multimodal/chat formatting was not evaluated. |
For text inference, supply a plain prompt compatible with the task rather than
assuming a chat template. A direct production vllm serve deployment was not
separately benchmarked by this run.
4.2 Transformers inspection path
The following reads model configuration only; it is not a weight-loading or serving validation path:
from transformers import AutoConfig
model_id = "INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound"
config = AutoConfig.from_pretrained(model_id, trust_remote_code=True)
print(config.model_type)
The checkpoint contains about 28.41 GB of indexed tensors. Loading full weights through an inspection-oriented Transformers path can use substantially more host/GPU memory; use the tested vLLM backend for inference.
5. Reproducibility
5.1 Hardware / OS
- GPUs: 3 × NVIDIA GeForce RTX 5090 used (compute capability 12.0; physical devices 1, 2 and 3); 32,607 MiB total memory reported per device.
- Host probe on 2026-09-28: NVIDIA driver 610.57.04; Intel Xeon 6767P (256 logical CPUs); 188.4 GiB RAM.
- OS: Ubuntu 24.04.4 LTS, kernel 6.8.0-139-generic, x86_64.
- CUDA: PyTorch build CUDA 13.0; runtime 13.0.88.
- Storage: model and run were on local filesystem paths; filesystem type was not recorded.
5.2 Environment
| Package | Version |
|---|---|
| Python | 3.12.14 |
| auto-round | 0.16.0.dev193+g22d6e28c |
| transformers | 5.17.0 |
| torch | 2.13.0+cu130 |
| vllm | 0.29.0 |
| flashinfer-python | 0.7.0 |
| compressed-tensors | 0.17.0 |
| lm-eval | 0.4.13 |
| autoquant-agent | 0.1.0 |
Versions are from the run manifest / frozen Conda environment, not a statement of minimum compatible versions.
5.3 Version selection
No packages were installed or upgraded for this run. These exact versions were used for the successful quantization and evaluation. Compatibility with other AutoRound, Transformers, vLLM or FlashInfer versions was not established here.
5.4 Artifact-level dependency warning
AutoRound 0.16.0.dev193 with Transformers 5.17.0 initially failed when its
fused-MoE meta skeleton could not materialize the new self_attn.k_scale and
self_attn.v_scale parameters. The successful run set
AR_DISABLE_META_LOAD=1, which disabled that meta-load path and allowed CPU
loading/initialization. Peak host RAM was 64.14 GiB. No source-code patch was
applied.
5.5 Environment variables
For quantization, set the tested workaround and map three verified idle physical devices:
export AR_DISABLE_META_LOAD=1
export CUDA_VISIBLE_DEVICES=<three-idle-physical-GPU-IDs>
AR_DISABLE_META_LOAD is a quantization workaround; it is not required for
inference. Do not put Hub tokens in this README or in a run YAML.
6. Reproduce the artifact
The original run was orchestrated with AQA; do not bypass it with a manually assembled AutoRound or lm_eval pipeline. Use fresh local paths and select three idle physical GPUs before starting.
6.1 Quantization YAML
The following is the core of the validated run.yaml; adapt only model,
environment and output paths:
model: /path/to/gemma-4-26B-A4B
quant:
scheme: MXFP8
method: model_free
export_format: llm_compressor
ignore_layers: "router,vision_tower,embed_vision,embed_tokens,mlp"
static_kv_fp8: true
device_map: "0,1,2"
eval:
backend: vllm
harness: lm_eval
benchmarks: [piqa, mmlu, hellaswag, gsm8k]
baseline: true
thinking: "off"
gen_max_toks: 2048
lm_eval:
group: true
seed: 42
apply_chat_template: false
fewshot_as_multiturn: false
model_args:
attention_backend: FLASHINFER
language_model_only: true
kv_cache_dtype: fp8
max_model_len: 131072
max_num_seqs: 1
gpu_memory_utilization: 0.85
tensor_parallel_size: 1
pipeline_parallel_size: 3
enable_thinking: false
enable_prefix_caching: false
gate:
max_accuracy_drop: 0.02
model_free uses the model-free RTN weight path, but static_kv_fp8 still
loads the checkpoint and performs calibration forwards for KV scales. The
successful log used the standard text calibration dataset
NeelNanda/pile-10k; the vision tower was not quantized.
6.2 Exact quantization arguments
These are the exact AutoRound flags recorded from the successful AQA builder
run; only output_dir is parameterized here. Reproduce the run through AQA as
shown below:
auto-round --model_name /models/gemma-4-26B-A4B --scheme MXFP8 --model_free \
--static_kv_dtype fp8 \
--ignore_layers router,vision_tower,embed_vision,embed_tokens,mlp \
--device_map 0,1,2 --format llm_compressor \
--output_dir <fresh-run-dir>/quantized
6.3 Validate and run
export AR_DISABLE_META_LOAD=1
export CUDA_VISIBLE_DEVICES=<three-idle-physical-GPU-IDs>
aqa validate run.yaml
aqa run run.yaml --dry-run
aqa run run.yaml
The tested run completed the full AQA baseline → quantization → evaluation pipeline. Its output had six safetensors shards and an index size of 28,412,087,908 bytes.
6.4 Evaluate
The measured commands used the same vLLM model arguments shown in §4.1. The GSM8K task used five few-shot examples, seed 42 and a 2,048-token generation budget, with chat-template and multiturn formatting disabled because of the tokenizer configuration observed in this run.
7. Known issues and caveats
- Do not infer long-context behavior from
max_model_len=131072; no 128K prompt was measured. - Do not infer vision/multimodal quality from the text-only results. The vision weights remain BF16, but their quality was not evaluated.
- PP=2 at the same length hit a CUDA graph warmup OOM on this host; the tested topology is TP=1 / PP=3 on three RTX 5090 GPUs.
- FlashInfer XQA did not support this model's 512 head dimension in the tested path; vLLM fell back to native FlashInfer decode. FP8 KV evaluation still completed.
- The token embedding warning is benign: it is not an AutoRound
Lineartarget, its key is present, and it remains BF16. - A previous BF16 GSM8K baseline repeat scored 0.7286 vs 0.7180 in the successful run. The cause is unknown; the reported gate compares baseline and MXFP8 only within the successful run.
- PIQA
acc_normdecreased by 0.54 pp even though the AQA primaryaccincreased. Consult both metric and uncertainty tables; do not call the quantization lossless.
8. License and attribution
The base model metadata identifies Apache 2.0 and links to the Gemma 4 license information. Users must comply with the base-model license and preserve its attribution. This quantized checkpoint is not an official Google release.
Quantization: Intel AutoRound.
Inference validation: vLLM and
FlashInfer. Evaluation:
lm-evaluation-harness.
Text calibration used NeelNanda/pile-10k; calibration samples are not included
in this repository. This repository contains the quantized weights,
tokenizer/processor configuration and model configuration; it does not
redistribute benchmark datasets.
Configuration
- Architecture
- Gemma4ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 30
- Hidden size
- 2,816
- Feed-forward size
- 2,112
- Attention heads
- 16
- Key/value heads
- 8
- Head dimension
- 256
- Vocabulary size
- 262,144
- Experts
- 128
- Sliding window (tokens)
- 1,024
- Model type
- gemma4
- Quantization
- compressed-tensors
Identity and Version
- Repository
- INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
- Publisher
- INC Optimized Models 4
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 25.8B parameters
- Languages
- Not stated by the source
- Revision
- dc8c7c2ba6d6d0fb803f8966052e594ae45f52b7
- First published
- 2026-09-28
- Last updated
- 2026-09-28
Files and Weights
15 files, 28.5 GB in total. The weights are 6 files totalling 28.4 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00006.safetensors | Weights | 5.4 GB | 88c867b5b419 |
| model-00002-of-00006.safetensors | Weights | 5.4 GB | bb608c31a522 |
| model-00003-of-00006.safetensors | Weights | 5.4 GB | a0b9b9e9fb8c |
| model-00004-of-00006.safetensors | Weights | 5.4 GB | 83ab3c5a52d8 |
| model-00005-of-00006.safetensors | Weights | 5.4 GB | 96c5c01a5726 |
| model-00006-of-00006.safetensors | Weights | 1.6 GB | 4fcc361091d7 |
| config.json | Configuration | 25.1 KB | — |
| generation_config.json | Configuration | 177 B | — |
| model.safetensors.index.json | Configuration | 2.5 MB | — |
| processor_config.json | Configuration | 1.7 KB | — |
| quantization_config.json | Configuration | 20.2 KB | — |
| README.md | Documentation | 19.9 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 32.2 MB | 12bac982b793 |
| tokenizer_config.json | Tokenizer | 1.5 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 28.4 GB
Released by INC Optimized Models 4 through its official repository on Hugging Face. Read the license.
Built From
- Derived from google/gemma-4-26B-A4B
- Quantized from google/gemma-4-26B-A4B
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 28.4 GB |
| 16-bit | 51.6 GB |
| 8-bit | 25.8 GB |
| 4-bit | 12.9 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
How much GPU memory does gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound need?
About 61.9 GB at 16-bit and 15.5 GB at 4-bit: the weights (25.8B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound commercially?
Yes. gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
Gemma 4 26B-A4B, post-trained with GRPO against a reward model learned from 1.2 million double-blind votes cast by HiWaifu users inside their own role-play conversations. Put back into the same arena, blind, it met GLM-5.1 in 1,430 battles and won 49.6% of the decided votes; against a 13-model field including Gemini, DeepSeek-v4 and Qwen's character models it won 54.7%. Most open role-play models are tuned on preferences that come from an LLM judge, from a handful of annotators, or from synthetic pairs. We had something rarer: a live arena where, inside ordinary chats on our platform, a user is occasionally shown two candidate replies and asked which one they want to continue with. Those…
Darwin-27B-RSI is Darwin-27B-Opus after Recursive Self-Improvement (RSI): the model was improved using only signal it produced itself. During self-improvement, the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically (agreement across its own samples and code execution). Under an identical evaluation protocol, Darwin-27B-RSI improves over its parent on graduate-level science reasoning — +5.24 points on GPQA Diamond (single sample) and +3.79 points with majority voting — with every gain statistically significant in paired tests. As the reasoning engine of Darwin-27B-JEV on…
Prism ML's ternary Ternary-Bonsai-2-27B build of Qwen/Qwen3.8-27B, repacked for chad, a Claude-Code-style local coding agent for Apple Silicon, with its speculative decoder bundled in. This is chad's default model. Created using Bonsai by Prism ML. with, already quantized. Nothing is built on first run. Every projection of Qwen3.8-27B (a dense qwen35 hybrid: 64 layers, 48 GatedDeltaNet + 16 full attention) is stored in a Hadamard-rotated basis: multiplied by a fixed sign vector and put through a blockwise Walsh-Hadamard transform offline, then quantized to 2-bit affine group-128 whose three levels reproduce the ternary set {−s, 0, +s}. The rotation costs no extra bits and no extra weight…
Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via JiRackDeltaNetTokenizer - JiRack DeltaNet understand video and images that best for Robotics also A fast and efficient 27B model optimized for CPU inference. Built on a Qwen3.8-style DeltaNet architecture (hybrid attention + SSM), with an updated tokenizer that includes Routing, Media, Vision, Sound, Tool call, and Robotics tags. Ready-to-run GGUF quantizations, and native Ollama support with reasoning disabled by default for fast, direct responses. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert…
Full 27B-class reasoning in ternary transformer weights — on everyday laptops - \~7.2 GB deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 95% of FP16 intelligence retained: 80.49 average across 15 thinking-mode benchmarks — a higher score than the conventional IQ2XXS build (72.73) at less than two-thirds of its footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01 - End-to-end ternary language weights across embeddings, attention projections, MLP…