Granite-4.2-3B — MXFP4 (GPTQ-calibrated): compressed-tensors + GGUF
An experimental MXFP4 quantization of
ibm-granite/granite-4.2-3b, made with
GPTQ calibration instead of the usual data-free rounding, and provided in two formats with
identical weights. It's published as a data point, not as the recommended build. For most
users, the Q4_K_M GGUF or
GPTQ int4 build is the better
choice.
| File(s) |
Format |
Size |
Mean KLD |
Perplexity |
Eval harness |
model.safetensors + config |
compressed-tensors (vLLM / transformers) |
2.6 GB |
0.1405 vs bf16 |
21.27 vs 19.94 (+6.68%) |
transformers |
granite-4.2-3b-MXFP4-gptq.gguf |
GGUF (llama.cpp) |
2.5 GB |
0.138 vs f16 |
+7.1% vs f16 |
llama.cpp |
Format: OCP MX — E2M1 (FP4) weights, group_size=32, one UE8M0 shared scale per group; weights
only, activations stay 16-bit. Embeddings, norms and lm_head are unquantized.
What the experiment shows
MXFP4 is usually produced data-free, by rounding weights straight to FP4 with no calibration
data. llm-compressor ships an MXFP4A16 scheme, but every example we found pairs GPTQModifier
with integer schemes only. So we tested whether GPTQ's calibration also helps MXFP4.
- GPTQ calibration clearly helps MXFP4. 0.1405 mean KLD beats the untuned int4 GPTQ build of
this same model (0.153). Our earlier MXFP4 results without calibration were far worse (0.233
data-free via llm-compressor, and 0.774 via
llama-quantize), but those were on granite-4.2-8b,
a different size, so treat that comparison as directional only.
- It still loses to integer 4-bit. The tuned GPTQ int4 build (0.110) and the Q4_K_M imatrix
GGUF (0.062) are both closer to the original.
- No speed benefit on most GPUs. MXFP4 is only faster on GPUs with native MXFP4 support. On
hardware without it, including the RDNA2 V620 we tested, you get no speedup.
This build is useful if you're studying MXFP4 quality, or targeting hardware where MXFP4 runs
natively. Otherwise, pick one of the other two.
How it was made
- Quantize: llm-compressor
GPTQModifier
with scheme="MXFP4A16", otherwise the same recipe as the GPTQ int4 build. Calibration:
neuralmagic/LLM_compression_calibration,
512 samples, max_seq_length=512. The run completed without errors, which confirms
GPTQModifier works with this float scheme.
- GGUF (byte-level transcode, no re-quantization). llama.cpp's
convert_hf_to_gguf.py doesn't
accept compressed-tensors MXFP4, and re-quantizing with llama-quantize would discard the
GPTQ-chosen values. Both formats store the same thing — 4-bit E2M1 codes plus one E8M0 scale per
32 weights — so a small script copies the codes and scales directly into llama.cpp's
block_mxfp4 layout. Only the nibble pairing within each block differs; the script also applies
llama.cpp's q/k row reordering. All 280 converted tensors decode bit-identically to the
compressed-tensors weights, and the GGUF's KLD (0.138, llama.cpp harness) matches the source
checkpoint (0.1405, transformers harness).
Usage
llama.cpp:
llama-server -m granite-4.2-3b-MXFP4-gptq.gguf
Loads and generates correctly with a recent mainline llama.cpp, on CPU and on ROCm (gfx1030). The
loader prints unknown type mxfp4 while guessing the file type; that warning is harmless. We
haven't benchmarked inference speed.
transformers: we loaded the compressed-tensors checkpoint with transformers (with
compressed-tensors installed) to evaluate it. We haven't tested serving it with vLLM on any
hardware.
All quantizations compared
| Variant |
Format |
Size |
Mean KLD vs unquantized ↓ |
Perplexity (wikitext-2) |
Eval harness |
| Original (bf16) |
safetensors (2 shards) |
6.8 GB |
— |
19.94 |
transformers |
| Q4_K_M + imatrix |
GGUF, 4.90 BPW |
2.1 GB |
0.062 |
not measured |
llama.cpp |
| GPTQ W4A16 (int4) |
compressed-tensors |
2.7 GB |
0.110 |
21.11 (+5.88%) |
transformers |
| MXFP4A16 |
compressed-tensors |
2.6 GB |
0.1405 |
21.27 (+6.68%) |
transformers |
| MXFP4 GGUF (same weights, transcoded) |
GGUF |
2.5 GB |
0.138 |
+7.1% vs f16 |
llama.cpp |
Lower KLD means the quantized model's next-token distribution stays closer to the original's.
How to read this table. The GGUF rows come from a different harness than the others.
- GGUF rows: llama-perplexity --kl-divergence against the f16 GGUF's logits, on the wikitext-2-raw
test split.
- GPTQ and MXFP4 (compressed-tensors): a transformers script that loads each model next to the bf16 original and
compares them over 51,100 tokens of the same wikitext-2 test split.
Both measure mean KL divergence on the same corpus, but the code isn't identical. Compare rows
within the same harness directly; across harnesses, treat gaps as directional, not decimal-precise.
The MXFP4 weights were measured both ways (0.1405 in transformers, 0.138 in llama.cpp), which
gives a feel for how closely the two harnesses agree.
Hardware: all measurements are on an AMD Radeon Pro V620 (RDNA2, gfx1030, 32 GB) with ROCm
7.14. Nothing here was tested on NVIDIA or other AMD GPUs.
Credits