Granite-4.2-3B — Q4_K_M GGUF (imatrix)
A 4-bit GGUF of ibm-granite/granite-4.2-3b,
quantized with an importance matrix built from a mixed prose and tool-calling calibration corpus.
Of the three quantizations we made of this model, it's the closest to the original and the
smallest.
| File |
Quant |
Size |
Bits per weight |
Mean KLD vs f16 |
Same top token |
granite-4.2-3b-Q4_K_M.gguf |
Q4_K_M + imatrix |
2.1 GB |
4.90 |
0.062012 |
87.88% |
"Same top token" is llama.cpp's Same top p statistic: the share of positions where the
quantized model's most likely next token matches the f16 model's.
How it was made
- Convert:
convert_hf_to_gguf.py (mainline llama.cpp), bf16 safetensors → f16 GGUF.
Granite's scaling parameters (attention_scale, embedding_scale, residual_scale,
logits_scale) are supported natively.
- Importance matrix:
llama-imatrix over
bartowski's v6 calibration dataset:
prose plus tool-calling conversations, rendered through Granite's own chat template.
573 chunks of 512 tokens.
- Quantize:
llama-quantize --imatrix <imatrix> <f16.gguf> <out.gguf> Q4_K_M.
- Evaluate:
llama-perplexity --kl-divergence against the f16 model's saved logits, on the
wikitext-2-raw test split.
The calibration corpus matters. This build beat our best GPTQ build (0.110 KLD) by about 44%.
Plain wikitext-2 is a common default for imatrix calibration, but a corpus in the model's own
chat format with tool use covers more of what the model actually does.
Usage
llama-server -m granite-4.2-3b-Q4_K_M.gguf
Any llama.cpp-based runtime that supports Granite 4 should load it. We haven't benchmarked this
file's inference speed yet. The speed figures in the
GPTQ card are for vLLM and don't
apply here.
Which one should I use?
- llama.cpp, Ollama, LM Studio, or similar: this GGUF. It's the best quality and the smallest
of the three.
- vLLM: the GPTQ W4A16 build,
which has measured serving numbers on RDNA2.
- MXFP4 research: the MXFP4
build (compressed-tensors and GGUF). It's an experiment, not the recommended pick.
All quantizations compared
| Variant |
Format |
Size |
Mean KLD vs unquantized ↓ |
Perplexity (wikitext-2) |
Eval harness |
| Original (bf16) |
safetensors (2 shards) |
6.8 GB |
— |
19.94 |
transformers |
| Q4_K_M + imatrix |
GGUF, 4.90 BPW |
2.1 GB |
0.062 |
not measured |
llama.cpp |
| GPTQ W4A16 (int4) |
compressed-tensors |
2.7 GB |
0.110 |
21.11 (+5.88%) |
transformers |
| MXFP4A16 |
compressed-tensors |
2.6 GB |
0.1405 |
21.27 (+6.68%) |
transformers |
| MXFP4 GGUF (same weights, transcoded) |
GGUF |
2.5 GB |
0.138 |
+7.1% vs f16 |
llama.cpp |
Lower KLD means the quantized model's next-token distribution stays closer to the original's.
How to read this table. The GGUF rows come from a different harness than the others.
- GGUF rows: llama-perplexity --kl-divergence against the f16 GGUF's logits, on the wikitext-2-raw
test split.
- GPTQ and MXFP4 (compressed-tensors): a transformers script that loads each model next to the bf16 original and
compares them over 51,100 tokens of the same wikitext-2 test split.
Both measure mean KL divergence on the same corpus, but the code isn't identical. Compare rows
within the same harness directly; across harnesses, treat gaps as directional, not decimal-precise.
The MXFP4 weights were measured both ways (0.1405 in transformers, 0.138 in llama.cpp), which
gives a feel for how closely the two harnesses agree.
Hardware: all measurements are on an AMD Radeon Pro V620 (RDNA2, gfx1030, 32 GB) with ROCm
7.14. Nothing here was tested on NVIDIA or other AMD GPUs.
Credits