Granite-4.2-3B — GPTQ W4A16 (int4, compressed-tensors)
A 4-bit weight-only GPTQ quantization of
ibm-granite/granite-4.2-3b, saved in
compressed-tensors format for vLLM. On an AMD Radeon Pro V620 (RDNA2), vLLM served it at
65 tok/s single-sequence and ~1,000 tok/s aggregate at 16 concurrent sequences.
| Metric |
Value |
| Size |
2.7 GB (model.safetensors 2.67 GB) |
| Mean KLD vs bf16 |
0.110 |
| Worst single-token KLD |
11.74 |
| Perplexity (wikitext-2) |
21.11 vs 19.94 for bf16 (+5.88%) |
How it was made
llm-compressor GPTQModifier:
- int4, symmetric, weight-only (activations stay 16-bit)
group_size=32, actorder="weight"
lm_head left unquantized
- calibration: neuralmagic/LLM_compression_calibration,
512 samples,
max_seq_length=512
The tuning mattered. With the default W4A16 preset (group_size=128), the same pipeline scored
0.153 mean KLD. Moving to group_size=32 with actorder="weight" cut that by about 28%, for
roughly 0.2 GB more disk. For comparison, an AWQ build (W4A16 asymmetric, group_size=128) from
the same pipeline scored 0.463, about 3x worse. We suspect that's because llm-compressor had to
guess AWQ's layer mappings for Granite, but we didn't confirm the cause.
Serving with vLLM
Tested on the AMD Radeon Pro V620 (gfx1030) with ROCm 7.14, using a self-built vLLM development
snapshot. vLLM picked RDNA2W4A16LinearKernel for this checkpoint.
vllm serve quark75/granite-4.2-3b-W4A16-GPTQ-gs32 --enforce-eager --dtype float16
--dtype float16 matches how the numbers below were measured; RDNA2 has no bf16 dot-product
hardware.
| Load |
Throughput (--enforce-eager) |
| 1 sequence |
65.04 tok/s |
| 16 concurrent sequences |
1003.56 tok/s aggregate |
Use --enforce-eager on RDNA2 (gfx1030). On this model and GPU, CUDA/HIP graph capture
produced wrong output, not an error. vLLM's own graph capture and a separate
torch.cuda.graph() attempt both failed the same way: under greedy decoding, 16 identical
prompts gave different outputs, and one degenerated into repeated tokens. Graph mode looked
faster (95 tok/s single-sequence), but its output was wrong. We haven't found the root cause.
We haven't tested other GPUs, so this may or may not apply to yours. If you enable graph
capture anywhere, check that greedy output is deterministic first.
All quantizations compared
| Variant |
Format |
Size |
Mean KLD vs unquantized ↓ |
Perplexity (wikitext-2) |
Eval harness |
| Original (bf16) |
safetensors (2 shards) |
6.8 GB |
— |
19.94 |
transformers |
| Q4_K_M + imatrix |
GGUF, 4.90 BPW |
2.1 GB |
0.062 |
not measured |
llama.cpp |
| GPTQ W4A16 (int4) |
compressed-tensors |
2.7 GB |
0.110 |
21.11 (+5.88%) |
transformers |
| MXFP4A16 |
compressed-tensors |
2.6 GB |
0.1405 |
21.27 (+6.68%) |
transformers |
| MXFP4 GGUF (same weights, transcoded) |
GGUF |
2.5 GB |
0.138 |
+7.1% vs f16 |
llama.cpp |
Lower KLD means the quantized model's next-token distribution stays closer to the original's.
How to read this table. The GGUF rows come from a different harness than the others.
- GGUF rows: llama-perplexity --kl-divergence against the f16 GGUF's logits, on the wikitext-2-raw
test split.
- GPTQ and MXFP4 (compressed-tensors): a transformers script that loads each model next to the bf16 original and
compares them over 51,100 tokens of the same wikitext-2 test split.
Both measure mean KL divergence on the same corpus, but the code isn't identical. Compare rows
within the same harness directly; across harnesses, treat gaps as directional, not decimal-precise.
The MXFP4 weights were measured both ways (0.1405 in transformers, 0.138 in llama.cpp), which
gives a feel for how closely the two harnesses agree.
Hardware: all measurements are on an AMD Radeon Pro V620 (RDNA2, gfx1030, 32 GB) with ROCm
7.14. Nothing here was tested on NVIDIA or other AMD GPUs.
Credits
- Base model: ibm-granite/granite-4.2-3b
by IBM, Apache 2.0. All credit for the model itself goes to IBM's Granite team. These are
unofficial quantizations, not affiliated with or endorsed by IBM.
- Tools: llm-compressor, vLLM.