Qwen3-32B-MXFP4-W4A4 · Model Card
Qwen3-32B-MXFP4-W4A4: Model Card
Written by Cezar, published under apache-2.0, revision a49f042511ac, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.
Benchmark code, raw data and results: github.com/cezarc1/decodebench (see RESULTS.md).
MXFP4 W4A4 quantization of
Qwen/Qwen3-32B
(revision 9216db5781bf21249d130ec9da846c4624c16137, BF16). It was made for a controlled NVFP4-vs-MXFP4 decode
benchmark on NVIDIA B200 (fp4bench), not as a general-purpose release: the NVFP4 and MXFP4
checkpoints share the model, the tool, the recipe, the calibration data and the layer coverage,
and differ only in the format.
| Weights | MXFP4: FP4 E2M1 values in 32-element blocks, E8M0 (power-of-two) block scales, 4.25 bits per weight |
| Activations | 4-bit, quantized dynamically per 32-element block at runtime; this format has no static activation scales, so none are stored |
| Coverage | all Linear layers except lm_head (448 modules); embeddings, norms and lm_head stay BF16 |
| KV cache | BF16 at serve time (--kv-cache-dtype bfloat16); the checkpoint has no kv_cache_scheme |
How it was made
- Tool: llm-compressor 0.14.0, compressed-tensors 0.19.0.
Recipe:
QuantizationModifier(targets="Linear", scheme="MXFP4", ignore=["lm_head"]). The plain preset recipe, with no weight-rounding optimization (GPTQ, AutoRound and the like). - Calibration: 256 prompts from the first user turn of ShareGPT conversations,
anon8231489123/ShareGPT_Vicuna_unfilteredat revision192ab2185289094fc556ec8ce5ce1e8e587154ca, at most 1024 tokens each (33208 tokens in all, seed 3). MXFP4 has no static activation scales, so the data does not change this checkpoint's weights. The NVFP4 checkpoint uses the same set. - The source weights are the BF16 original, so this is the only quantization step.
- Quantized at 2026-10-04T23:56:31.271659+00:00.
Measured size
- Quantized Linear weights stored in this checkpoint (
weight_packed,weight_scaleand, for NVFP4,weight_global_scale): 16,577,986,560 bytes (15.44 GiB). - NVFP4 / MXFP4 stored bytes, measured on the pair: 1.0588 (format-implied 4.5 / 4.25 = 1.0588).
Accuracy reference
- The BF16 source model's mean NLL is 1.7730 nats/token,
teacher-forced on the benchmark's 64 NLL prompts of 512 tokens (sha256
dc20a83531303e1559a5eafe5d1f907cc77680363a3b6d05cd0438d8d8ffd55c), defined as vLLM'sprompt_logprobsdefines it. - This checkpoint's own NLL on those prompts is measured by the benchmark run, not by this card.
Serving
vllm serve ggamecrazy/Qwen3-32B-MXFP4-W4A4 --kv-cache-dtype bfloat16 --linear-backend flashinfer_cutedsl
Tested with vLLM v0.31.0 on B200 (SM100). The benchmark pins the kernel
with --linear-backend.
License
Apache 2.0, same as the base model.