Qwen3-32B-NVFP4-W4A4 · Model Card
Qwen3-32B-NVFP4-W4A4: Model Card
Written by Cezar, published under apache-2.0, revision a274b18165ad, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.
Benchmark code, raw data and results: github.com/cezarc1/decodebench (see RESULTS.md).
NVFP4 W4A4 quantization of
Qwen/Qwen3-32B
(revision 9216db5781bf21249d130ec9da846c4624c16137, BF16). It was made for a controlled NVFP4-vs-MXFP4 decode
benchmark on NVIDIA B200 (fp4bench), not as a general-purpose release: the NVFP4 and MXFP4
checkpoints share the model, the tool, the recipe, the calibration data and the layer coverage,
and differ only in the format.
| Weights | NVFP4: FP4 E2M1 values in 16-element blocks, E4M3 block scales plus one FP32 per-tensor global scale, 4.5 bits per weight |
| Activations | 4-bit, quantized per 16-element block at runtime, with a static per-tensor activation scale (input_global_scale) calibrated on the data below |
| Coverage | all Linear layers except lm_head (448 modules); embeddings, norms and lm_head stay BF16 |
| KV cache | BF16 at serve time (--kv-cache-dtype bfloat16); the checkpoint has no kv_cache_scheme |
How it was made
- Tool: llm-compressor 0.14.0, compressed-tensors 0.19.0.
Recipe:
QuantizationModifier(targets="Linear", scheme="NVFP4", ignore=["lm_head"]). The plain preset recipe, with no weight-rounding optimization (GPTQ, AutoRound and the like). - Calibration: 256 prompts from the first user turn of ShareGPT conversations,
anon8231489123/ShareGPT_Vicuna_unfilteredat revision192ab2185289094fc556ec8ce5ce1e8e587154ca, at most 1024 tokens each (33208 tokens in all, seed 3). The data sets NVFP4's static per-tensor activation scale (input_global_scale). The MXFP4 checkpoint uses the same set. - The source weights are the BF16 original, so this is the only quantization step.
- Quantized at 2026-10-05T00:01:03.284345+00:00.
Measured size
- Quantized Linear weights stored in this checkpoint (
weight_packed,weight_scaleand, for NVFP4,weight_global_scale): 17,553,164,032 bytes (16.35 GiB). - NVFP4 / MXFP4 stored bytes, measured on the pair: 1.0588 (format-implied 4.5 / 4.25 = 1.0588).
Accuracy reference
- The BF16 source model's mean NLL is 1.7730 nats/token,
teacher-forced on the benchmark's 64 NLL prompts of 512 tokens (sha256
dc20a83531303e1559a5eafe5d1f907cc77680363a3b6d05cd0438d8d8ffd55c), defined as vLLM'sprompt_logprobsdefines it. - This checkpoint's own NLL on those prompts is measured by the benchmark run, not by this card.
Serving
vllm serve ggamecrazy/Qwen3-32B-NVFP4-W4A4 --kv-cache-dtype bfloat16 --linear-backend flashinfer_cutedsl
Tested with vLLM v0.31.0 on B200 (SM100). The benchmark pins the kernel
with --linear-backend.
License
Apache 2.0, same as the base model.