SAVRN
Search Contact SAVRN

Qwen3-32B-MXFP4-W4A4 · Model Card

Qwen3-32B-MXFP4-W4A4: Model Card

Written by Cezar, published under apache-2.0, revision a49f042511ac, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.

Benchmark code, raw data and results: github.com/cezarc1/decodebench (see RESULTS.md).

MXFP4 W4A4 quantization of Qwen/Qwen3-32B (revision 9216db5781bf21249d130ec9da846c4624c16137, BF16). It was made for a controlled NVFP4-vs-MXFP4 decode benchmark on NVIDIA B200 (fp4bench), not as a general-purpose release: the NVFP4 and MXFP4 checkpoints share the model, the tool, the recipe, the calibration data and the layer coverage, and differ only in the format.

Weights MXFP4: FP4 E2M1 values in 32-element blocks, E8M0 (power-of-two) block scales, 4.25 bits per weight
Activations 4-bit, quantized dynamically per 32-element block at runtime; this format has no static activation scales, so none are stored
Coverage all Linear layers except lm_head (448 modules); embeddings, norms and lm_head stay BF16
KV cache BF16 at serve time (--kv-cache-dtype bfloat16); the checkpoint has no kv_cache_scheme

How it was made

  • Tool: llm-compressor 0.14.0, compressed-tensors 0.19.0. Recipe: QuantizationModifier(targets="Linear", scheme="MXFP4", ignore=["lm_head"]). The plain preset recipe, with no weight-rounding optimization (GPTQ, AutoRound and the like).
  • Calibration: 256 prompts from the first user turn of ShareGPT conversations, anon8231489123/ShareGPT_Vicuna_unfiltered at revision 192ab2185289094fc556ec8ce5ce1e8e587154ca, at most 1024 tokens each (33208 tokens in all, seed 3). MXFP4 has no static activation scales, so the data does not change this checkpoint's weights. The NVFP4 checkpoint uses the same set.
  • The source weights are the BF16 original, so this is the only quantization step.
  • Quantized at 2026-10-04T23:56:31.271659+00:00.

Measured size

  • Quantized Linear weights stored in this checkpoint (weight_packed, weight_scale and, for NVFP4, weight_global_scale): 16,577,986,560 bytes (15.44 GiB).
  • NVFP4 / MXFP4 stored bytes, measured on the pair: 1.0588 (format-implied 4.5 / 4.25 = 1.0588).

Accuracy reference

  • The BF16 source model's mean NLL is 1.7730 nats/token, teacher-forced on the benchmark's 64 NLL prompts of 512 tokens (sha256 dc20a83531303e1559a5eafe5d1f907cc77680363a3b6d05cd0438d8d8ffd55c), defined as vLLM's prompt_logprobs defines it.
  • This checkpoint's own NLL on those prompts is measured by the benchmark run, not by this card.

Serving

vllm serve ggamecrazy/Qwen3-32B-MXFP4-W4A4 --kv-cache-dtype bfloat16 --linear-backend flashinfer_cutedsl

Tested with vLLM v0.31.0 on B200 (SM100). The benchmark pins the kernel with --linear-backend.

License

Apache 2.0, same as the base model.