SAVRN
Search Contact SAVRN

Qwen3-32B-NVFP4-W4A4 · Model Card

Qwen3-32B-NVFP4-W4A4: Model Card

Written by Cezar, published under apache-2.0, revision a274b18165ad, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.

Benchmark code, raw data and results: github.com/cezarc1/decodebench (see RESULTS.md).

NVFP4 W4A4 quantization of Qwen/Qwen3-32B (revision 9216db5781bf21249d130ec9da846c4624c16137, BF16). It was made for a controlled NVFP4-vs-MXFP4 decode benchmark on NVIDIA B200 (fp4bench), not as a general-purpose release: the NVFP4 and MXFP4 checkpoints share the model, the tool, the recipe, the calibration data and the layer coverage, and differ only in the format.

Weights NVFP4: FP4 E2M1 values in 16-element blocks, E4M3 block scales plus one FP32 per-tensor global scale, 4.5 bits per weight
Activations 4-bit, quantized per 16-element block at runtime, with a static per-tensor activation scale (input_global_scale) calibrated on the data below
Coverage all Linear layers except lm_head (448 modules); embeddings, norms and lm_head stay BF16
KV cache BF16 at serve time (--kv-cache-dtype bfloat16); the checkpoint has no kv_cache_scheme

How it was made

  • Tool: llm-compressor 0.14.0, compressed-tensors 0.19.0. Recipe: QuantizationModifier(targets="Linear", scheme="NVFP4", ignore=["lm_head"]). The plain preset recipe, with no weight-rounding optimization (GPTQ, AutoRound and the like).
  • Calibration: 256 prompts from the first user turn of ShareGPT conversations, anon8231489123/ShareGPT_Vicuna_unfiltered at revision 192ab2185289094fc556ec8ce5ce1e8e587154ca, at most 1024 tokens each (33208 tokens in all, seed 3). The data sets NVFP4's static per-tensor activation scale (input_global_scale). The MXFP4 checkpoint uses the same set.
  • The source weights are the BF16 original, so this is the only quantization step.
  • Quantized at 2026-10-05T00:01:03.284345+00:00.

Measured size

  • Quantized Linear weights stored in this checkpoint (weight_packed, weight_scale and, for NVFP4, weight_global_scale): 17,553,164,032 bytes (16.35 GiB).
  • NVFP4 / MXFP4 stored bytes, measured on the pair: 1.0588 (format-implied 4.5 / 4.25 = 1.0588).

Accuracy reference

  • The BF16 source model's mean NLL is 1.7730 nats/token, teacher-forced on the benchmark's 64 NLL prompts of 512 tokens (sha256 dc20a83531303e1559a5eafe5d1f907cc77680363a3b6d05cd0438d8d8ffd55c), defined as vLLM's prompt_logprobs defines it.
  • This checkpoint's own NLL on those prompts is measured by the benchmark run, not by this card.

Serving

vllm serve ggamecrazy/Qwen3-32B-NVFP4-W4A4 --kv-cache-dtype bfloat16 --linear-backend flashinfer_cutedsl

Tested with vLLM v0.31.0 on B200 (SM100). The benchmark pins the kernel with --linear-backend.

License

Apache 2.0, same as the base model.