An experimental MXFP4 quantization of ibm-granite/granite-4.2-3b, made with GPTQ calibration instead of the usual data-free rounding, and provided in two formats with identical weights. It's published as a data point, not as the recommended build. For most users, the Q4KM GGUF or GPTQ int4 build is the better choice. Format: OCP MX — E2M1 (FP4) weights, groupsize=32, one UE8M0 shared scale per group; weights only, activations stay 16-bit. Embeddings, norms and lmhead are unquantized. MXFP4 is usually produced data-free, by rounding weights straight to FP4 with no calibration data. llm-compressor ships an MXFP4A16 scheme, but every example we found pairs GPTQModifier with integer schemes…
Open weights
apache-2.0
3.7B parameters
131,072 tokens
transformers
A 4-bit weight-only GPTQ quantization of ibm-granite/granite-4.2-3b, saved in compressed-tensors format for vLLM. On an AMD Radeon Pro V620 (RDNA2), vLLM served it at 65 tok/s single-sequence and ~1,000 tok/s aggregate at 16 concurrent sequences. llm-compressor GPTQModifier: - int4, symmetric, weight-only (activations stay 16-bit) - groupsize=32, actorder="weight" - lmhead left unquantized 512 samples, maxseqlength=512 The tuning mattered. With the default W4A16 preset (groupsize=128), the same pipeline scored 0.153 mean KLD. Moving to groupsize=32 with actorder="weight" cut that by about 28%, for roughly 0.2 GB more disk. For comparison, an AWQ build (W4A16 asymmetric, groupsize=128) from…
Open weights
apache-2.0
3.7B parameters
131,072 tokens
transformers
A 4-bit GGUF of ibm-granite/granite-4.2-3b, quantized with an importance matrix built from a mixed prose and tool-calling calibration corpus. Of the three quantizations we made of this model, it's the closest to the original and the smallest. "Same top token" is llama.cpp's Same top p statistic: the share of positions where the quantized model's most likely next token matches the f16 model's. 1. Convert: converthftogguf.py (mainline llama.cpp), bf16 safetensors → f16 GGUF. Granite's scaling parameters (attentionscale, embeddingscale, residualscale, logitsscale) are supported natively. 2. Importance matrix: llama-imatrix over prose plus tool-calling conversations, rendered through Granite's…
Open weights
apache-2.0