# GSQ: Highly-Accurate Low-Precision Scalar Quantization for
Source: https://savrn.com/papers/gsq-highly-accurate-low-precision-scalar-quantization-for-llms-via-gumbel-softmax-sampling
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Abstract

Weight quantization has become a standard tool for efficient LLM deployment, especially for local inference, where models are now routinely served at 2-3 bits per parameter. The state of the art is currently split into two sets of methods: simple scalar quantization techniques, such as GPTQ or AWQ, which are widely deployed but plateau in accuracy at 3-4 bits per parameter (bpp), and "second-generation" vector- or trellis-quantized methods, such as QTIP, GPTVQ and AQLM, which push the accuracy frontier at low bit-widths but are notoriously hard to implement and to scale, and have gained relatively less traction. In this paper, we ask whether this gap is fundamental, or whether a carefully optimized scalar quantizer can recover most of it. We answer in the affirmative, by introducing GSQ (Gumbel-Softmax Quantization), a post-training scalar quantization method which jointly learns the per-coordinate grid assignments and the per-group scales using a Gumbel-Softmax relaxation of the discrete grid. GSQ matches the cardinality of the relaxation to the small number of levels available in the target bit-width regime (e.g., 3-8 levels for ternary and 3 bpp, respectively), making the relaxation tight and the optimization tractable. Practically, on the standard Llama-3.1-8B/70B-Instruct models, GSQ closes most of the gap between scalar quantization and the QTIP frontier at 2 and 3 bits, while using a symmetric scalar grid with group-wise quantization, and thus fully compatible with existing scalar inference kernels. We further show that GSQ scales to trillion-scale Mixture-of-Experts models such as Kimi-K2.5, where vector-quantized methods are difficult to apply.

[Full paper on arXiv](https://arxiv.org/abs/2604.18556) · [Code](https://github.com/https://github.com/IST-DASLab/GSQ)

## Details

arXiv identifier

2604.18556

Published

2026-04-20

Authors

Alireza Dadgarnia, Soroush Tabesh, Mahdi Nikdan, Michael Helcig, Eldar Kurtic, Dan Alistarh

## Open Models Built on This Paper

Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.

| Model | Task | Size | License | Monthly downloads | Cheapest setup at 16-bit |
| --- | --- | --- | --- | --- | --- |
| [Qwen3.8-27B-GSQ-RCO-GGUF](https://savrn.com/models/qwen3-8-27b-gsq-rco-gguf) IST Austria Distributed Algorithms and Systems Lab | [Image and text to text](https://savrn.com/models/tasks/image-and-text-to-text) | — | apache-2.0 | 1.7M | — |
| [Dirk-Qwen3.8-27B-GGUF](https://savrn.com/models/dirk-qwen3-8-27b-gguf) Saga | [Image and text to text](https://savrn.com/models/tasks/image-and-text-to-text) | — | apache-2.0 | 1.3M | — |
| [Qwen3.8-Flash-Next-GSQ-RCO-GGUF](https://savrn.com/models/qwen3-8-flash-next-gsq-rco-gguf) IST Austria Distributed Algorithms and Systems Lab | [Image and text to text](https://savrn.com/models/tasks/image-and-text-to-text) | — | apache-2.0 | 1.1M | — |
| [GLM-5.3-Flash-GSQ-RCO-3.0bit-Q4Kattn-GGUF](https://savrn.com/models/glm-5-3-flash-gsq-rco-3-0bit-q4kattn-gguf) Neuralll | — | — | mit | — | — |

## Explore More

- [All research papers](https://savrn.com/papers)
- [The model directory](https://savrn.com/models)
- [Datasets](https://savrn.com/datasets)

## Source

- arXiv metadata, read 2026-10-03.
- [How the hub is built](https://savrn.com/model-hub/methodology)
