gemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8 · Model Card
gemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8: Model Card
Written by Sleepyj, published under apache-2.0, revision 593a63f3f4f5, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.
This is an EXL3 build of llmfan46/gemma-4-31B-it-uncensored-heretic
for ExLlamaV3 and TabbyAPI. The vision tower is quantized to 6 bits. It loads only with ExLlamaV3 or TabbyAPI's exllamav3 loader,
not with Transformers, vLLM or llama.cpp.
This is not the coder3101 heretic. That 4.00bpw quant is
sjoe1244/gemma-4-31B-it-heretic-exl3-4.00bpw-h6.
On a 24 GB RTX 4090 this quant tops out at about 67K prompt with an 8-bit KV cache, or about 118K with a 4-bit KV cache. It does not fit 139,264 context at all. If you need 128K+ on one 24 GB card, use the 4.15bpw-h6 or 4.00bpw-h6 quant instead. See RTX 4090 (24 GB) limits below for the measured numbers.
Quantization
| Method | EXL3, ExLlamaV3 1.5.3 |
| Weights | 4.50 bpw (measured 4.50 bpw over the decoder layers) |
Head (lm_head) |
8 bits |
| Vision tower | 6 bits |
| Codebook | mul1, out_scales always |
| Calibration | 250 rows x 2048 cols (ExLlamaV3 default set) |
| Source | llmfan46/gemma-4-31B-it-uncensored-heretic, BF16 (Gemma4ForConditionalGeneration) |
| File | Bytes |
|---|---|
model-00001-of-00003.safetensors |
8,437,700,726 |
model-00002-of-00003.safetensors |
8,402,120,533 |
model-00003-of-00003.safetensors |
4,342,120,961 |
| Total | 21,181,942,220 (19.73 GiB) |
RTX 4090 (24 GB) limits
All measured on one RTX 4090 (24,564 MiB total) used as a server (the desktop on that card uses only a few MiB).
TabbyAPI with ExLlamaV3 1.5.3, max_batch_size 2, chunk_size 1024, vision: false. VRAM is the whole-GPU
nvidia-smi reading. "Largest pool that loads" was bisected in 2,048-token steps against TabbyAPI's
Insufficient VRAM in split check.
How much context fits
KV cache (cache_mode) |
Largest pool that loads | Pool that passed the stress test | Idle VRAM after load | Peak VRAM in the stress test | KL vs FP16 cache |
|---|---|---|---|---|---|
| 8,8 | 73,728 | 69,632 | 23,089 MiB | 24,009 MiB (48K + 18K prompts at once) | 0.053 |
| 4,8 (K4 / V8) | 92,160 | 83,968 | ~22.9 GB | 23.90 GB (63K + 18K at once) | 0.324 |
| 4,4 | 120,832 (122,880 fails) | 120,832 | ~22.9 GB | 23.9 GB (one 116K prompt) | 0.427 |
| 4,4 at 139,264 | does not load | - | - | - | - |
- The loader is not the real limit. At 4,8, a 92,160 pool loads but crashes with a CUDA out-of-memory error during prefill when two ~63K prompts run at once. Two long prompts at once needed about 0.9 GB over idle VRAM for prefill activations (8,8: 23,089 idle, 24,009 peak), so stay at the stress-tested sizes above.
- The KV cache costs more quality than the weights here. KL vs an FP16 cache, measured with ExLlamaV3's
eval/model_diff.py(20 x 2048 rows of WikiText-2):
| cache (K,V) | 8,8 | 4,8 | 8,4 | 4,4 | 4,3 | 3,4 | 3,3 |
|---|---|---|---|---|---|---|---|
| KL | 0.053 | 0.324 | 0.341 | 0.427 | 0.612 | 0.615 | 0.747 |
For comparison, the 4.50bpw weights cost 0.018 vs BF16. At 8 vs 4 bits Gemma 4 is more sensitive to V than to K,
so if you cut one half, use K4 / V8 rather than K8 / V4. At 3 bits that difference disappears: 4,3 and 3,4 both
land near 0.61, for only about 12% less cache memory than 4,4. We don't recommend either.
- A smaller chunk_size does not help. 512 reserves more VRAM in ExLlamaV3 than 1024 and fails to load at a
pool that loads with 1024. 2048 gives about 8% faster prefill, but a second request waits about twice as long
behind a long prefill.
- Vision was not tested on this quant. On our 3.0bpw 31B quant, turning vision on cost about 460 MiB at load
and about 730 MiB once images had been processed. Shrink the pool to match if you enable it.
- If your desktop or other apps also use this card, that VRAM comes out of the pool, so step down a size.
Speed (8,8 cache, pool 69,632, one request at a time)
| Prompt | Prefill | Decode |
|---|---|---|
| 154 tokens | - | 41.0 tok/s |
| 31K tokens | 19.4 s (~1,600 tok/s) | 37.5 tok/s |
| 58K tokens | 44.9 s (~1,300 tok/s) | 35.5 tok/s |
With two long prompts at once, combined prefill was about 1,000 tok/s (4,8 cache). At 4,4 a single 116K prompt prefilled at 913 tok/s.
Long uptime
After about 46 hours of uptime at the 8,8 / 69,632 setting, every prompt over about 30K tokens started failing
right away with torch.OutOfMemoryError (an 84 MiB allocation), while short prompts still worked. VRAM had crept
from 23,089 to about 24,050 MiB. Restarting TabbyAPI fixed it. We now set
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True in TabbyAPI's environment. With it, the stress test (three rounds
of 48K + 18K + a short prompt at once, plus a single 62K prompt) passed with VRAM steady at about 24,030 MiB. If
you run this close to the limit, set that variable and restart TabbyAPI if long prompts start failing.
Recommended TabbyAPI settings for a 4090
Best quality at about 67K prompt (what we run day to day):
model:
max_seq_len: 67584
cache_size: 69632
cache_mode: 8,8
chunk_size: 1024
max_batch_size: 2
vision: false
prompt_template: gemma4
memory:
sysmem_kv_cache: 8192 # optional: system-RAM tier, so evicted prefixes are restored instead of re-prefilled
For longer context, use cache_mode: 4,4 with cache_size: 120832 and max_seq_len up to 118784. This trades a
KV-cache KL of about 0.43 for roughly 1.7x the context.
The 139K test
With max_seq_len and cache_size 139264, cache_mode 4,4, max_batch_size 2 and chunk_size 2048, TabbyAPI stops
at load with Insufficient VRAM in split for model and cache. The weights are about 2 GiB larger than the 4.00bpw-h6
quant, which already peaks at 22.15 GiB at this context. For comparison, the coder3101 4.00bpw-h6 quant measured
20.59 GiB after load, 22.13 GiB peak and 45.6 tok/s on the same setup.
Quality vs the BF16 source
| Metric | This quant | BF16 |
|---|---|---|
| KL divergence vs BF16, WikiText-2 test, chat-templated (20 x 2048 rows) | 0.0181 (median 0.0029, p90 0.0303) | — |
| KL divergence vs BF16, self-sampled chat trace (~8.8K response tokens) | 0.0071 (p90 0.0147) | — |
| Perplexity, WikiText-2 test, chat-templated | 16.62 | 16.59 |
| Perplexity, chat trace | 1.190 | 1.199 |
These numbers come from ExLlamaV3's eval/qbench.py, with BF16 reference logits from the unquantized source and the same token IDs for every model.
KL divergence is the main quality metric here. At this bitrate, perplexity differences between the quants fall within run-to-run noise.
About raw perplexity: WikiText-2 perplexity scored as raw text with no chat framing is not meaningful for Gemma 4 instruct models. The BF16 source itself scores about 950 that way without a BOS token, and about 12,500 with one, so an earlier revision of this card that listed a perplexity of about 900 was removed. The chat trace was sampled from the 4.50bpw-h8 quant, which favours that quant slightly on the trace perplexity. The WikiText-2 rows have no such bias.
All three quants
| Quant | Size | KL vs BF16 (wiki2, chat-templated) | KL vs BF16 (chat trace) | PPL wiki2 (chat-templated) | VRAM after load | Peak VRAM @ ~135K | Decode tok/s | 139K ctx on 24 GB |
|---|---|---|---|---|---|---|---|---|
| 4.00bpw-h6 | 17.70 GiB | 0.0262 | 0.0099 | 16.65 | 20.59 GiB | 22.15 GiB | 46.0 | fits |
| 4.15bpw-h6 | 18.21 GiB | 0.0250 | 0.0087 | 16.58 | 21.08 GiB | 22.64 GiB | 44.6 | fits |
| 4.50bpw-h8 | 19.73 GiB | 0.0181 | 0.0071 | 16.62 | does not load at 139K | n/a | n/a | does not fit (max ~67K at 8,8, ~118K at 4,4) |
| BF16 (reference) | 57.2 GiB | — | — | 16.59 | — | — | — | — |
TabbyAPI settings used for the comparison table
The other two quants were measured with this config. This quant does not load with it (see above).
model:
max_seq_len: 139264
cache_size: 139264
cache_mode: 4,4
chunk_size: 2048
max_batch_size: 2
vision: false
prompt_template: gemma4
Credit and license
Licensed under Apache 2.0, matching google/gemma-4-31B-it and the source model. The decensoring credit belongs to
llmfan46. This repo only adds the EXL3 export, made with
ExLlamaV3.