SAVRN
Search Contact SAVRN

gemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8 · Model Card

gemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8: Model Card

Written by Sleepyj, published under apache-2.0, revision 593a63f3f4f5, read 2026-10-09. Shown as written; SAVRN's own facts about this model are on its page.

This is an EXL3 build of llmfan46/gemma-4-31B-it-uncensored-heretic for ExLlamaV3 and TabbyAPI. The vision tower is quantized to 6 bits. It loads only with ExLlamaV3 or TabbyAPI's exllamav3 loader, not with Transformers, vLLM or llama.cpp.

This is not the coder3101 heretic. That 4.00bpw quant is sjoe1244/gemma-4-31B-it-heretic-exl3-4.00bpw-h6.

On a 24 GB RTX 4090 this quant tops out at about 67K prompt with an 8-bit KV cache, or about 118K with a 4-bit KV cache. It does not fit 139,264 context at all. If you need 128K+ on one 24 GB card, use the 4.15bpw-h6 or 4.00bpw-h6 quant instead. See RTX 4090 (24 GB) limits below for the measured numbers.

Quantization

Method EXL3, ExLlamaV3 1.5.3
Weights 4.50 bpw (measured 4.50 bpw over the decoder layers)
Head (lm_head) 8 bits
Vision tower 6 bits
Codebook mul1, out_scales always
Calibration 250 rows x 2048 cols (ExLlamaV3 default set)
Source llmfan46/gemma-4-31B-it-uncensored-heretic, BF16 (Gemma4ForConditionalGeneration)
File Bytes
model-00001-of-00003.safetensors 8,437,700,726
model-00002-of-00003.safetensors 8,402,120,533
model-00003-of-00003.safetensors 4,342,120,961
Total 21,181,942,220 (19.73 GiB)

RTX 4090 (24 GB) limits

All measured on one RTX 4090 (24,564 MiB total) used as a server (the desktop on that card uses only a few MiB). TabbyAPI with ExLlamaV3 1.5.3, max_batch_size 2, chunk_size 1024, vision: false. VRAM is the whole-GPU nvidia-smi reading. "Largest pool that loads" was bisected in 2,048-token steps against TabbyAPI's Insufficient VRAM in split check.

How much context fits

KV cache (cache_mode) Largest pool that loads Pool that passed the stress test Idle VRAM after load Peak VRAM in the stress test KL vs FP16 cache
8,8 73,728 69,632 23,089 MiB 24,009 MiB (48K + 18K prompts at once) 0.053
4,8 (K4 / V8) 92,160 83,968 ~22.9 GB 23.90 GB (63K + 18K at once) 0.324
4,4 120,832 (122,880 fails) 120,832 ~22.9 GB 23.9 GB (one 116K prompt) 0.427
4,4 at 139,264 does not load - - - -
  • The loader is not the real limit. At 4,8, a 92,160 pool loads but crashes with a CUDA out-of-memory error during prefill when two ~63K prompts run at once. Two long prompts at once needed about 0.9 GB over idle VRAM for prefill activations (8,8: 23,089 idle, 24,009 peak), so stay at the stress-tested sizes above.
  • The KV cache costs more quality than the weights here. KL vs an FP16 cache, measured with ExLlamaV3's eval/model_diff.py (20 x 2048 rows of WikiText-2):
cache (K,V) 8,8 4,8 8,4 4,4 4,3 3,4 3,3
KL 0.053 0.324 0.341 0.427 0.612 0.615 0.747

For comparison, the 4.50bpw weights cost 0.018 vs BF16. At 8 vs 4 bits Gemma 4 is more sensitive to V than to K, so if you cut one half, use K4 / V8 rather than K8 / V4. At 3 bits that difference disappears: 4,3 and 3,4 both land near 0.61, for only about 12% less cache memory than 4,4. We don't recommend either. - A smaller chunk_size does not help. 512 reserves more VRAM in ExLlamaV3 than 1024 and fails to load at a pool that loads with 1024. 2048 gives about 8% faster prefill, but a second request waits about twice as long behind a long prefill. - Vision was not tested on this quant. On our 3.0bpw 31B quant, turning vision on cost about 460 MiB at load and about 730 MiB once images had been processed. Shrink the pool to match if you enable it. - If your desktop or other apps also use this card, that VRAM comes out of the pool, so step down a size.

Speed (8,8 cache, pool 69,632, one request at a time)

Prompt Prefill Decode
154 tokens - 41.0 tok/s
31K tokens 19.4 s (~1,600 tok/s) 37.5 tok/s
58K tokens 44.9 s (~1,300 tok/s) 35.5 tok/s

With two long prompts at once, combined prefill was about 1,000 tok/s (4,8 cache). At 4,4 a single 116K prompt prefilled at 913 tok/s.

Long uptime

After about 46 hours of uptime at the 8,8 / 69,632 setting, every prompt over about 30K tokens started failing right away with torch.OutOfMemoryError (an 84 MiB allocation), while short prompts still worked. VRAM had crept from 23,089 to about 24,050 MiB. Restarting TabbyAPI fixed it. We now set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True in TabbyAPI's environment. With it, the stress test (three rounds of 48K + 18K + a short prompt at once, plus a single 62K prompt) passed with VRAM steady at about 24,030 MiB. If you run this close to the limit, set that variable and restart TabbyAPI if long prompts start failing.

Recommended TabbyAPI settings for a 4090

Best quality at about 67K prompt (what we run day to day):

model:
  max_seq_len: 67584
  cache_size: 69632
  cache_mode: 8,8
  chunk_size: 1024
  max_batch_size: 2
  vision: false
  prompt_template: gemma4
memory:
  sysmem_kv_cache: 8192   # optional: system-RAM tier, so evicted prefixes are restored instead of re-prefilled

For longer context, use cache_mode: 4,4 with cache_size: 120832 and max_seq_len up to 118784. This trades a KV-cache KL of about 0.43 for roughly 1.7x the context.

The 139K test

With max_seq_len and cache_size 139264, cache_mode 4,4, max_batch_size 2 and chunk_size 2048, TabbyAPI stops at load with Insufficient VRAM in split for model and cache. The weights are about 2 GiB larger than the 4.00bpw-h6 quant, which already peaks at 22.15 GiB at this context. For comparison, the coder3101 4.00bpw-h6 quant measured 20.59 GiB after load, 22.13 GiB peak and 45.6 tok/s on the same setup.

Quality vs the BF16 source

Metric This quant BF16
KL divergence vs BF16, WikiText-2 test, chat-templated (20 x 2048 rows) 0.0181 (median 0.0029, p90 0.0303) —
KL divergence vs BF16, self-sampled chat trace (~8.8K response tokens) 0.0071 (p90 0.0147) —
Perplexity, WikiText-2 test, chat-templated 16.62 16.59
Perplexity, chat trace 1.190 1.199

These numbers come from ExLlamaV3's eval/qbench.py, with BF16 reference logits from the unquantized source and the same token IDs for every model. KL divergence is the main quality metric here. At this bitrate, perplexity differences between the quants fall within run-to-run noise.

About raw perplexity: WikiText-2 perplexity scored as raw text with no chat framing is not meaningful for Gemma 4 instruct models. The BF16 source itself scores about 950 that way without a BOS token, and about 12,500 with one, so an earlier revision of this card that listed a perplexity of about 900 was removed. The chat trace was sampled from the 4.50bpw-h8 quant, which favours that quant slightly on the trace perplexity. The WikiText-2 rows have no such bias.

All three quants

Quant Size KL vs BF16 (wiki2, chat-templated) KL vs BF16 (chat trace) PPL wiki2 (chat-templated) VRAM after load Peak VRAM @ ~135K Decode tok/s 139K ctx on 24 GB
4.00bpw-h6 17.70 GiB 0.0262 0.0099 16.65 20.59 GiB 22.15 GiB 46.0 fits
4.15bpw-h6 18.21 GiB 0.0250 0.0087 16.58 21.08 GiB 22.64 GiB 44.6 fits
4.50bpw-h8 19.73 GiB 0.0181 0.0071 16.62 does not load at 139K n/a n/a does not fit (max ~67K at 8,8, ~118K at 4,4)
BF16 (reference) 57.2 GiB — — 16.59 — — — —

TabbyAPI settings used for the comparison table

The other two quants were measured with this config. This quant does not load with it (see above).

model:
  max_seq_len: 139264
  cache_size: 139264
  cache_mode: 4,4
  chunk_size: 2048
  max_batch_size: 2
  vision: false
  prompt_template: gemma4

Credit and license

Licensed under Apache 2.0, matching google/gemma-4-31B-it and the source model. The decensoring credit belongs to llmfan46. This repo only adds the EXL3 export, made with ExLlamaV3.