SAVRN
Search Contact SAVRN

Independent publisher

Sleepyj

sjoe1244

Models in Library1
Datasets in Library0
Models on Hugging Face15
Followers1

Models

This is an EXL3 build of llmfan46/gemma-4-31B-it-uncensored-heretic for ExLlamaV3 and TabbyAPI. The vision tower is quantized to 6 bits. It loads only with ExLlamaV3 or TabbyAPI's exllamav3 loader, not with Transformers, vLLM or llama.cpp. This is not the coder3101 heretic. That 4.00bpw quant is On a 24 GB RTX 4090 this quant tops out at about 67K prompt with an 8-bit KV cache, or about 118K with a 4-bit KV cache. It does not fit 139,264 context at all. If you need 128K+ on one 24 GB card, use the 4.00bpw-h6 quant instead. See RTX 4090 (24 GB) limits below for the measured numbers. All measured on one RTX 4090 (24,564 MiB total) used as a server (the desktop on that card uses only a few…

Open weights apache-2.0 10.6B parameters 262,144 tokens exllamav3