# Sleepyj: Open-Weight Models and Datasets
Source: https://savrn.com/model-publishers/sjoe1244
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Models

Model

### [gemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8](https://savrn.com/models/gemma-4-31b-it-uncensored-heretic-exl3-4-50bpw-h8)

[Sleepyj](https://savrn.com/model-publishers/sjoe1244)

This is an EXL3 build of llmfan46/gemma-4-31B-it-uncensored-heretic for ExLlamaV3 and TabbyAPI. The vision tower is quantized to 6 bits. It loads only with ExLlamaV3 or TabbyAPI's exllamav3 loader, not with Transformers, vLLM or llama.cpp. This is not the coder3101 heretic. That 4.00bpw quant is On a 24 GB RTX 4090 this quant tops out at about 67K prompt with an 8-bit KV cache, or about 118K with a 4-bit KV cache. It does not fit 139,264 context at all. If you need 128K+ on one 24 GB card, use the 4.00bpw-h6 quant instead. See RTX 4090 (24 GB) limits below for the measured numbers. All measured on one RTX 4090 (24,564 MiB total) used as a server (the desktop on that card uses only a few…

Open weights apache-2.0 10.6B parameters 262,144 tokens exllamav3

[View model](https://savrn.com/models/gemma-4-31b-it-uncensored-heretic-exl3-4-50bpw-h8)

## Explore More

- [All model publishers](https://savrn.com/model-publishers)
- [The model directory](https://savrn.com/models)
- [The dataset directory](https://savrn.com/datasets)

## Source

- Listed from their public repositories, read 2026-10-09.
- [Hugging Face profile](https://huggingface.co/sjoe1244)
- [How the hub is built](https://savrn.com/model-hub/methodology)
