Open-weight model
Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16
by Johannes Uusikuu Johneeee/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16
This model was quantized using oQ (oMLX v0.7.0.dev2) mixed-precision quantization.
Runs On
What it takes to serve Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16 (26.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 53.8 GB | 64.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 26.9 GB | 32.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 13.4 GB | 16.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
This model was quantized using oQ (oMLX v0.7.0.dev2) mixed-precision quantization.
Excerpt from the card by Johannes Uusikuu.
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 64
- Hidden size
- 5,120
- Feed-forward size
- 17,408
- Attention heads
- 24
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Stored precision
- float16
- Model type
- qwen3_5
Identity and Version
- Repository
- Johneeee/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16
- Publisher
- Johannes Uusikuu
- Task
- Not stated by the source
- Modality
- Other
- Library
- mlx
- Parameters
- 26.9B parameters
- Languages
- mlx, oq
- Revision
- ed065c4f9525e827f98ee833a2a0c4fdeeb1ca00
- First published
- 2026-09-18
- Last updated
- 2026-09-18
Files and Weights
14 files, 19.2 GB in total. The weights are 4 files totalling 19.2 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00004.safetensors | Weights | 5.0 GB | 9136b4998087 |
| model-00002-of-00004.safetensors | Weights | 5.0 GB | fe3df26d2123 |
| model-00003-of-00004.safetensors | Weights | 5.0 GB | 6571a4b4442f |
| model-00004-of-00004.safetensors | Weights | 4.1 GB | 9aac628b2de6 |
| config.json | Configuration | 15.0 KB | — |
| generation_config.json | Configuration | 213 B | — |
| model.safetensors.index.json | Configuration | 182.3 KB | — |
| oq_imatrix_report.json | Configuration | 30.9 KB | — |
| README.md | Documentation | 373 B | — |
| chat_template.jinja | Other | 17.1 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 0997f410c57a |
| tokenizer_config.json | Tokenizer | 8.7 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- Not stated by the source
- Access
- Open weights, no gate
- Download size
- 19.2 GB
Released by Johannes Uusikuu through its official repository on Hugging Face.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 19.2 GB |
| 16-bit | 53.8 GB |
| 8-bit | 26.9 GB |
| 4-bit | 13.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16
How much GPU memory does Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16 need?
About 64.6 GB at 16-bit and 16.1 GB at 4-bit: the weights (26.9B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What is Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-ULTRA-HERETIC-Uncensored-oQ5e-fp16's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.