This model was quantized using oQ (oMLX v0.6.3rc3) mixed-precision quantization.
Runs On
What it takes to serve Qwen3.8-Flash-Next-oQ4e-mtp (180B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 360.0 GB | 432.0 GB | 2x MI325X (256 GB) Vultr |
$4.00 | 2x MI355X $5.18 · 3x MI300X $5.55 |
| 8-bit | 180.0 GB | 216.0 GB | 1x MI325X (256 GB) Vultr |
$2.00 | 1x MI355X $2.59 · 2x MI300X $3.70 |
| 4-bit | 90.0 GB | 108.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x MI325X $2.00 · 1x MI355X $2.59 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
This model was quantized using oQ (oMLX v0.6.3rc3) mixed-precision quantization.
Excerpt from the card by Robot Haus.
Configuration
- Architecture
- Qwen4ExpForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 48
- Hidden size
- 2,560
- Attention heads
- 24
- Key/value heads
- 2
- Head dimension
- 256
- Vocabulary size
- 248,320
- Experts
- 512
- Experts active per token
- 10
- Model type
- qwen4_exp
Identity and Version
- Repository
- Robot-Haus/Qwen3.8-Flash-Next-oQ4e-mtp
- Publisher
- Robot Haus
- Task
- Not stated by the source
- Modality
- Other
- Library
- mlx
- Parameters
- 180B parameters
- Languages
- mlx, oq
- Revision
- 33049b79e87fe16a6b32700e7f3883e966af5da8
- First published
- 2026-09-18
- Last updated
- 2026-09-18
Files and Weights
33 files, 106.3 GB in total. The weights are 21 files totalling 106.3 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00021.safetensors | Weights | 5.2 GB | a6cf3bfbe6fb |
| model-00002-of-00021.safetensors | Weights | 5.0 GB | db06472d04ba |
| model-00003-of-00021.safetensors | Weights | 5.0 GB | 2b9227ee8660 |
| model-00004-of-00021.safetensors | Weights | 5.0 GB | 4915e5847b9c |
| model-00005-of-00021.safetensors | Weights | 5.0 GB | 750fa33d4187 |
| model-00006-of-00021.safetensors | Weights | 5.0 GB | c030f56b7c18 |
| model-00007-of-00021.safetensors | Weights | 5.0 GB | 0f3e15d9ecdf |
| model-00008-of-00021.safetensors | Weights | 5.1 GB | b355190385c3 |
| model-00009-of-00021.safetensors | Weights | 5.2 GB | 56590b6cf020 |
| model-00010-of-00021.safetensors | Weights | 5.2 GB | b294ba05bf87 |
| model-00011-of-00021.safetensors | Weights | 5.2 GB | c5ba00b14125 |
| model-00012-of-00021.safetensors | Weights | 5.2 GB | 9d10275984cd |
| model-00013-of-00021.safetensors | Weights | 5.2 GB | 2a894851d5da |
| model-00014-of-00021.safetensors | Weights | 5.2 GB | ddc7e07e1f65 |
| model-00015-of-00021.safetensors | Weights | 5.2 GB | 2685967ee4cf |
| model-00016-of-00021.safetensors | Weights | 5.2 GB | e804c5fa0973 |
| model-00017-of-00021.safetensors | Weights | 5.2 GB | 36b1a9604392 |
| model-00018-of-00021.safetensors | Weights | 5.2 GB | 390935475275 |
| model-00019-of-00021.safetensors | Weights | 5.2 GB | b8d136f5780a |
| model-00020-of-00021.safetensors | Weights | 5.2 GB | d5b16c8abcba |
| model-00021-of-00021.safetensors | Weights | 3.8 GB | e4339a78fe23 |
| config.json | Configuration | 183.2 KB | — |
| generation_config.json | Configuration | 202 B | — |
| model.safetensors.index.json | Configuration | 405.2 KB | — |
| oq_imatrix_report.json | Configuration | 67.1 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| README.md | Documentation | 321 B | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 0997f410c57a |
| tokenizer_config.json | Tokenizer | 17.9 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- Not stated by the source
- Access
- Open weights, no gate
- Download size
- 106.3 GB
Released by Robot Haus through its official repository on Hugging Face.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 106.3 GB |
| 16-bit | 360.0 GB |
| 8-bit | 180.0 GB |
| 4-bit | 90.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Qwen3.8-Flash-Next-oQ4e-mtp
How much GPU memory does Qwen3.8-Flash-Next-oQ4e-mtp need?
About 432 GB at 16-bit and 108 GB at 4-bit: the weights (180B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Qwen3.8-Flash-Next-oQ4e-mtp on?
At 16-bit, 2x MI325X from $4.00 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What is Qwen3.8-Flash-Next-oQ4e-mtp's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.