Open-weight model
Qwen3.8-Flash-Next-oQ4e-fp16-mtp
by Robot Haus Robot-Haus/Qwen3.8-Flash-Next-oQ4e-fp16-mtp
Cloned from monroewilliams/Qwen3.8-Flash-Next-oQ4e-fp16-mtp This model was converted from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp using this script. It was not requantized, just processed to convert all bf16 parts to fp16, for better performance on M1/M2 machines.
Model Card
Cloned from monroewilliams/Qwen3.8-Flash-Next-oQ4e-fp16-mtp This model was converted from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp using this script. It was not requantized, just processed to convert all bf16 parts to fp16, for better performance on M1/M2 machines. An FP16 conversion of Vontra/Qwen3.8-Flash-Next-MLX-oQ3-MTP, reprocessed with omlx-fp16-clone. Weights removed after benchmarking showed no performance benefit on M1 Ultra. FP16 was not faster than BF16 on M1 Ultra. The results show a clear split: FP16 prefill is significantly faster, but FP16 decode is slower and memory usage is higher. The omlx-fp16-clone script must promote all vision/audio passthrough tensors from BF16 to FP32 (an…
Excerpt from the card by Robot Haus.
Configuration
- Architecture
- Qwen4ExpForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 48
- Hidden size
- 2,560
- Attention heads
- 24
- Key/value heads
- 2
- Head dimension
- 256
- Vocabulary size
- 248,320
- Experts
- 512
- Experts active per token
- 10
- Model type
- qwen4_exp
Identity and Version
- Repository
- Robot-Haus/Qwen3.8-Flash-Next-oQ4e-fp16-mtp
- Publisher
- Robot Haus
- Task
- Not stated by the source
- Modality
- Other
- Library
- mlx
- Parameters
- Not stated by the source
- Languages
- mlx, oq
- Revision
- e6be1266d5e90e2da09e1e7ecabad3ddd4a1629f
- First published
- 2026-08-30
- Last updated
- 2026-09-18
Files and Weights
11 files, 23.2 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| config.json | Configuration | 183.2 KB | — |
| generation_config.json | Configuration | 202 B | — |
| oq_imatrix_report.json | Configuration | 67.1 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| README.md | Documentation | 4.4 KB | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 0997f410c57a |
| tokenizer_config.json | Tokenizer | 17.9 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- Not stated by the source
- Access
- Open weights, no gate
Released by Robot Haus through its official repository on Hugging Face.
Built From
- Derived from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp
- Quantized from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp
Questions About Qwen3.8-Flash-Next-oQ4e-fp16-mtp
What is Qwen3.8-Flash-Next-oQ4e-fp16-mtp's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.