SAVRN
Search Contact SAVRN

Open-weight model

Qwen3.8-Flash-Next-oQ4e-fp16-mtp

by Robot Haus Robot-Haus/Qwen3.8-Flash-Next-oQ4e-fp16-mtp

Cloned from monroewilliams/Qwen3.8-Flash-Next-oQ4e-fp16-mtp This model was converted from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp using this script. It was not requantized, just processed to convert all bf16 parts to fp16, for better performance on M1/M2 machines.

Parameters
Context262,144
Weights23.2 MB
License
AccessOpen weights
Monthly Downloads378

Model Card

Cloned from monroewilliams/Qwen3.8-Flash-Next-oQ4e-fp16-mtp This model was converted from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp using this script. It was not requantized, just processed to convert all bf16 parts to fp16, for better performance on M1/M2 machines. An FP16 conversion of Vontra/Qwen3.8-Flash-Next-MLX-oQ3-MTP, reprocessed with omlx-fp16-clone. Weights removed after benchmarking showed no performance benefit on M1 Ultra. FP16 was not faster than BF16 on M1 Ultra. The results show a clear split: FP16 prefill is significantly faster, but FP16 decode is slower and memory usage is higher. The omlx-fp16-clone script must promote all vision/audio passthrough tensors from BF16 to FP32 (an…

Excerpt from the card by Robot Haus.

Configuration

Architecture
Qwen4ExpForConditionalGeneration
Context length (tokens)
262,144
Layers
48
Hidden size
2,560
Attention heads
24
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
512
Experts active per token
10
Model type
qwen4_exp

Identity and Version

Repository
Robot-Haus/Qwen3.8-Flash-Next-oQ4e-fp16-mtp
Publisher
Robot Haus
Task
Not stated by the source
Modality
Other
Library
mlx
Parameters
Not stated by the source
Languages
mlx, oq
Revision
e6be1266d5e90e2da09e1e7ecabad3ddd4a1629f
First published
2026-08-30
Last updated
2026-09-18

Files and Weights

11 files, 23.2 MB in total.

Configuration4 files · 250.8 KB
Tokenizer4 files · 22.9 MB
Documentation1 file · 4.4 KB
Other1 file · 9.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
config.jsonConfiguration183.2 KB
generation_config.jsonConfiguration202 B
oq_imatrix_report.jsonConfiguration67.1 KB
preprocessor_config.jsonConfiguration390 B
README.mdDocumentation4.4 KB
chat_template.jinjaOther9.0 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer3.4 MB
tokenizer.jsonTokenizer12.8 MB 0997f410c57a
tokenizer_config.jsonTokenizer17.9 KB
vocab.jsonTokenizer6.7 MB

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download from Robot Haus

Released by Robot Haus through its official repository on Hugging Face.

Built From

  • Derived from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp
  • Quantized from Jundot/Qwen3.8-Flash-Next-oQ4e-mtp

Questions About Qwen3.8-Flash-Next-oQ4e-fp16-mtp

What is Qwen3.8-Flash-Next-oQ4e-fp16-mtp's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.