Runs on SparkInfer, SGLang and vLLM unmodified — configs for all three are below. Serving many users at once? See concurrency. SparkInfer × this NVFP4 build × the DSpark v2 drafter — an engine, a checkpoint, and a speculative drafter optimized against each other, compounding to 4.3×. The drafter never changes what the model says: the target verifies every drafted token. GeForce RTX 5090–specific NVFP4 checkpoint of Qwen/Qwen3.8-27B, quantized with NVIDIA Model Optimizer. Serves the full native 262,144-token context on 32 GB. With the DSpark v2 drafter: 264.8 tok/s overall — up to 420 on code — on SparkInfer (its bench harness; the HTTP server is autoregressive-only today) and 161.7 tok/s on…
Open weights
apache-2.0
14.6B parameters
262,144 tokens
transformers
GLM-5.3-Flash with its routed experts quantized to NestQuant, a nested two-level format. Each expert has a 1.5-bit base plus a residual plane that lifts it to 4 bits. At serving time an expert moves from 1.5 bit to 4 bit by loading its residual plane on top of the base, and the base bytes stay the same. This build targets a 128 GB Mac with 96 GB usable for weights, KV cache and runtime. The native vision tower is included, so the model still takes images. Status: TODO. Encoding in progress. This repo cannot be loaded with stock transformers, vLLM or MLX. The serving kernel and loader will be published separately. - 67 of 288 experts per layer (23%) run at 4 bit. The rest run at 1.5 bit.…
Open weights
mit
16.9B parameters
1,048,576 tokens
Quantized version of https://huggingface.co/Qwen/Qwen3.8-27B
Open weights
apache-2.0
17.6B parameters
262,144 tokens
transformers
The RadixArk Qwen3.8-27B-NVFP4 model is the quantized version of Qwen/Qwen3.8-27B. The quantization was produced at RadixArk using NVIDIA Model Optimizer, following a mixed NVFP4 W4A4 recipe. Run on SGLang: launch command and per-platform recipes in the Qwen3.8-27B cookbook. This model is not owned or developed by RadixArk. It is a quantized derivative of Qwen's model; see the upstream Qwen3.8-27B model card for the source model's capabilities, training information, limitations, and license. Global Developers looking to deploy an off-the-shelf, pre-quantized model in AI agent systems, chatbots, RAG systems, and other AI-powered applications. Hugging Face 08/14/2026 via…
Open weights
apache-2.0
18.2B parameters
262,144 tokens
Model Optimizer
NVFP4 checkpoint of an abliterated Swift-Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: The Swift 1.5 version is source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support (tested on 0.29.0). No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Tested on an RTX 5090 (32 GB) with vLLM 0.29.0: NVFP4 layers on…
Open weights
other
18.2B parameters
262,144 tokens
vllm
NVFP4 checkpoint of an abliterated Swift 1.5 Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: (measured on the BF16 source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. The module split matches UkisAI's Swift 1.5 NVFP4 exactly. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support. No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Same format, recipe, module…
Open weights
other
18.2B parameters
262,144 tokens
vllm