GGUF quants of ajgazin/Swift-Qwen3.8-27B-Uncensored-MTP, an abliterated Swift-Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and version is size, with Unsloth's importance matrix. - MTP head included in every main GGUF, for self-speculative decoding in llama.cpp. Each quant is one file, Swift-Qwen3.8-27B-Uncensored-Dynamic-MTP-.gguf; the BF16 is Swift-Qwen3.8-27B-Uncensored-MTP-BF16.gguf. Each quant compared with the BF16 it was made from, token by token: wikitext-2 test set, 36 × 8192 tokens, f16 KV cache, llama.cpp 94659b076. Lower KLD and higher top-1 are better. - UD-IQ3S and UD-IQ4XS were added after this measurement run, so they have no numbers yet and…
Open weights
other
gguf
An abliterated Swift-Qwen3.8-27B, UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B. It applies the single-direction refusal ablation of (Arditi et al. 2024), with orcarouter's own direction, to Swift's weights. The vision tower is untouched and the MTP head is kept and edited consistently, so self-speculative decoding works. GGUF (llama.cpp, Unsloth-dynamic Q2 to Q8) and NVFP4 (vLLM, SGLang). The same edit on Swift 1.5 is All four rows are our measurements with Heretic's built-in evaluation (evaluatemodel, BF16): keyword-based refusal detector. the original model. - Thinking is closed immediately with a response prefix ("\n \n\n"), so answers are scored, not reasoning. - Refusal counts…
Open weights
other
27.8B parameters
262,144 tokens
transformers
NVFP4 checkpoint of an abliterated Swift-Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: The Swift 1.5 version is source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support (tested on 0.29.0). No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Tested on an RTX 5090 (32 GB) with vLLM 0.29.0: NVFP4 layers on…
Open weights
other
18.2B parameters
262,144 tokens
vllm
NVFP4 checkpoint of an abliterated Swift 1.5 Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and SGLang. GGUFs for llama.cpp: (measured on the BF16 source). - Swift's own NVFP4 recipe, unmodified, from ukisai/Swift-Qwen3.8-27B-NVFP4, calibrated with NVIDIA ModelOpt. The module split matches UkisAI's Swift 1.5 NVFP4 exactly. - MTP head and vision tower in BF16, bit-identical to the source. 21.9 GB, NVIDIA ModelOpt mixed-precision format. Needs a vLLM with ModelOpt mixed-precision support. No --quantization flag. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default. Same format, recipe, module…
Open weights
other
18.2B parameters
262,144 tokens
vllm
GGUF quants of ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP, an abliterated Swift 1.5 Qwen3.8-27B (UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B). For vLLM and size, with Unsloth's importance matrix. - MTP head included in every main GGUF, for self-speculative decoding in llama.cpp. Each quant is one file, Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-.gguf; the BF16 is Swift-1.5-Qwen3.8-27B-Uncensored-MTP-BF16.gguf. --spec-type draft-mtp needs a llama.cpp build with MTP support for qwen35. The MTP head loads from the main GGUF; there is no separate draft file. Sampling, as for Swift and Qwen: temperature 1.0, topp 0.95, topk 20, minp 0. The model thinks before answering by default.…
Open weights
other
gguf