SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Dirk-Qwen3.8-27B-GGUF

by Saga peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF

Dirk-Qwen3.8-27B-GGUF is an open-weight model for image and text to text from Saga, released under Apache License 2.0. Its published files total 239.7 GB. It draws 1.3M downloads a month.

Dirk is the Qwen3.8-27B that gets straight to the point. With our Sharp chat template, MTP, and vision baked in, the model answers lean and stays on-task out of the box. No template wrangling: download, point llama.cpp at it, go.

Parameters—
Context—
Weights239.7 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads1.3M

Model Card

By Saga, published under apache-2.0, revision 52cb3e759635.

Dirk is the Qwen3.8-27B that gets straight to the point. With our Sharp chat template, MTP, and vision baked in, the model answers lean and stays on-task out of the box. No template wrangling: download, point llama.cpp at it, go. If you want it to think deeper, set the effort level through chattemplatekwargs: Levels: low, medium, xhigh — high is accepted but is an alias for xhigh, not a step below it. Omit it for Dirk's lean default (medium). Turn thinking off entirely with "enablethinking": false. Dynamic 3.0 UD quants — their current generation, not an older ladder. Below that, GSQ-RCO quants from IST-DASLab, which hold up substantially better at 2–3 bpw (see the note under the file…

Read Saga's full model card

Dirk is the Qwen3.8-27B that gets straight to the point.

With our Sharp chat template, MTP, and vision baked in, the model answers lean and stays on-task out of the box. No template wrangling: download, point llama.cpp at it, go. If you want it to think deeper, set the effort level through chat_template_kwargs:

{"messages": [...], "chat_template_kwargs": {"reasoning_effort": "high"}}

Levels: low, medium, xhigh — high is accepted but is an alias for xhigh, not a step below it. Omit it for Dirk's lean default (medium). Turn thinking off entirely with "enable_thinking": false.

What it is

  • Base: Qwen/Qwen3.8-27B, a dense 27B vision-language model (vision preserved).
  • Quant: two quantizers, each where it is strongest. From 3 bpw up, Unsloth's Dynamic 3.0 UD quants — their current generation, not an older ladder. Below that, GSQ-RCO quants from IST-DASLab, which hold up substantially better at 2–3 bpw (see the note under the file table). Every tier keeps the model's MTP (nextn) head: runtimes with multi-token-prediction speculative decoding can use it for faster generation.
  • Template: the Sharp chat template (Qwen 3.8-aware) — froggeric's fixed Qwen template plus a terseness system prompt, and turning off the xhigh thinking default. Every tier supports the terseness opt-out (chat_template_kwargs: {"terse": false}). The GSQ-RCO- tiers carry v22.4.1, which also stands down when the runtime injects its own tool protocol (an LM Studio fix); the UD- tiers carry v22.4.0 and are otherwise identical — they pick up v22.4.1 on the next pass. It is byte-swapped into the GGUF metadata; the weights and the MTP tensors are untouched.

The only thing Dirk changes versus the stock quant is the template. Same weights, asked better.

Proven on Nail and Dagger

Dirk is new, but the template is not. The identical terseness edit, measured on Dagger's base (ThinkingCap-27B, same weights, only the template swapped):

stock template Sharp template change
Claw-Eval, answer component 59.3 66.7 +7.4
Claw-Eval answer tokens 5393 2217 −59%
MMLU-Pro tokens per correct answer 1601 1248 −22%

Roughly: the same answers in a bit over half the words, with accuracy moving up. That is what Dirk inherits — and its own SWE-bench-Live and MMLU-Pro numbers, shown above, bear it out.

Thinking effort

Stock Qwen3.8-27B forces reasoning_effort=xhigh on every call — always-on maximum-effort reasoning. Dirk removes that default, so it runs at the model's native medium effort: in both the official and Unsloth templates, medium is the setting that injects no reasoning instruction (only xhigh and low add one), and Dirk simply leaves it there. So Dirk thinks at the baseline and answers terse, instead of being pushed to the ceiling on every request. Set reasoning_effort yourself (low, medium, xhigh; high maps to xhigh), per request, through chat_template_kwargs — the OpenAI-style top-level reasoning_effort field is dropped by llama.cpp and oMLX, so it must go there (see the JSON example above).

Run it

file size notes
Dirk-Qwen3.8-27B-GSQ-RCO-IQ2_XS.gguf 8.8 GB smallest tier — the 12 GB card pick, with room for real context
Dirk-Qwen3.8-27B-UD-Q2_K_XL.gguf 9.8 GB 2-bit UD; kept for continuity — prefer GSQ-RCO-IQ2_S just below it, which is smaller and better
Dirk-Qwen3.8-27B-GSQ-RCO-IQ2_S.gguf 9.6 GB the 12 GB pick — 2-bit that still tracks the base model closely
Dirk-Qwen3.8-27B-GSQ-RCO-IQ3_XXS.gguf 10.4 GB fits 16 GB with room to spare, and a 12 GB card at shorter context
Dirk-Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf 12.1 GB 3-bit at near-base quality; the value pick if 16 GB is your ceiling
Dirk-Qwen3.8-27B-UD-Q3_K_XL.gguf 13.1 GB 3-bit with headroom to spare on 16 GB; prefer IQ4_XS below unless you need the extra ~1 GB for context
Dirk-Qwen3.8-27B-UD-IQ4_XS.gguf 14.3 GB the 16 GB pick — 4-bit quality with room for real context, where Q4_K_S leaves almost none
Dirk-Qwen3.8-27B-UD-Q4_K_S.gguf 15.4 GB tight 4-bit; useful when Q4_K_XL will not fit alongside your context
Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf 17.6 GB start here — the 24 GB-card default; best size/quality balance
Dirk-Qwen3.8-27B-UD-Q5_K_XL.gguf 20.9 GB the recommended 24 GB pick — dynamic + imatrix-calibrated, and small enough to leave real room for context
Dirk-Qwen3.8-27B-UD-Q6_K.gguf 22.0 GB 6-bit — the largest that still fits 24 GB, with tighter headroom than UD-Q5_K_XL
Dirk-Qwen3.8-27B-UD-Q6_K_XL.gguf 25.3 GB near-max quality; wants ~32 GB
Dirk-Qwen3.8-27B-UD-Q8_K_L.gguf 28.0 GB 8-bit, near-lossless — fits 48 GB with room for 256k context; a touch leaner than Q8_K_XL
Dirk-Qwen3.8-27B-UD-Q8_K_XL.gguf 31.5 GB 8-bit, effectively lossless

Every file carries the Sharp template and the MTP (nextn) head, and all share mmproj-F16.gguf for vision — you need only one copy of it.

Why two quantizers. UD- tiers are Unsloth Dynamic 3.0. GSQ-RCO- tiers come from IST-DASLab — Alistarh's lab, the GPTQ group — and are built by a genuinely different method: GSQ (arXiv) learns each tensor's quantization grid through a Gumbel-Softmax relaxation instead of rounding to it, and RCO (arXiv) then picks a per-tensor quantization type under an exact size budget by gradient descent on the task loss, rather than from a hand-tuned table. Below ~3 bpw that buys a lot: at a matched 8.4 GB, ISTA measure it well ahead of the equivalent UD file on wikitext perplexity and on AIME25 / GPQA-Diamond / LiveCodeBench v6. By ~3.5 bpw the two methods converge to within noise, which is exactly why the ladder switches over at 3 bpw and stays on UD above it. Those are ISTA's measurements, not ours — we have re-templated their files, not re-benchmarked them.

Let llama.cpp fetch it — pass a :quant tag from the table (:Q4_K_XL, :IQ4_XS, :Q6_K_XL, …). The tag is required: this repo has no Q4_K_M, so a bare -hf with no tag falls back to the wrong file. The mmproj rides along in the manifest, so vision works from the same tag — no second download.

# text — auto-downloads to llama.cpp's own cache (24 GB-card default shown)
llama-server   -hf peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF:Q4_K_XL -ngl 99   # or llama-cli
# vision — same tag; the mmproj is pulled automatically
llama-mtmd-cli -hf peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF:Q4_K_XL -ngl 99 --image photo.jpg

Prefer to keep the files yourself? Download explicitly, then point -m at the local path:

hf download peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf \
  mmproj-F16.gguf --local-dir Dirk
llama-cli      -m Dirk/Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf -ngl 99                              # text
llama-mtmd-cli -m Dirk/Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf --mmproj Dirk/mmproj-F16.gguf -ngl 99  # vision

llama.cpp applies the embedded Sharp template automatically — nothing to pass.

Driving it from a coding agent? Add --reasoning-format deepseek to llama-server. It returns the model's <think> block in the OpenAI reasoning_content field instead of inline in content, so the agent never sees raw thinking tokens in the text stream. Current llama.cpp already defaults to this (--reasoning-format auto is defined as "same as deepseek"), so it is a no-op on a recent build and insurance on an older one. Just don't pass --reasoning-format none — that is the one that leaves the tags inline.

Pick your weapon

Qwen3.8-27B may be the new intelligence density frontier for local models that run on consumer hardware, but the already battle-tested Dagger and Nail, joined by the newer TielCoder, each have their own use cases, in an arsenal that contains all four.

  • Nail-35B-A3B generates tokens 3–4× faster than 27B models, while still being very good at routine coding, debugging, knowledge work, and many other kinds of tasks — which means that for tasks that aren't too hard for it, it writes the unit test and regression test, and implements the feature in the time it takes 3.8-27B to get out of the gate. Reach for Nail when you need volume routine work done right and fast.
  • TielCoder-35B-A3B is the dedicated coder: Nail's 35B-A3B speed class, rebuilt on Ornith-1.5 with the Sharp template and pointed at one job. It fixes 12 of 25 on SWE-bench-Live — level with Opus 4.6, four clear of Sonnet 5 (medium) — at the lowest mean time per attempt of the 35B-A3B family. It pays for that in general knowledge: 73.7 on MMLU-Pro against Nail's 84.0. Reach for TielCoder when the work is code; reach for Nail when the same session also has to know things.
  • Dagger-27B is — unlike 3.8-27B — specifically tuned to minimize the number of thinking tokens while sacrificing minimal accuracy, which might still give it the advantage in speed-to-answer and multi-turn stamina under the context ceiling. Reach for Dagger when you need a session to survive 100 turns.
  • Dirk-27B is what you reach for when the task is genuinely hard and you want the strongest local answer without filler — accepting that Nail reaches an answer faster on work it can handle, and that a marathon session running 100 turns under the context ceiling is Dagger's domain, not Dirk's.

Dagger, Nail and TielCoder might still be your go-to workhorses for long and short tasks within their ability bands, due to their advantage in speed and stamina.

Credits

  • Qwen — the Qwen3.8-27B weights.
  • Unsloth — the UD Dynamic 3.0 quants (MTP-preserving) this repo redistributes.
  • IST-DASLab — the GSQ-RCO quants (MTP-preserving) behind the 2–3 bpw tiers, redistributed here with only the chat template changed. Method: GSQ (Dadgarnia, Tabesh, Nikdan, Helcig, Kurtic, Kleinegger, Alistarh) and RCO (Helcig, Alistarh); originals at ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF.
  • froggeric — the fixed chat template the Sharp template builds on.

Apache-2.0, matching upstream.

Identity and Version

Repository
peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF
Publisher
Saga
Task
Image and text to text
Modality
Image and text
Library
gguf
Parameters
Not stated by the source
Languages
en, zh
Revision
52cb3e759635ab4605e08790b6c47df8adcf0744
First published
2026-08-14
Last updated
2026-09-02

Files and Weights

21 files, 239.7 GB in total. The weights are 15 files totalling 239.7 GB in gguf.

Weights15 files · 239.7 GB
Documentation1 file · 12.6 KB
Other4 files · 3.1 MB
Repository1 file · 2.9 KB
Every file
FileTypeSizeSHA-256
Dirk-Qwen3.8-27B-GSQ-RCO-IQ2_S.ggufWeights9.6 GB 79bf1148f37b
Dirk-Qwen3.8-27B-GSQ-RCO-IQ2_XS.ggufWeights8.8 GB b9ba59f8dd3e
Dirk-Qwen3.8-27B-GSQ-RCO-IQ3_S.ggufWeights12.1 GB 108cb89de873
Dirk-Qwen3.8-27B-GSQ-RCO-IQ3_XXS.ggufWeights10.4 GB 9c31d3f1d48b
Dirk-Qwen3.8-27B-UD-IQ4_XS.ggufWeights14.3 GB f14c59d94d86
Dirk-Qwen3.8-27B-UD-Q2_K_XL.ggufWeights9.8 GB 8813b11dfaaf
Dirk-Qwen3.8-27B-UD-Q3_K_XL.ggufWeights13.1 GB 87f707acd8a7
Dirk-Qwen3.8-27B-UD-Q4_K_S.ggufWeights15.4 GB 1fe42f846bd4
Dirk-Qwen3.8-27B-UD-Q4_K_XL.ggufWeights17.6 GB d1ad2472a147
Dirk-Qwen3.8-27B-UD-Q5_K_XL.ggufWeights20.9 GB 43d3b23be6e2
Dirk-Qwen3.8-27B-UD-Q6_K.ggufWeights22.0 GB 06601c59c3dd
Dirk-Qwen3.8-27B-UD-Q6_K_XL.ggufWeights25.3 GB 3ea8eebc1ec4
Dirk-Qwen3.8-27B-UD-Q8_K_L.ggufWeights28.0 GB a4e68db88401
Dirk-Qwen3.8-27B-UD-Q8_K_XL.ggufWeights31.5 GB be2f08a26002
mmproj-F16.ggufWeights927.6 MB cbb841a9ee06
README.mdDocumentation12.6 KB —
assets/card_dirk_mmlu_live.pngOther865.6 KB b99769fa4832
assets/card_swe_bounds.pngOther1.1 MB 4c4e7cb87eb8
assets/card_swe_sharp.pngOther791.3 KB 01229425b122
assets/dirk_banner_eyebrow.pngOther356.8 KB 52d0495836cf
.gitattributesRepository2.9 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
239.7 GB
Download from Saga

Released by Saga through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published239.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Dirk-Qwen3.8-27B-GGUF

Can I use Dirk-Qwen3.8-27B-GGUF commercially?

Yes. Dirk-Qwen3.8-27B-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

in 8 bit and over 718 arc-c in 4 bit. This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality. In otherwords while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more. This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants. and other quant versions (also see "Quantized" in the "model tree" too (lower right)). The strongest, smartest open source multi-stage model fine tune for consumer hardware ever and BUILT on consumer hardware via Unsloth. The first model of this size/type to breach "730" ARC-C…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-Swift series come from ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 22 to 52 (0-based indexing) have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. This is just a test/validation. The newly added Huihui-Qwen3.8-27B-abliterated-Ternary series come from prism-ml/Ternary-Bonsai-2-27B-gguf have been ablated, while the other layers…

Open weights apache-2.0 transformers

and it does so in 4bit and 8bit. Regular and MTP (fast) NEO IMATRIX GGUFs provided. (this model is part of the Qwen 3.6 27B Fable Fusion 711 pipelines: 2200+ likes, 3 million + downloads) instruct modes (2 new - Spoon / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model ctrl at the chat/message level). Model name has "plusIQ" in the name. A 12+12 (12 reasoning and 12 instruct) model with interactive optimization/help system will be releasing shortly too. BF16/16-bit MTP GGUF also avail. (there is also a extra robust "tools" version too - Q6 and Q8.) Extreme intelligence in a small package. Jaw dropping performance. Superior…

Open weights apache-2.0

Non-uniform GGUF quantizations produced with GSQ and RCO, with a vision projector for multimodal use. This repository provides GGUF quantizations of Qwen3.8-27B at four sizes, together with the model's vision projector (mmproj) for multimodal use. In contrast to uniform quantization, which applies a single quantization type to all weight tensors, each model here assigns a separate quantization type to every tensor. The assignment is obtained by a gradient-based search that allocates precision according to per-tensor sensitivity, subject to a total size budget. The resulting files are standard GGUF and run unmodified in llama.cpp, Ollama, and LM Studio. Both methods were developed at the…

Open weights apache-2.0 gguf