humanizer 12B: GGUF quantizations
GGUF files of jialinyyzz/humanizer (v2), a 12B model that rewrites AI-written drafts in English and Chinese so they read like a person wrote them, while keeping every number, date, name and quote. This page is about the quantized files only: what each one costs in quality, and how they were made. For the model itself (training, full results, examples), see the main model card and GitHub.
The files here are byte-for-byte the same as the GGUF files in the main repository; download from either.
Files
| File |
Bits |
Size |
Memory needed ¹ |
KL vs. bf16, EN / ZH ² |
Top-1 same as bf16, EN / ZH ² |
Standard llama.cpp imatrix build of the same class: KL EN / ZH (top-1) ³ |
Quality |
Use it for |
humanizer-12b-bf16.gguf |
16 (bf16) |
23.8 GB |
about 24.8 GB (est.) |
0 (reference) |
100% (reference) |
Reference |
Quantizing it yourself; reference runs |
humanizer-12b-Q8_0.gguf |
8 (Q8_0) |
12.7 GB |
13.7 GB |
0.0017 / 0.0017 |
98.45% / 98.25% |
same file: Q8_0 is a standard build |
Best |
32 GB of memory or more. Recommended. |
humanizer-12b-Q6_K.gguf |
6 (Q6_K) |
10.0 GB |
about 11.0 GB (est.) |
0.0031 / 0.0035 |
97.95% / 97.48% |
same file: Q6_K is a standard build |
No measurable loss |
16 GB of memory |
humanizer-12b-Q4_K_M.gguf |
4 (Q4_K_M), quantization-aware trained |
7.6 GB |
10.0 GB |
0.0136 / 0.0146 |
95.62% / 94.77% |
Q4_K_M, 7.6 GB: 0.0203 / 0.0225 (94.46% / 93.46%) |
Slight loss |
About 14 GB machines, or when disk is tight |
humanizer-12b-IQ3_XXS-QAT.gguf (the Q3 size) |
Q3: header type IQ3_XXS, but most bytes are Q4_K / Q6_K, about 3.9 bits per weight on average ⁴; quantization-aware trained |
5.6 GB |
8.0 GB |
0.0300 / 0.0318 ² |
93.63% / 92.64% |
IQ3_XXS (3-bit), 4.7 GB: 0.138 / 0.190 (85.72% / 82.55%) |
Small loss; a few more fact slips in English |
12 GB machines; proofread numbers and names |
humanizer-12b-IQ2_XS-QAT.gguf |
2 (mostly IQ2_XS) ⁴, quantization-aware trained |
3.9 GB |
6.2 GB |
0.106 / 0.106 |
87.72% / 87.07% |
IQ2_XS, 3.8 GB: 0.474 / 0.769 (74.43% / 65.53%) |
Lowest AI-detector score; a few more fact slips |
Smallest machines (8 GB); proofread numbers and names |
prompt_format.json (tiny) holds the instruction and separator verbatim; you need it unless you use the chat template (see Usage).
¹ Peak resident memory of llama-server (llama.cpp, Metal) on an Apple M5 Max with the app's settings (8,192-token context, one request at a time) while rewriting. Q4_K_M, Q3 and 2-bit were measured side by side with the app's llama.cpp build on four real drafts; generation ran at about 52, 65 and 69 tokens per second. Q8_0 was measured in an earlier run on one 470-token draft, which read about 1.3 GB lower for the same file (2-bit: 4.9 GB there, 6.2 GB here), so it may need a little more than shown; Q6_K and bf16 are estimated (est.). A 32,768-token context adds about 0.4 GB. Leave room for the system and other apps (the app keeps about 4 GB free).
² llama-perplexity --kl-divergence, 30 chunks of 2,048 tokens per language (about 30,700 scored tokens each) of held-out drafts and rewrites (no overlap with calibration or training data), reference = the bf16 GGUF, measured on A100 and A30 GPUs (the same file measures within about 3% across GPU models, far below the gaps between files). KL is how far a file's next-token probabilities drift from bf16, averaged over every token; lower is better. "Top-1 same" is how often the file and bf16 would pick the same most likely next token.
³ The default method: a plain llama-quantize run of the same type with the same importance matrix (imatrix) our builds started from, no extra training. Q8_0 and Q6_K here are such standard builds. The IQ2_XS and IQ3_XXS comparison builds use the full vocabulary and 4-bit embeddings; the comparison below uses llama-quantize's default tensor types instead.
⁴ Both low-bit files are named after the type in their header, as llama.cpp reports it: the 2-bit file reports IQ2_XS and the Q3 file IQ3_XXS, but neither is a plain build of that type. humanizer-12b-IQ3_XXS-QAT.gguf here is the same file as humanizer-12b-Q3-QAT.gguf in the main repository (same sha256); the app downloads that copy. See How the files were made.
How it compares to standard quantization
The blue line is what plain llama.cpp quantization gives for this model: 13 types from IQ1_M to Q8_0, each a llama-quantize run with the same importance matrix our 2-bit and Q3 files started from, full vocabulary, every other setting at its default. The green line is the third-party imatrix GGUF of the same model (mradermacher/humanizer-i1-GGUF, 7 of its files); it follows the blue line, slightly worse in Chinese at 1–2 bits. The stars are the quantization-aware trained + distilled files in this repo. "Same size" below means the standard curve at that file size, counting only the best standard build at each size (KL interpolated on the log scale).
| File here |
Size |
KL EN / ZH |
Standard build at the same size: KL EN / ZH |
Lower by |
A standard build reaches the same KL at |
Top-1 EN / ZH (standard at the same size) |
humanizer-12b-IQ2_XS-QAT.gguf (2-bit) |
3.9 GB |
0.106 / 0.106 |
0.47 / 0.75 (IQ2_XS is 3.9 GB: 0.47 / 0.76) |
4.4× / 7.1× |
about 5.1 / 5.3 GB (+1.2 / +1.4 GB) |
87.7% / 87.1% (74.5% / 66.3%) |
humanizer-12b-IQ3_XXS-QAT.gguf (Q3) |
5.6 GB |
0.030 / 0.032 |
0.065 / 0.079 |
2.2× / 2.5× |
about 6.5 GB (+0.9 / +1.0 GB) |
93.6% / 92.6% (90.1% / 88.1%) |
humanizer-12b-Q4_K_M.gguf |
7.6 GB |
0.0136 / 0.0146 |
0.018 / 0.020 |
1.3× / 1.4× |
about 8.1 GB (+0.5 GB) |
95.6% / 94.8% (94.7% / 93.9%) |
- 2-bit: the 3.9 GB file drifts from bf16 less than a standard IQ3_XXS that is 1 GB larger (4.8 GB: 0.136 / 0.185), and in Chinese, where plain 2-bit quantization breaks down, it is 7× closer than the standard file of the same size.
- Q3: less than half the KL of a standard build of its size; the smallest standard file that does as well is IQ4_XS at 6.6 GB (0.026 / 0.029), 1.0 GB larger.
- Q4_K_M: the gain is real but small: about a quarter less KL than a standard build of the same size, the same as a standard build about 0.5 GB larger. A standard Q5_K_M (8.5 GB, 0.011 / 0.011) is still closer to bf16. (The standard Q4_K_M on the curve is 7.4 GB because llama-quantize's default keeps the embeddings at 6-bit; the one in the table above keeps them at 8-bit like ours, 7.6 GB, 0.0203 / 0.0225.)
How it was measured: every point uses the same bf16 reference and the same ~30,700 held-out tokens per language (llama-perplexity --kl-divergence, as in footnote ²). The Q3 and 2-bit files are measured against a bf16 with the same pruned vocabulary (pruning on its own: KL 0.0013 / 0.0002). Points were measured on A100 and A30 GPUs; cross-GPU difference ≈3%, far below the gaps shown. All numbers: assets/quant-curve-data.csv.
Run it in the app
humanizer also comes as a desktop app for macOS (Apple silicon) and Windows (x64) that runs these GGUF files locally. It is a small local web app: double-click it and it opens in your browser at http://127.0.0.1, runs llama.cpp in the background (Metal on Mac; CUDA, Vulkan or CPU on Windows), and works offline once the model is downloaded. Your text never leaves your computer. Paste a draft on the left, get the rewrite on the right, with new wording highlighted and replaced wording struck through.
- Download:
Humanizer-0.3.1-macos-arm64.dmg or Humanizer-0.3.1-windows-x64-setup.exe (or the portable .zip) from app 0.3.1 on GitHub Releases (later versions: latest release). The app is not code-signed yet; the one-time first-launch fix is in the install guide.
- Picking a size: on first launch the setup page shows your memory and lists the sizes with a short quality note each, marking one as Recommended: 32 GB or more → Q8_0 ("Best quality"), 16 GB → Q6_K ("Nearly identical to Q8_0"), 14 GB → Q4_K_M ("Slight loss"), 12 GB → Q3 ("Small loss; a few more fact slips in English"), less → 2-bit ("Lowest AI-detector score; a few more fact slips"). The app picks the largest size that leaves about 4 GB for the system and your browser. You can pick any of them and download it once (downloads can be paused and resumed; Hugging Face or the hf-mirror.com mirror). To switch later, open the ⋯ menu → Change model size.
- Which files the app offers: all five quantized files in the table, Q8_0, Q6_K, Q4_K_M, Q3 and 2-bit (app 0.3.0 and later; 0.2.0 had only the first three). The bf16 file is not in the app; use it with llama.cpp as below.
How the files were made
- Standard types, our own calibration. Q8_0 and Q6_K are plain llama.cpp quantizations; Q6_K, Q4_K_M, Q3 and the 2-bit file use an importance matrix (imatrix) computed on our own English and Chinese rewriting data, not on generic web text.
- Embeddings and output layer stay at 8-bit in Q8_0, Q6_K and Q4_K_M (the model ties them, so this is one tensor).
- Quantization-aware training + distillation (Q4_K_M, Q3 and 2-bit). Starting from the standard imatrix build, the quantized weights are tuned block by block against the bf16 model, and the quantization scales are then distilled from bf16 on English and Chinese rewriting data. The output is still an ordinary GGUF of the same type: no custom kernels, any recent llama.cpp runs it. For Q4_K_M this lowers KL by about a third at the same size (0.0136 vs. 0.0203 in English).
- Mixed precision by sensitivity (Q3 and 2-bit). Each tensor's sensitivity to quantization was measured, and bits were spread accordingly. The 2-bit file is IQ2_XS, with a few of the most sensitive tensors given a slightly larger type. The Q3 file spreads them more widely: by bytes it is about half Q4_K and a fifth Q6_K, with IQ3_XXS (17%) and IQ2_XS (11%) on the least sensitive tensors, about 3.9 bits per weight on average. Its header says IQ3_XXS, but most of its bytes are 4- and 6-bit.
- Vocabulary pruning (Q3 and 2-bit). The files keep about 130,000 of the 262,144 tokens: the ones English and Chinese text actually uses, plus everything needed to spell any input. Any text still encodes and decodes exactly; rare symbols, emoji and other scripts just take a few more tokens. Pruning on its own changes the model by a KL of only about 0.001 (0.0013 English, 0.0002 Chinese); the Q3 and 2-bit KL in the table are measured against a bf16 with the same pruned vocabulary.
The result: the 2-bit file is closer to bf16 than a standard 3-bit IQ3_XXS that is 0.8 GB larger, and Chinese, which plain 2-bit quantization hurts most (KL 0.769 vs. 0.474 for English), comes out the same as English. The Q3 file, 0.9 GB larger than that standard IQ3_XXS, drifts about a fifth as much (KL 0.030 vs. 0.138 in English, 0.032 vs. 0.190 in Chinese) and is 2 GB smaller than Q4_K_M.
English vs. Chinese, by token confidence
Each token of the rewrite is put in a bin by how sure the bf16 model is of its top choice (x-axis), and the mean KL is compared between English and Chinese within the same bin. This plot counts only the tokens of the rewrite itself, so its means are lower than the table's; panel (c) is the standard Q4_K_M build, not the quantization-aware file here. With the standard IQ2_XS (d), Chinese drifts 1.3 to 6 times more than English at the same confidence, even on tokens bf16 is almost sure of: Chinese is genuinely quantized worse, not just harder. In the 2-bit file (e), the two languages overlap, as they do for Q6_K and Q4_K_M; panel (f) shows the ratio per bin.
Quality on the task
Fact check. A strict LLM judge (GLM-5.3, one vote per rewrite) compared every rewrite with its draft: 420 English and 204 Chinese rewrites from the held-out evaluation set (2 per draft), run on each file with llama.cpp. A second pass re-read each flagged rewrite and listed every problem ("spots").
| File |
English: rewrites flagged, spots listed |
Chinese: rewrites flagged, spots listed |
| bf16 |
52 of 420, 160 spots |
50 of 200 ⁵, 216 spots |
| Q8_0 |
44 of 420, 135 spots |
54 of 204, 236 spots |
| Q6_K |
56 of 420, 153 spots |
56 of 204, 255 spots |
| Q4_K_M (quantization-aware) |
58 of 420, 162 spots |
51 of 204, 194 spots |
| Q4_K_M, standard build (for comparison) |
56 of 420, 163 spots |
48 of 204, 218 spots |
Q3 (IQ3_XXS-QAT) |
64 of 420, 210 spots |
53 of 204, 233 spots |
2-bit (IQ2_XS-QAT) |
70 of 420, 265 spots |
66 of 204, 361 spots |
⁵ 4 Chinese rewrites from bf16 could not be judged.
Most spots are one word, number or short phrase, about 9 in 10 for every file (Q3: 191 of its 210 English spots; 2-bit: 243 of its 265 English spots and 330 of 361 Chinese). Compared draft by draft, Q8_0, Q6_K and Q4_K_M are within noise of each other and of bf16. The 2-bit file is not: in English it was flagged on 70 rewrites against 44 for Q8_0 on the same drafts, a real difference (in Chinese, 66 vs. 54, within noise). Q3 sits in between: in English 64 rewrites flagged, against 52 for bf16 and 58 for Q4_K_M; in Chinese 53, on par with bf16 (50). The judge is deliberately strict and some flags are harmless rewording, but with any file, read numbers, dates and names before you send.
AI detection. On the same 60 English drafts (Originality.ai, strictest setting), the 2-bit file had 0 of 60 rewrites flagged as AI, against 7 of 60 for Q8_0 and 4 of 60 for bf16. That is the best result of any version, and against Q8_0 it is not noise: all 7 drafts flagged for Q8_0 passed with the 2-bit file, and none went the other way (paired p = 0.016). It looks like the small drift that quantization adds makes the wording a little less predictable, so it reads less machine-like. The Q3 file had 4 of 60 flagged, the same as bf16 (3 drafts flagged only with Q3, 3 only with bf16). Q6_K and Q4_K_M were not measured. Detectors change; this is one measurement on one date.
Usage
Text completion, not chat. The model was trained on one plain prompt: the instruction from prompt_format.json, a blank line, your draft, then \n\n### Rewritten:\n\n. No system prompt, no chat turns. Sampling: temperature 1.0, top-p 0.95, and nothing else (top-k 0, min-p 0, repetition penalty 1.0); stop on EOS only, no stop strings.
brew install llama.cpp # or: winget install llama.cpp / a zip from github.com/ggml-org/llama.cpp/releases
pip install -U "huggingface_hub[cli]"
hf download jialinyyzz/humanizer-GGUF humanizer-12b-Q8_0.gguf prompt_format.json --local-dir ./humanizer-model
# 16 GB: humanizer-12b-Q6_K.gguf · tight: humanizer-12b-Q4_K_M.gguf · tighter: humanizer-12b-IQ3_XXS-QAT.gguf · smallest: humanizer-12b-IQ2_XS-QAT.gguf
llama-server -m ./humanizer-model/humanizer-12b-Q8_0.gguf -c 8192 -np 1 -ngl 99 --host 127.0.0.1 --port 8080
import json, urllib.request
PF = json.load(open("humanizer-model/prompt_format.json", encoding="utf-8"))
def humanize(draft: str, url: str = "http://127.0.0.1:8080") -> str:
body = {"prompt": PF["instr"] + "\n\n" + draft.strip() + PF["sep"],
"temperature": 1.0, "top_p": 0.95, "top_k": 0, "min_p": 0, "repeat_penalty": 1.0,
"n_predict": 2048} # no "stop": the model ends at EOS
req = urllib.request.Request(url + "/completion", json.dumps(body).encode("utf-8"),
{"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=900) as r:
return json.load(r)["content"].strip()
print(humanize(open("draft.txt", encoding="utf-8").read()))
Chat template and defaults. Every file here carries a chat template that builds exactly this prompt from the last user message (system prompts and earlier turns are ignored), and stores the sampling settings above as defaults. So llama-server's /v1/chat/completions (with --jinja, the default in recent builds) and chat apps that use the file's template, such as LM Studio, work with one draft per message. Ollama ignores the GGUF's template; use the Modelfile in USAGE.md, section 7. Keep instruction + draft + rewrite within 8,192 tokens and split long documents at paragraph breaks, or use the hz command-line tool, which does it for you.
Wrong language, occasionally. On short, informal English drafts with technical jargon, the model occasionally writes the whole rewrite in Chinese. The app (0.3.1 and later) and hz check the language and sample again automatically. If you call the model yourself: when an English draft comes back with more than a few Chinese characters, sample once more with the same settings.
Full guide for every runtime, a batch script, long documents, Chinese and troubleshooting: USAGE.md. sha256 checksums:
| File |
Bytes |
sha256 |
humanizer-12b-bf16.gguf |
23,832,049,568 |
47d79b44c3e15ea2540f4edb63556f2b7e456067d7dd252a42ff103a26a51be9 |
humanizer-12b-Q8_0.gguf |
12,669,630,368 |
8d7a457b56de6530eaaf0151ccfa7550a4b20dab979737e259da9c63e960e0b0 |
humanizer-12b-Q6_K.gguf |
10,029,799,584 |
c98f03bb9e71456181f99b0e1d3391e07ce1afc357f9db4c33d6d379b8dd9f0d |
humanizer-12b-Q4_K_M.gguf |
7,625,160,864 |
2229574dec5178629575ee4a153dfee7d9e924ab997d0ad9d2ac622b67e44834 |
humanizer-12b-IQ3_XXS-QAT.gguf |
5,587,794,816 |
307bbfdf66fb22bf98aaa93fe7d54a713bf167e646073dc0cf10870b1525eb94 |
humanizer-12b-IQ2_XS-QAT.gguf |
3,893,632,896 |
383e5ca8f1f48ab5f65013adbc1965fa70d1d1afa6c45d932c34b854e38edbb5 |
Links
License
Apache License 2.0. Fine-tuned from google/gemma-4-12B, which Google releases under Apache 2.0. This project is not affiliated with or endorsed by Google.