SAVRN
Search Contact SAVRN

Open-weight model · Text generation

humanizer-GGUF

by Stephen Yu jialinyyzz/humanizer-GGUF

humanizer-GGUF is an open-weight model for text generation from Stephen Yu, released under Apache License 2.0. Its published files total 63.6 GB. It draws 56 downloads a month.

GGUF files of jialinyyzz/humanizer (v2), a 12B model that rewrites AI-written drafts in English and Chinese so they read like a person wrote them, while keeping every number, date, name and quote.

Parameters—
Context—
Weights63.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads56

Model Card

By Stephen Yu, published under apache-2.0, revision 066a62f45691.

GGUF files of jialinyyzz/humanizer (v2), a 12B model that rewrites AI-written drafts in English and Chinese so they read like a person wrote them, while keeping every number, date, name and quote. This page is about the quantized files only: what each one costs in quality, and how they were made. For the model itself (training, full results, examples), see the main model card and GitHub. The files here are byte-for-byte the same as the GGUF files in the main repository; download from either. promptformat.json (tiny) holds the instruction and separator verbatim; you need it unless you use the chat template (see Usage). ¹ Peak resident memory of llama-server (llama.cpp, Metal) on an Apple M5…

Read Stephen Yu's full model card

humanizer 12B: GGUF quantizations

GGUF files of jialinyyzz/humanizer (v2), a 12B model that rewrites AI-written drafts in English and Chinese so they read like a person wrote them, while keeping every number, date, name and quote. This page is about the quantized files only: what each one costs in quality, and how they were made. For the model itself (training, full results, examples), see the main model card and GitHub.

The files here are byte-for-byte the same as the GGUF files in the main repository; download from either.

Files

File Bits Size Memory needed ¹ KL vs. bf16, EN / ZH ² Top-1 same as bf16, EN / ZH ² Standard llama.cpp imatrix build of the same class: KL EN / ZH (top-1) ³ Quality Use it for
humanizer-12b-bf16.gguf 16 (bf16) 23.8 GB about 24.8 GB (est.) 0 (reference) 100% (reference) Reference Quantizing it yourself; reference runs
humanizer-12b-Q8_0.gguf 8 (Q8_0) 12.7 GB 13.7 GB 0.0017 / 0.0017 98.45% / 98.25% same file: Q8_0 is a standard build Best 32 GB of memory or more. Recommended.
humanizer-12b-Q6_K.gguf 6 (Q6_K) 10.0 GB about 11.0 GB (est.) 0.0031 / 0.0035 97.95% / 97.48% same file: Q6_K is a standard build No measurable loss 16 GB of memory
humanizer-12b-Q4_K_M.gguf 4 (Q4_K_M), quantization-aware trained 7.6 GB 10.0 GB 0.0136 / 0.0146 95.62% / 94.77% Q4_K_M, 7.6 GB: 0.0203 / 0.0225 (94.46% / 93.46%) Slight loss About 14 GB machines, or when disk is tight
humanizer-12b-IQ3_XXS-QAT.gguf (the Q3 size) Q3: header type IQ3_XXS, but most bytes are Q4_K / Q6_K, about 3.9 bits per weight on average ⁴; quantization-aware trained 5.6 GB 8.0 GB 0.0300 / 0.0318 ² 93.63% / 92.64% IQ3_XXS (3-bit), 4.7 GB: 0.138 / 0.190 (85.72% / 82.55%) Small loss; a few more fact slips in English 12 GB machines; proofread numbers and names
humanizer-12b-IQ2_XS-QAT.gguf 2 (mostly IQ2_XS) ⁴, quantization-aware trained 3.9 GB 6.2 GB 0.106 / 0.106 87.72% / 87.07% IQ2_XS, 3.8 GB: 0.474 / 0.769 (74.43% / 65.53%) Lowest AI-detector score; a few more fact slips Smallest machines (8 GB); proofread numbers and names

prompt_format.json (tiny) holds the instruction and separator verbatim; you need it unless you use the chat template (see Usage).

¹ Peak resident memory of llama-server (llama.cpp, Metal) on an Apple M5 Max with the app's settings (8,192-token context, one request at a time) while rewriting. Q4_K_M, Q3 and 2-bit were measured side by side with the app's llama.cpp build on four real drafts; generation ran at about 52, 65 and 69 tokens per second. Q8_0 was measured in an earlier run on one 470-token draft, which read about 1.3 GB lower for the same file (2-bit: 4.9 GB there, 6.2 GB here), so it may need a little more than shown; Q6_K and bf16 are estimated (est.). A 32,768-token context adds about 0.4 GB. Leave room for the system and other apps (the app keeps about 4 GB free).

² llama-perplexity --kl-divergence, 30 chunks of 2,048 tokens per language (about 30,700 scored tokens each) of held-out drafts and rewrites (no overlap with calibration or training data), reference = the bf16 GGUF, measured on A100 and A30 GPUs (the same file measures within about 3% across GPU models, far below the gaps between files). KL is how far a file's next-token probabilities drift from bf16, averaged over every token; lower is better. "Top-1 same" is how often the file and bf16 would pick the same most likely next token.

³ The default method: a plain llama-quantize run of the same type with the same importance matrix (imatrix) our builds started from, no extra training. Q8_0 and Q6_K here are such standard builds. The IQ2_XS and IQ3_XXS comparison builds use the full vocabulary and 4-bit embeddings; the comparison below uses llama-quantize's default tensor types instead.

⁴ Both low-bit files are named after the type in their header, as llama.cpp reports it: the 2-bit file reports IQ2_XS and the Q3 file IQ3_XXS, but neither is a plain build of that type. humanizer-12b-IQ3_XXS-QAT.gguf here is the same file as humanizer-12b-Q3-QAT.gguf in the main repository (same sha256); the app downloads that copy. See How the files were made.

How it compares to standard quantization

The blue line is what plain llama.cpp quantization gives for this model: 13 types from IQ1_M to Q8_0, each a llama-quantize run with the same importance matrix our 2-bit and Q3 files started from, full vocabulary, every other setting at its default. The green line is the third-party imatrix GGUF of the same model (mradermacher/humanizer-i1-GGUF, 7 of its files); it follows the blue line, slightly worse in Chinese at 1–2 bits. The stars are the quantization-aware trained + distilled files in this repo. "Same size" below means the standard curve at that file size, counting only the best standard build at each size (KL interpolated on the log scale).

File here Size KL EN / ZH Standard build at the same size: KL EN / ZH Lower by A standard build reaches the same KL at Top-1 EN / ZH (standard at the same size)
humanizer-12b-IQ2_XS-QAT.gguf (2-bit) 3.9 GB 0.106 / 0.106 0.47 / 0.75 (IQ2_XS is 3.9 GB: 0.47 / 0.76) 4.4× / 7.1× about 5.1 / 5.3 GB (+1.2 / +1.4 GB) 87.7% / 87.1% (74.5% / 66.3%)
humanizer-12b-IQ3_XXS-QAT.gguf (Q3) 5.6 GB 0.030 / 0.032 0.065 / 0.079 2.2× / 2.5× about 6.5 GB (+0.9 / +1.0 GB) 93.6% / 92.6% (90.1% / 88.1%)
humanizer-12b-Q4_K_M.gguf 7.6 GB 0.0136 / 0.0146 0.018 / 0.020 1.3× / 1.4× about 8.1 GB (+0.5 GB) 95.6% / 94.8% (94.7% / 93.9%)
  • 2-bit: the 3.9 GB file drifts from bf16 less than a standard IQ3_XXS that is 1 GB larger (4.8 GB: 0.136 / 0.185), and in Chinese, where plain 2-bit quantization breaks down, it is 7× closer than the standard file of the same size.
  • Q3: less than half the KL of a standard build of its size; the smallest standard file that does as well is IQ4_XS at 6.6 GB (0.026 / 0.029), 1.0 GB larger.
  • Q4_K_M: the gain is real but small: about a quarter less KL than a standard build of the same size, the same as a standard build about 0.5 GB larger. A standard Q5_K_M (8.5 GB, 0.011 / 0.011) is still closer to bf16. (The standard Q4_K_M on the curve is 7.4 GB because llama-quantize's default keeps the embeddings at 6-bit; the one in the table above keeps them at 8-bit like ours, 7.6 GB, 0.0203 / 0.0225.)

How it was measured: every point uses the same bf16 reference and the same ~30,700 held-out tokens per language (llama-perplexity --kl-divergence, as in footnote ²). The Q3 and 2-bit files are measured against a bf16 with the same pruned vocabulary (pruning on its own: KL 0.0013 / 0.0002). Points were measured on A100 and A30 GPUs; cross-GPU difference ≈3%, far below the gaps shown. All numbers: assets/quant-curve-data.csv.

Run it in the app

humanizer also comes as a desktop app for macOS (Apple silicon) and Windows (x64) that runs these GGUF files locally. It is a small local web app: double-click it and it opens in your browser at http://127.0.0.1, runs llama.cpp in the background (Metal on Mac; CUDA, Vulkan or CPU on Windows), and works offline once the model is downloaded. Your text never leaves your computer. Paste a draft on the left, get the rewrite on the right, with new wording highlighted and replaced wording struck through.

  • Download: Humanizer-0.3.1-macos-arm64.dmg or Humanizer-0.3.1-windows-x64-setup.exe (or the portable .zip) from app 0.3.1 on GitHub Releases (later versions: latest release). The app is not code-signed yet; the one-time first-launch fix is in the install guide.
  • Picking a size: on first launch the setup page shows your memory and lists the sizes with a short quality note each, marking one as Recommended: 32 GB or more → Q8_0 ("Best quality"), 16 GB → Q6_K ("Nearly identical to Q8_0"), 14 GB → Q4_K_M ("Slight loss"), 12 GB → Q3 ("Small loss; a few more fact slips in English"), less → 2-bit ("Lowest AI-detector score; a few more fact slips"). The app picks the largest size that leaves about 4 GB for the system and your browser. You can pick any of them and download it once (downloads can be paused and resumed; Hugging Face or the hf-mirror.com mirror). To switch later, open the ⋯ menu → Change model size.
  • Which files the app offers: all five quantized files in the table, Q8_0, Q6_K, Q4_K_M, Q3 and 2-bit (app 0.3.0 and later; 0.2.0 had only the first three). The bf16 file is not in the app; use it with llama.cpp as below.

How the files were made

  • Standard types, our own calibration. Q8_0 and Q6_K are plain llama.cpp quantizations; Q6_K, Q4_K_M, Q3 and the 2-bit file use an importance matrix (imatrix) computed on our own English and Chinese rewriting data, not on generic web text.
  • Embeddings and output layer stay at 8-bit in Q8_0, Q6_K and Q4_K_M (the model ties them, so this is one tensor).
  • Quantization-aware training + distillation (Q4_K_M, Q3 and 2-bit). Starting from the standard imatrix build, the quantized weights are tuned block by block against the bf16 model, and the quantization scales are then distilled from bf16 on English and Chinese rewriting data. The output is still an ordinary GGUF of the same type: no custom kernels, any recent llama.cpp runs it. For Q4_K_M this lowers KL by about a third at the same size (0.0136 vs. 0.0203 in English).
  • Mixed precision by sensitivity (Q3 and 2-bit). Each tensor's sensitivity to quantization was measured, and bits were spread accordingly. The 2-bit file is IQ2_XS, with a few of the most sensitive tensors given a slightly larger type. The Q3 file spreads them more widely: by bytes it is about half Q4_K and a fifth Q6_K, with IQ3_XXS (17%) and IQ2_XS (11%) on the least sensitive tensors, about 3.9 bits per weight on average. Its header says IQ3_XXS, but most of its bytes are 4- and 6-bit.
  • Vocabulary pruning (Q3 and 2-bit). The files keep about 130,000 of the 262,144 tokens: the ones English and Chinese text actually uses, plus everything needed to spell any input. Any text still encodes and decodes exactly; rare symbols, emoji and other scripts just take a few more tokens. Pruning on its own changes the model by a KL of only about 0.001 (0.0013 English, 0.0002 Chinese); the Q3 and 2-bit KL in the table are measured against a bf16 with the same pruned vocabulary.

The result: the 2-bit file is closer to bf16 than a standard 3-bit IQ3_XXS that is 0.8 GB larger, and Chinese, which plain 2-bit quantization hurts most (KL 0.769 vs. 0.474 for English), comes out the same as English. The Q3 file, 0.9 GB larger than that standard IQ3_XXS, drifts about a fifth as much (KL 0.030 vs. 0.138 in English, 0.032 vs. 0.190 in Chinese) and is 2 GB smaller than Q4_K_M.

English vs. Chinese, by token confidence

Each token of the rewrite is put in a bin by how sure the bf16 model is of its top choice (x-axis), and the mean KL is compared between English and Chinese within the same bin. This plot counts only the tokens of the rewrite itself, so its means are lower than the table's; panel (c) is the standard Q4_K_M build, not the quantization-aware file here. With the standard IQ2_XS (d), Chinese drifts 1.3 to 6 times more than English at the same confidence, even on tokens bf16 is almost sure of: Chinese is genuinely quantized worse, not just harder. In the 2-bit file (e), the two languages overlap, as they do for Q6_K and Q4_K_M; panel (f) shows the ratio per bin.

Quality on the task

Fact check. A strict LLM judge (GLM-5.3, one vote per rewrite) compared every rewrite with its draft: 420 English and 204 Chinese rewrites from the held-out evaluation set (2 per draft), run on each file with llama.cpp. A second pass re-read each flagged rewrite and listed every problem ("spots").

File English: rewrites flagged, spots listed Chinese: rewrites flagged, spots listed
bf16 52 of 420, 160 spots 50 of 200 ⁵, 216 spots
Q8_0 44 of 420, 135 spots 54 of 204, 236 spots
Q6_K 56 of 420, 153 spots 56 of 204, 255 spots
Q4_K_M (quantization-aware) 58 of 420, 162 spots 51 of 204, 194 spots
Q4_K_M, standard build (for comparison) 56 of 420, 163 spots 48 of 204, 218 spots
Q3 (IQ3_XXS-QAT) 64 of 420, 210 spots 53 of 204, 233 spots
2-bit (IQ2_XS-QAT) 70 of 420, 265 spots 66 of 204, 361 spots

⁵ 4 Chinese rewrites from bf16 could not be judged.

Most spots are one word, number or short phrase, about 9 in 10 for every file (Q3: 191 of its 210 English spots; 2-bit: 243 of its 265 English spots and 330 of 361 Chinese). Compared draft by draft, Q8_0, Q6_K and Q4_K_M are within noise of each other and of bf16. The 2-bit file is not: in English it was flagged on 70 rewrites against 44 for Q8_0 on the same drafts, a real difference (in Chinese, 66 vs. 54, within noise). Q3 sits in between: in English 64 rewrites flagged, against 52 for bf16 and 58 for Q4_K_M; in Chinese 53, on par with bf16 (50). The judge is deliberately strict and some flags are harmless rewording, but with any file, read numbers, dates and names before you send.

AI detection. On the same 60 English drafts (Originality.ai, strictest setting), the 2-bit file had 0 of 60 rewrites flagged as AI, against 7 of 60 for Q8_0 and 4 of 60 for bf16. That is the best result of any version, and against Q8_0 it is not noise: all 7 drafts flagged for Q8_0 passed with the 2-bit file, and none went the other way (paired p = 0.016). It looks like the small drift that quantization adds makes the wording a little less predictable, so it reads less machine-like. The Q3 file had 4 of 60 flagged, the same as bf16 (3 drafts flagged only with Q3, 3 only with bf16). Q6_K and Q4_K_M were not measured. Detectors change; this is one measurement on one date.

Usage

Text completion, not chat. The model was trained on one plain prompt: the instruction from prompt_format.json, a blank line, your draft, then \n\n### Rewritten:\n\n. No system prompt, no chat turns. Sampling: temperature 1.0, top-p 0.95, and nothing else (top-k 0, min-p 0, repetition penalty 1.0); stop on EOS only, no stop strings.

brew install llama.cpp            # or: winget install llama.cpp / a zip from github.com/ggml-org/llama.cpp/releases
pip install -U "huggingface_hub[cli]"
hf download jialinyyzz/humanizer-GGUF humanizer-12b-Q8_0.gguf prompt_format.json --local-dir ./humanizer-model
#   16 GB: humanizer-12b-Q6_K.gguf · tight: humanizer-12b-Q4_K_M.gguf · tighter: humanizer-12b-IQ3_XXS-QAT.gguf · smallest: humanizer-12b-IQ2_XS-QAT.gguf
llama-server -m ./humanizer-model/humanizer-12b-Q8_0.gguf -c 8192 -np 1 -ngl 99 --host 127.0.0.1 --port 8080
import json, urllib.request

PF = json.load(open("humanizer-model/prompt_format.json", encoding="utf-8"))

def humanize(draft: str, url: str = "http://127.0.0.1:8080") -> str:
    body = {"prompt": PF["instr"] + "\n\n" + draft.strip() + PF["sep"],
            "temperature": 1.0, "top_p": 0.95, "top_k": 0, "min_p": 0, "repeat_penalty": 1.0,
            "n_predict": 2048}                    # no "stop": the model ends at EOS
    req = urllib.request.Request(url + "/completion", json.dumps(body).encode("utf-8"),
                                 {"Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=900) as r:
        return json.load(r)["content"].strip()

print(humanize(open("draft.txt", encoding="utf-8").read()))

Chat template and defaults. Every file here carries a chat template that builds exactly this prompt from the last user message (system prompts and earlier turns are ignored), and stores the sampling settings above as defaults. So llama-server's /v1/chat/completions (with --jinja, the default in recent builds) and chat apps that use the file's template, such as LM Studio, work with one draft per message. Ollama ignores the GGUF's template; use the Modelfile in USAGE.md, section 7. Keep instruction + draft + rewrite within 8,192 tokens and split long documents at paragraph breaks, or use the hz command-line tool, which does it for you.

Wrong language, occasionally. On short, informal English drafts with technical jargon, the model occasionally writes the whole rewrite in Chinese. The app (0.3.1 and later) and hz check the language and sample again automatically. If you call the model yourself: when an English draft comes back with more than a few Chinese characters, sample once more with the same settings.

Full guide for every runtime, a batch script, long documents, Chinese and troubleshooting: USAGE.md. sha256 checksums:

File Bytes sha256
humanizer-12b-bf16.gguf 23,832,049,568 47d79b44c3e15ea2540f4edb63556f2b7e456067d7dd252a42ff103a26a51be9
humanizer-12b-Q8_0.gguf 12,669,630,368 8d7a457b56de6530eaaf0151ccfa7550a4b20dab979737e259da9c63e960e0b0
humanizer-12b-Q6_K.gguf 10,029,799,584 c98f03bb9e71456181f99b0e1d3391e07ce1afc357f9db4c33d6d379b8dd9f0d
humanizer-12b-Q4_K_M.gguf 7,625,160,864 2229574dec5178629575ee4a153dfee7d9e924ab997d0ad9d2ac622b67e44834
humanizer-12b-IQ3_XXS-QAT.gguf 5,587,794,816 307bbfdf66fb22bf98aaa93fe7d54a713bf167e646073dc0cf10870b1525eb94
humanizer-12b-IQ2_XS-QAT.gguf 3,893,632,896 383e5ca8f1f48ab5f65013adbc1965fa70d1d1afa6c45d932c34b854e38edbb5

Links

License

Apache License 2.0. Fine-tuned from google/gemma-4-12B, which Google releases under Apache 2.0. This project is not affiliated with or endorsed by Google.

Identity and Version

Repository
jialinyyzz/humanizer-GGUF
Publisher
Stephen Yu
Task
Text generation
Modality
Text
Library
gguf
Parameters
Not stated by the source
Languages
en, zh
Revision
066a62f4569173f738d46bdea3688709a8141506
First published
2026-10-06
Last updated
2026-10-07

Files and Weights

15 files, 63.6 GB in total. The weights are 6 files totalling 63.6 GB in gguf.

Weights6 files · 63.6 GB
Configuration1 file · 553 B
Documentation1 file · 20.6 KB
Other6 files · 1.5 MB
Repository1 file · 2.3 KB
Every file
FileTypeSizeSHA-256
humanizer-12b-IQ2_XS-QAT.ggufWeights3.9 GB 383e5ca8f1f4
humanizer-12b-IQ3_XXS-QAT.ggufWeights5.6 GB 307bbfdf66fb
humanizer-12b-Q4_K_M.ggufWeights7.6 GB 2229574dec51
humanizer-12b-Q6_K.ggufWeights10.0 GB c98f03bb9e71
humanizer-12b-Q8_0.ggufWeights12.7 GB 8d7a457b56de
humanizer-12b-bf16.ggufWeights23.8 GB 47d79b44c3e1
prompt_format.jsonConfiguration553 B —
README.mdDocumentation20.6 KB —
assets/app-en.pngOther522.1 KB 96775d01e86b
assets/banner-en.pngOther194.7 KB 17d5fe5e800d
assets/kl-en-zh-by-confidence.pngOther299.7 KB f56c534f390c
assets/quant-curve-data.csvOther4.7 KB —
assets/quant-curve-kl.pngOther239.9 KB 5baf6c3ea189
assets/quant-curve-top1.pngOther234.5 KB 7c4f76340268
.gitattributesRepository2.3 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
63.6 GB
Download from Stephen Yu

Released by Stephen Yu through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published63.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About humanizer-GGUF

Can I use humanizer-GGUF commercially?

Yes. humanizer-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers