SAVRN
Search Contact SAVRN

Open-weight model · Any to any

kai-os_Grug-12B-GGUF

by Bartowski bartowski/kai-os_Grug-12B-GGUF

Using llama.cpp release b10068 for quantization. All quants made using imatrix option with dataset from here Run them in your choice of tools: Note: if it's a newly supported model, you may need to wait for an update from the developers.

Parameters
Context
Weights197.5 GB
Licenseother
AccessOpen weights
Monthly Downloads250.7k

Model Card

Using llama.cpp release b10068 for quantization. All quants made using imatrix option with dataset from here Run them in your choice of tools: Note: if it's a newly supported model, you may need to wait for an update from the developers. Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings and output weights quantized to Q80 instead of what they would normally default to. First, make sure you have huggingface-cli installed: Then, you can target the specific file you want: If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run: You can either specify a new local-dir…

Excerpt from the card by Bartowski, licensed other.

Identity and Version

Repository
bartowski/kai-os_Grug-12B-GGUF
Publisher
Bartowski
Task
Any to any
Modality
Multimodal
Library
transformers
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
ea2c68a94a48fc76ea40da176ab8a26c675f2ca7
First published
2026-07-20
Last updated
2026-07-20

Files and Weights

30 files, 197.5 GB in total. The weights are 28 files totalling 197.5 GB in gguf.

Weights28 files · 197.5 GB
Documentation1 file · 14.6 KB
Repository1 file · 3.3 KB
Every file
FileTypeSizeSHA-256
kai-os_Grug-12B-IQ2_M.ggufWeights4.9 GB 99073872f4f2
kai-os_Grug-12B-IQ2_S.ggufWeights4.7 GB 0ca588fdd442
kai-os_Grug-12B-IQ3_M.ggufWeights6.0 GB 588d6235da55
kai-os_Grug-12B-IQ3_XS.ggufWeights5.5 GB 68eb7ad313f6
kai-os_Grug-12B-IQ3_XXS.ggufWeights5.1 GB a0bd33c68164
kai-os_Grug-12B-IQ4_NL.ggufWeights7.1 GB 30a85d09dcfe
kai-os_Grug-12B-IQ4_XS.ggufWeights6.8 GB 13dfea2c8cc9
kai-os_Grug-12B-Q2_K.ggufWeights5.1 GB 5e711e475223
kai-os_Grug-12B-Q2_K_L.ggufWeights5.3 GB acb94b7e6d2e
kai-os_Grug-12B-Q3_K_L.ggufWeights6.7 GB 65495152ba92
kai-os_Grug-12B-Q3_K_M.ggufWeights6.3 GB 83898fd43216
kai-os_Grug-12B-Q3_K_S.ggufWeights5.7 GB ce7a335bfa9f
kai-os_Grug-12B-Q3_K_XL.ggufWeights6.9 GB 8a138cda84dc
kai-os_Grug-12B-Q4_0.ggufWeights7.1 GB 2e75b6d214f3
kai-os_Grug-12B-Q4_1.ggufWeights7.8 GB ca62278851ce
kai-os_Grug-12B-Q4_K_L.ggufWeights7.9 GB 20d19935e6fa
kai-os_Grug-12B-Q4_K_M.ggufWeights7.7 GB 5110bdafeb92
kai-os_Grug-12B-Q4_K_S.ggufWeights7.2 GB 9873d2322882
kai-os_Grug-12B-Q5_K_L.ggufWeights9.0 GB 7121f469596a
kai-os_Grug-12B-Q5_K_M.ggufWeights8.8 GB a79c4e1d9c05
kai-os_Grug-12B-Q5_K_S.ggufWeights8.4 GB 0e3fdd734ad5
kai-os_Grug-12B-Q6_K.ggufWeights10.2 GB 8bda1bec23bf
kai-os_Grug-12B-Q6_K_L.ggufWeights10.5 GB 64403dc2eeb8
kai-os_Grug-12B-Q8_0.ggufWeights12.7 GB 654e42ce66ff
kai-os_Grug-12B-bf16.ggufWeights23.8 GB 6944eb71ee57
kai-os_Grug-12B-imatrix.ggufWeights7.5 MB dc525c600193
mmproj-kai-os_Grug-12B-bf16.ggufWeights175.1 MB 3b40b68b9889
mmproj-kai-os_Grug-12B-f16.ggufWeights122.0 MB 0d59c7571a59
README.mdDocumentation14.6 KB
.gitattributesRepository3.3 KB

License and Download

License
other
Access
Open weights, no gate
Download size
197.5 GB
Download from Bartowski

Released by Bartowski through its official repository on Hugging Face.

Built From

  • Derived from kai-os/Grug-12B
  • Quantized from kai-os/Grug-12B
  • Trained on (disclosed) CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
  • Trained on (disclosed) HSH-Intelligence/verified-math-reasoning-3k
  • Trained on (disclosed) Madarabr/cortex-adaptive-thinking
  • Trained on (disclosed) Scale-or-Reason/general-reasoning-ift-pairs
  • Trained on (disclosed) hotdogs/uka-glm-5.2
  • Trained on (disclosed) kd13/CodeDebug-Instruct-v2-Reasoning
  • Trained on (disclosed) samcheng0/lumia-reasoning-sft-v1

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Local 36-row math reasoning eval Task EOS-only local math reasoning proxyMetric Grug-12B average generated tokensComparison conditions not established 68.9444 bartowski
Publisher reported
Evaluated revision not stated
Local 36-row math reasoning eval Task EOS-only local math reasoning proxyMetric Grug-12B proxy accuracyComparison conditions not established 1 bartowski
Publisher reported
Evaluated revision not stated
Local 36-row math reasoning eval Task EOS-only local math reasoning proxyMetric Grug-12B total generated tokensComparison conditions not established 2482 bartowski
Publisher reported
Evaluated revision not stated

Memory Requirements

PrecisionWeights in memory
As published197.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About kai-os_Grug-12B-GGUF

What license is kai-os_Grug-12B-GGUF released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Any to any

gemma-4-E4B-it-GGUF

GGML Org

Run with https://llama.app - https://huggingface.co/google/gemma-4-E4B-it - https://huggingface.co/google/gemma-4-E4B-it-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized - add info - add dflash

Open weights apache-2.0

Model · Any to any

gemma-4-12B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-12B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-12B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-E4B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0

Model · Any to any

gemma-4-E4B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E4B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers

Model · Any to any

gemma-4-E2B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E2B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers