SAVRN
Search Contact SAVRN

Open-weight model · Any to any

TheDrummer_Orion-26B-A4B-v1-GGUF

by Bartowski bartowski/TheDrummer_Orion-26B-A4B-v1-GGUF

Using llama.cpp release b10630 for quantization. Don't know which to choose? Grab Q4KM (17.04GB) - usually a good mix of size and performance.

Parameters
Context
Weights435.7 GB
License
AccessOpen weights
Monthly Downloads24.6k

Model Card

Using llama.cpp release b10630 for quantization. Don't know which to choose? Grab Q4KM (17.04GB) - usually a good mix of size and performance. Download instructions available here First, make sure you have the Hugging Face CLI installed: These quants run with llama.cpp - installable in one line via llama.app: llama-server includes a built-in chat web UI, served at http://localhost:8080 by default. These quants were made with llama.cpp release b10630 - if this model's architecture is newly supported, you'll need that release or newer to run them. They also work in: model supports image and audio input. Alongside the quants, this repo includes the multimodal projector files…

Excerpt from the card by Bartowski.

Identity and Version

Repository
bartowski/TheDrummer_Orion-26B-A4B-v1-GGUF
Publisher
Bartowski
Task
Any to any
Modality
Multimodal
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
5a12f5b78e5a9018b1721fd9ceed1fdb7c5543dc
First published
2026-08-26
Last updated
2026-08-27

Files and Weights

32 files, 435.7 GB in total. The weights are 29 files totalling 435.7 GB in gguf.

Weights29 files · 435.7 GB
Documentation1 file · 15.0 KB
Other1 file · 1.0 MB
Repository1 file · 3.7 KB
Every file
FileTypeSizeSHA-256
TheDrummer_Orion-26B-A4B-v1-IQ2_M.ggufWeights10.7 GB e30f7e789104
TheDrummer_Orion-26B-A4B-v1-IQ2_S.ggufWeights10.2 GB 3d0b6a0b7f29
TheDrummer_Orion-26B-A4B-v1-IQ2_XS.ggufWeights10.1 GB f0cc1872785f
TheDrummer_Orion-26B-A4B-v1-IQ3_M.ggufWeights13.3 GB cb7572a6be41
TheDrummer_Orion-26B-A4B-v1-IQ3_XS.ggufWeights12.4 GB 1f2c176abff7
TheDrummer_Orion-26B-A4B-v1-IQ3_XXS.ggufWeights12.2 GB 0ef6795453b3
TheDrummer_Orion-26B-A4B-v1-IQ4_NL.ggufWeights14.7 GB e4ac738dc920
TheDrummer_Orion-26B-A4B-v1-IQ4_XS.ggufWeights14.2 GB 318ca310fbb0
TheDrummer_Orion-26B-A4B-v1-Q2_K.ggufWeights11.0 GB 79bc026ff326
TheDrummer_Orion-26B-A4B-v1-Q2_K_L.ggufWeights11.1 GB 4496641a1d1a
TheDrummer_Orion-26B-A4B-v1-Q3_K_L.ggufWeights13.2 GB 84aba33d7f1e
TheDrummer_Orion-26B-A4B-v1-Q3_K_M.ggufWeights13.0 GB 1da65eb8c0b2
TheDrummer_Orion-26B-A4B-v1-Q3_K_S.ggufWeights12.6 GB 3f5c4f54d3ad
TheDrummer_Orion-26B-A4B-v1-Q3_K_XL.ggufWeights13.4 GB 0c044e6f820d
TheDrummer_Orion-26B-A4B-v1-Q4_0.ggufWeights14.8 GB ff5bdada1339
TheDrummer_Orion-26B-A4B-v1-Q4_1.ggufWeights16.1 GB 1375bc1d9303
TheDrummer_Orion-26B-A4B-v1-Q4_K_L.ggufWeights17.2 GB 04f64451b19a
TheDrummer_Orion-26B-A4B-v1-Q4_K_M.ggufWeights17.0 GB 036c70d3a8ff
TheDrummer_Orion-26B-A4B-v1-Q4_K_S.ggufWeights15.8 GB 887b4b8f7c5a
TheDrummer_Orion-26B-A4B-v1-Q5_K_L.ggufWeights19.5 GB 6b64e468aad0
TheDrummer_Orion-26B-A4B-v1-Q5_K_M.ggufWeights19.3 GB 623e36eea9c9
TheDrummer_Orion-26B-A4B-v1-Q5_K_S.ggufWeights18.1 GB 6fb8b14f7086
TheDrummer_Orion-26B-A4B-v1-Q6_K.ggufWeights22.9 GB b5203988c668
TheDrummer_Orion-26B-A4B-v1-Q6_K_L.ggufWeights23.0 GB 3c2964c5fb44
TheDrummer_Orion-26B-A4B-v1-Q8_0.ggufWeights26.9 GB a61a8f5555e2
TheDrummer_Orion-26B-A4B-v1-bf16.ggufWeights50.5 GB 3eb03d7bf91b
TheDrummer_Orion-26B-A4B-v1-imatrix.ggufWeights56.9 MB 69f29bf28b83
mmproj-TheDrummer_Orion-26B-A4B-v1-bf16.ggufWeights1.2 GB 21459d78a7d7
mmproj-TheDrummer_Orion-26B-A4B-v1-f16.ggufWeights1.2 GB d9432954aba5
README.mdDocumentation15.0 KB
TheDrummer_Orion-26B-A4B-v1-calibration-v6.txtOther1.0 MB
.gitattributesRepository3.7 KB

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
435.7 GB
Download from Bartowski

Released by Bartowski through its official repository on Hugging Face.

Built From

  • Derived from TheDrummer/Orion-26B-A4B-v1
  • Quantized from TheDrummer/Orion-26B-A4B-v1

Memory Requirements

PrecisionWeights in memory
As published435.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Similar Models

Model · Any to any

gemma-4-E4B-it-GGUF

GGML Org

Run with https://llama.app - https://huggingface.co/google/gemma-4-E4B-it - https://huggingface.co/google/gemma-4-E4B-it-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized - add info - add dflash

Open weights apache-2.0

Model · Any to any

gemma-4-12B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-12B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-12B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-E4B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0

Model · Any to any

gemma-4-E4B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E4B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers

Model · Any to any

gemma-4-E2B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E2B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers