SAVRN
Search Contact SAVRN

Open-weight model · Any to any

TheDrummer_Artemis-31B-v1.1-GGUF

by Bartowski bartowski/TheDrummer_Artemis-31B-v1.1-GGUF

Using llama.cpp release b10262 for quantization. Don't know which to choose? Grab Q4KM (19.60GB) - usually a good mix of size and performance.

Parameters
Context
Weights526.9 GB
License
AccessOpen weights
Monthly Downloads22.7k

Model Card

Using llama.cpp release b10262 for quantization. Don't know which to choose? Grab Q4KM (19.60GB) - usually a good mix of size and performance. Download instructions available here First, make sure you have the Hugging Face CLI installed: The files marked true in the Split column above are stored as multiple parts in a folder. To download all the parts to a local folder, run: You can either specify a new local-dir (TheDrummerArtemis-31B-v1.1-bf16) or download them all in place (./) These quants run with llama.cpp - installable in one line via llama.app: llama-server includes a built-in chat web UI, served at http://localhost:8080 by default. These quants were made with llama.cpp release…

Excerpt from the card by Bartowski.

Identity and Version

Repository
bartowski/TheDrummer_Artemis-31B-v1.1-GGUF
Publisher
Bartowski
Task
Any to any
Modality
Multimodal
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
853656cc5ba892c87a4b63b6c51e36f07e840f53
First published
2026-08-06
Last updated
2026-08-06

Files and Weights

33 files, 526.9 GB in total. The weights are 31 files totalling 526.9 GB in gguf.

Weights31 files · 526.9 GB
Documentation1 file · 13.3 KB
Repository1 file · 4.0 KB
Every file
FileTypeSizeSHA-256
TheDrummer_Artemis-31B-v1.1-IQ2_M.ggufWeights12.6 GB f27f17af9a39
TheDrummer_Artemis-31B-v1.1-IQ2_S.ggufWeights12.1 GB 073207d5709d
TheDrummer_Artemis-31B-v1.1-IQ2_XS.ggufWeights11.5 GB 143aa1094583
TheDrummer_Artemis-31B-v1.1-IQ2_XXS.ggufWeights10.8 GB d00a3e5c4ca3
TheDrummer_Artemis-31B-v1.1-IQ3_M.ggufWeights15.1 GB 837f9af2ac4f
TheDrummer_Artemis-31B-v1.1-IQ3_XS.ggufWeights13.8 GB 8296eb96d219
TheDrummer_Artemis-31B-v1.1-IQ3_XXS.ggufWeights13.0 GB 35cec3b908b3
TheDrummer_Artemis-31B-v1.1-IQ4_NL.ggufWeights18.0 GB 32a2f369a68e
TheDrummer_Artemis-31B-v1.1-IQ4_XS.ggufWeights17.2 GB ccbe3e3d8dad
TheDrummer_Artemis-31B-v1.1-Q2_K.ggufWeights12.6 GB 260cde15b9aa
TheDrummer_Artemis-31B-v1.1-Q2_K_L.ggufWeights13.0 GB e085b11bbd56
TheDrummer_Artemis-31B-v1.1-Q3_K_L.ggufWeights16.8 GB 6f6d5cf3568c
TheDrummer_Artemis-31B-v1.1-Q3_K_M.ggufWeights15.9 GB 68331ce9840a
TheDrummer_Artemis-31B-v1.1-Q3_K_S.ggufWeights14.3 GB 897e781eed10
TheDrummer_Artemis-31B-v1.1-Q3_K_XL.ggufWeights17.2 GB f98c572d9b97
TheDrummer_Artemis-31B-v1.1-Q4_0.ggufWeights18.1 GB 9acff1b23e47
TheDrummer_Artemis-31B-v1.1-Q4_1.ggufWeights19.8 GB 03975e46757a
TheDrummer_Artemis-31B-v1.1-Q4_K_L.ggufWeights19.9 GB 895525b4c846
TheDrummer_Artemis-31B-v1.1-Q4_K_M.ggufWeights19.6 GB 0637f8949910
TheDrummer_Artemis-31B-v1.1-Q4_K_S.ggufWeights18.2 GB 533206da2419
TheDrummer_Artemis-31B-v1.1-Q5_K_L.ggufWeights23.0 GB 90e4555b654b
TheDrummer_Artemis-31B-v1.1-Q5_K_M.ggufWeights22.6 GB 2cd61f47abef
TheDrummer_Artemis-31B-v1.1-Q5_K_S.ggufWeights21.5 GB 71ed4737458b
TheDrummer_Artemis-31B-v1.1-Q6_K.ggufWeights26.7 GB 7daf6d080117
TheDrummer_Artemis-31B-v1.1-Q6_K_L.ggufWeights27.1 GB 08ee8eb4159b
TheDrummer_Artemis-31B-v1.1-Q8_0.ggufWeights32.6 GB a3559a26f8f6
TheDrummer_Artemis-31B-v1.1-bf16/TheDrummer_Artemis-31B-v1.1-bf16-00001-of-00002.ggufWeights40.0 GB ce0ab5f2e4a9
TheDrummer_Artemis-31B-v1.1-bf16/TheDrummer_Artemis-31B-v1.1-bf16-00002-of-00002.ggufWeights21.4 GB 2f6f9be3f4de
TheDrummer_Artemis-31B-v1.1-imatrix.ggufWeights13.8 MB afd8a94b6729
mmproj-TheDrummer_Artemis-31B-v1.1-bf16.ggufWeights1.2 GB 780d66e554f4
mmproj-TheDrummer_Artemis-31B-v1.1-f16.ggufWeights1.2 GB ecd2c4de5246
README.mdDocumentation13.3 KB
.gitattributesRepository4.0 KB

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
526.9 GB
Download from Bartowski

Released by Bartowski through its official repository on Hugging Face.

Built From

  • Derived from TheDrummer/Artemis-31B-v1.1
  • Quantized from TheDrummer/Artemis-31B-v1.1

Memory Requirements

PrecisionWeights in memory
As published526.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Similar Models

Model · Any to any

gemma-4-E4B-it-GGUF

GGML Org

Run with https://llama.app - https://huggingface.co/google/gemma-4-E4B-it - https://huggingface.co/google/gemma-4-E4B-it-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized-assistant - https://huggingface.co/google/gemma-4-E4B-it-qat-q40-unquantized - add info - add dflash

Open weights apache-2.0

Model · Any to any

gemma-4-12B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-12B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-12B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 transformers

Model · Any to any

gemma-4-E4B-it-qat-q4_0-gguf

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0

Model · Any to any

gemma-4-E4B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E4B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers

Model · Any to any

gemma-4-E2B-it-qat-GGUF

Unsloth AI

This model ships a Multi-Token Prediction drafter at the repo root (mtp-gemma-4-E2B-it.gguf, a near-lossless smart Q40). A recent llama.cpp auto-discovers it from -hf, so you do not pass --model-draft: The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context…

Open weights apache-2.0 131,072 tokens transformers