SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

XYZAILab_XYZ-Aquila-mini-GGUF

by Bartowski bartowski/XYZAILab_XYZ-Aquila-mini-GGUF

Using llama.cpp release b10142 for quantization. All quants made using imatrix option with dataset from here Run them in your choice of tools: Note: if it's a newly supported model, you may need to wait for an update from the developers.

Parameters
Context
Weights570.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads883.9k

Model Card

By Bartowski, published under apache-2.0, revision c1359d500cd5.

Using llama.cpp release b10142 for quantization. All quants made using imatrix option with dataset from here Run them in your choice of tools: Note: if it's a newly supported model, you may need to wait for an update from the developers. Some of these quants (Q3KXL, Q4KL etc) are the standard quantization method with the embeddings and output weights quantized to Q80 instead of what they would normally default to. First, make sure you have huggingface-cli installed: Then, you can target the specific file you want: If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run: You can either specify a new local-dir…

Read Bartowski's full model card

Llamacpp imatrix Quantizations of XYZ-Aquila-mini by XYZAILab

Using llama.cpp release b10142 for quantization.

Original model: https://huggingface.co/XYZAILab/XYZ-Aquila-mini

All quants made using imatrix option with dataset from here

Run them in your choice of tools:

Note: if it's a newly supported model, you may need to wait for an update from the developers.

Prompt format

<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
<think>

Download a file (not the whole branch) from below:

Filename Quant type File Size Split Description
XYZAILab_XYZ-Aquila-mini-bf16.gguf bf16 69.38GB true Full BF16 weights.
XYZAILab_XYZ-Aquila-mini-Q8_0.gguf Q8_0 36.91GB false Extremely high quality, generally unneeded but max available quant.
XYZAILab_XYZ-Aquila-mini-Q6_K_L.gguf Q6_K_L 30.30GB false Uses Q8_0 for embed and output weights. Very high quality, near perfect, recommended.
XYZAILab_XYZ-Aquila-mini-Q6_K.gguf Q6_K 30.05GB false Very high quality, near perfect, recommended.
XYZAILab_XYZ-Aquila-mini-Q5_K_L.gguf Q5_K_L 25.33GB false Uses Q8_0 for embed and output weights. High quality, recommended.
XYZAILab_XYZ-Aquila-mini-Q5_K_M.gguf Q5_K_M 25.02GB false High quality, recommended.
XYZAILab_XYZ-Aquila-mini-Q5_K_S.gguf Q5_K_S 24.16GB false High quality, recommended.
XYZAILab_XYZ-Aquila-mini-Q4_1.gguf Q4_1 21.97GB false Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon.
XYZAILab_XYZ-Aquila-mini-Q4_K_L.gguf Q4_K_L 21.77GB false Uses Q8_0 for embed and output weights. Good quality, recommended.
XYZAILab_XYZ-Aquila-mini-Q4_K_M.gguf Q4_K_M 21.39GB false Good quality, default size for most use cases, recommended.
XYZAILab_XYZ-Aquila-mini-Q4_K_S.gguf Q4_K_S 20.59GB false Slightly lower quality with more space savings, recommended.
XYZAILab_XYZ-Aquila-mini-Q4_0.gguf Q4_0 19.94GB false Legacy format, offers online repacking for ARM and AVX CPU inference.
XYZAILab_XYZ-Aquila-mini-IQ4_NL.gguf IQ4_NL 19.86GB false Similar to IQ4_XS, but slightly larger. Offers online repacking for ARM CPU inference.
XYZAILab_XYZ-Aquila-mini-IQ4_XS.gguf IQ4_XS 18.81GB false Decent quality, smaller than Q4_K_S with similar performance, recommended.
XYZAILab_XYZ-Aquila-mini-Q3_K_XL.gguf Q3_K_XL 17.33GB false Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability.
XYZAILab_XYZ-Aquila-mini-IQ3_M.gguf IQ3_M 16.90GB false Medium-low quality, new method with decent performance comparable to Q3_K_M.
XYZAILab_XYZ-Aquila-mini-Q3_K_L.gguf Q3_K_L 16.89GB false Lower quality but usable, good for low RAM availability.
XYZAILab_XYZ-Aquila-mini-Q3_K_M.gguf Q3_K_M 16.23GB false Low quality.
XYZAILab_XYZ-Aquila-mini-IQ3_XS.gguf IQ3_XS 16.22GB false Lower quality, new method with decent performance, slightly better than Q3_K_S.
XYZAILab_XYZ-Aquila-mini-Q3_K_S.gguf Q3_K_S 15.51GB false Low quality, not recommended.
XYZAILab_XYZ-Aquila-mini-IQ3_XXS.gguf IQ3_XXS 14.87GB false Lower quality, new method with decent performance, comparable to Q3 quants.
XYZAILab_XYZ-Aquila-mini-Q2_K_L.gguf Q2_K_L 13.11GB false Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable.
XYZAILab_XYZ-Aquila-mini-Q2_K.gguf Q2_K 12.62GB false Very low quality but surprisingly usable.
XYZAILab_XYZ-Aquila-mini-IQ2_M.gguf IQ2_M 12.07GB false Relatively low quality, uses SOTA techniques to be surprisingly usable.
XYZAILab_XYZ-Aquila-mini-IQ2_S.gguf IQ2_S 11.01GB false Low quality, uses SOTA techniques to be usable.
XYZAILab_XYZ-Aquila-mini-IQ2_XS.gguf IQ2_XS 10.80GB false Low quality, uses SOTA techniques to be usable.
XYZAILab_XYZ-Aquila-mini-IQ2_XXS.gguf IQ2_XXS 9.78GB false Very low quality, uses SOTA techniques to be usable.

Embed/output weights

Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.

Downloading using huggingface-cli

Click to view download instructions First, make sure you have huggingface-cli installed:
pip install -U "huggingface_hub[cli]"
Then, you can target the specific file you want:
huggingface-cli download bartowski/XYZAILab_XYZ-Aquila-mini-GGUF --include "XYZAILab_XYZ-Aquila-mini-Q4_K_M.gguf" --local-dir ./
If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:
huggingface-cli download bartowski/XYZAILab_XYZ-Aquila-mini-GGUF --include "XYZAILab_XYZ-Aquila-mini-Q8_0/*" --local-dir ./
You can either specify a new local-dir (XYZAILab_XYZ-Aquila-mini-Q8_0) or download them all in place (./)

ARM/AVX information

Previously, you would download Q4_0_4_4/4_8/8_8, and these would have their weights interleaved in memory in order to improve performance on ARM and AVX machines by loading up more data in one pass.

Now, however, there is something called "online repacking" for weights. details in this PR. If you use Q4_0 and your hardware would benefit from repacking weights, it will do it automatically on the fly.

As of llama.cpp build b4282 you will not be able to run the Q4_0_X_X files and will instead need to use Q4_0.

Additionally, if you want to get slightly better quality, you can use IQ4_NL thanks to this PR which will also repack the weights for ARM, though only the 4_4 for now. The loading time may be slower but it will result in an overall speed increase.

Click to view Q4_0_X_X information (deprecated) I'm keeping this section to show the potential theoretical uplift in performance from using the Q4_0 with online repacking.
Click to view benchmarks on an AVX2 system (EPYC7702) | model | size | params | backend | threads | test | t/s | % (vs Q4_0) | | ------------------------------ | ---------: | ---------: | ---------- | ------: | ------------: | -------------------: |-------------: | | qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp512 | 204.03 ± 1.03 | 100% | | qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp1024 | 282.92 ± 0.19 | 100% | | qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | pp2048 | 259.49 ± 0.44 | 100% | | qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg128 | 39.12 ± 0.27 | 100% | | qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg256 | 39.31 ± 0.69 | 100% | | qwen2 3B Q4_0 | 1.70 GiB | 3.09 B | CPU | 64 | tg512 | 40.52 ± 0.03 | 100% | | qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp512 | 301.02 ± 1.74 | 147% | | qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp1024 | 287.23 ± 0.20 | 101% | | qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | pp2048 | 262.77 ± 1.81 | 101% | | qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg128 | 18.80 ± 0.99 | 48% | | qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg256 | 24.46 ± 3.04 | 83% | | qwen2 3B Q4_K_M | 1.79 GiB | 3.09 B | CPU | 64 | tg512 | 36.32 ± 3.59 | 90% | | qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp512 | 271.71 ± 3.53 | 133% | | qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp1024 | 279.86 ± 45.63 | 100% | | qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | pp2048 | 320.77 ± 5.00 | 124% | | qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg128 | 43.51 ± 0.05 | 111% | | qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg256 | 43.35 ± 0.09 | 110% | | qwen2 3B Q4_0_8_8 | 1.69 GiB | 3.09 B | CPU | 64 | tg512 | 42.60 ± 0.31 | 105% | Q4_0_8_8 offers a nice bump to prompt processing and a small bump to text generation

Which file should I choose?

Click here for details A great write up with charts showing various performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9) The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have. If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM. If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total. Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'. If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M. If you want to get more into the weeds, you can check out this extremely useful feature chart: [llama.cpp feature matrix](https://github.com/ggml-org/llama.cpp/wiki/Feature-matrix) But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size. These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.

Credits

Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.

Thank you ZeroWw for the inspiration to experiment with embed/output.

Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

Identity and Version

Repository
bartowski/XYZAILab_XYZ-Aquila-mini-GGUF
Publisher
Bartowski
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
c1359d500cd53588fdb5a7f35b6e61e9ad3e5fab
First published
2026-07-28
Last updated
2026-07-28

Files and Weights

33 files, 570.8 GB in total. The weights are 31 files totalling 570.8 GB in gguf.

Weights31 files · 570.8 GB
Documentation1 file · 14.9 KB
Repository1 file · 3.9 KB
Every file
FileTypeSizeSHA-256
XYZAILab_XYZ-Aquila-mini-IQ2_M.ggufWeights12.1 GB 7804f08e4a4b
XYZAILab_XYZ-Aquila-mini-IQ2_S.ggufWeights11.0 GB b8eda693e3d7
XYZAILab_XYZ-Aquila-mini-IQ2_XS.ggufWeights10.8 GB 29e2c41015ab
XYZAILab_XYZ-Aquila-mini-IQ2_XXS.ggufWeights9.8 GB 490f4e58843c
XYZAILab_XYZ-Aquila-mini-IQ3_M.ggufWeights16.9 GB 75bfbfabd194
XYZAILab_XYZ-Aquila-mini-IQ3_XS.ggufWeights16.2 GB c06fe241303f
XYZAILab_XYZ-Aquila-mini-IQ3_XXS.ggufWeights14.9 GB d36e7289a2c6
XYZAILab_XYZ-Aquila-mini-IQ4_NL.ggufWeights19.9 GB 8073cb5be9b0
XYZAILab_XYZ-Aquila-mini-IQ4_XS.ggufWeights18.8 GB 5b145576a5fd
XYZAILab_XYZ-Aquila-mini-Q2_K.ggufWeights12.6 GB 64bbf736f347
XYZAILab_XYZ-Aquila-mini-Q2_K_L.ggufWeights13.1 GB 1f9eab19c651
XYZAILab_XYZ-Aquila-mini-Q3_K_L.ggufWeights16.9 GB da9487a17bd8
XYZAILab_XYZ-Aquila-mini-Q3_K_M.ggufWeights16.2 GB 272167fc9149
XYZAILab_XYZ-Aquila-mini-Q3_K_S.ggufWeights15.5 GB 7156cd1cc5ad
XYZAILab_XYZ-Aquila-mini-Q3_K_XL.ggufWeights17.3 GB 3e26abf2fc40
XYZAILab_XYZ-Aquila-mini-Q4_0.ggufWeights19.9 GB b590735735df
XYZAILab_XYZ-Aquila-mini-Q4_1.ggufWeights22.0 GB 0f7289f00732
XYZAILab_XYZ-Aquila-mini-Q4_K_L.ggufWeights21.8 GB 9a291a1e650f
XYZAILab_XYZ-Aquila-mini-Q4_K_M.ggufWeights21.4 GB c96ee35332dc
XYZAILab_XYZ-Aquila-mini-Q4_K_S.ggufWeights20.6 GB bbe9717b6059
XYZAILab_XYZ-Aquila-mini-Q5_K_L.ggufWeights25.3 GB 063890994ea4
XYZAILab_XYZ-Aquila-mini-Q5_K_M.ggufWeights25.0 GB fbaddbee1abc
XYZAILab_XYZ-Aquila-mini-Q5_K_S.ggufWeights24.2 GB 12381d7cf317
XYZAILab_XYZ-Aquila-mini-Q6_K.ggufWeights30.1 GB c66485a0c0d4
XYZAILab_XYZ-Aquila-mini-Q6_K_L.ggufWeights30.3 GB 36d9fce40da1
XYZAILab_XYZ-Aquila-mini-Q8_0.ggufWeights36.9 GB f3d4b19c2409
XYZAILab_XYZ-Aquila-mini-bf16/XYZAILab_XYZ-Aquila-mini-bf16-00001-of-00002.ggufWeights39.8 GB d4efb5f3e62a
XYZAILab_XYZ-Aquila-mini-bf16/XYZAILab_XYZ-Aquila-mini-bf16-00002-of-00002.ggufWeights29.6 GB 05af7e3a37e3
XYZAILab_XYZ-Aquila-mini-imatrix.ggufWeights192.2 MB 328c579c0cf3
mmproj-XYZAILab_XYZ-Aquila-mini-bf16.ggufWeights902.8 MB 44375c62ac42
mmproj-XYZAILab_XYZ-Aquila-mini-f16.ggufWeights899.3 MB b80c0f20a6ca
README.mdDocumentation14.9 KB
.gitattributesRepository3.9 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
570.8 GB
Download from Bartowski

Released by Bartowski through its official repository on Hugging Face. Read the license.

Built From

  • Derived from XYZAILab/XYZ-Aquila-mini
  • Quantized from XYZAILab/XYZ-Aquila-mini

Memory Requirements

PrecisionWeights in memory
As published570.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About XYZAILab_XYZ-Aquila-mini-GGUF

Can I use XYZAILab_XYZ-Aquila-mini-GGUF commercially?

Yes. XYZAILab_XYZ-Aquila-mini-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-GSQ-RCO series come from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 23 to 51 have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. The size after conversion may differ from the original GGUF. The newly added Huihui-Qwen3.8-27B-abliterated-UD series come from unsloth/Qwen3.8-27B-GGUF. Only layers 18 to 51 have been ablated(Previously…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

Model · Image and text to text

Gemma-4-E4B-Uncensored-HauhauCS-Aggressive

HauhauCS

Gemma 4 E4B-IT uncensored by HauhauCS. 0/465 Refusals\ No changes to datasets or capabilities. Fully functional, 100% of what the original authors intended - just without the refusals. These are meant to be the best lossless uncensored models out there. Stronger uncensoring — model is fully unlocked and won't refuse prompts. May occasionally append short disclaimers (baked into base model training, not refusals) but full content is always generated. For a more conservative uncensor that keeps some safety guardrails, check the Balanced variant when it's available. All quants generated with importance matrix (imatrix) for optimal quality preservation on abliterated weights. KP ("Perfect")…

Open weights gemma

Model · Image and text to text

Qwen3.5-9B-GGUF

Unsloth AI

You can now also fine-tune the model locally with Unsloth. - Read our Qwen3.5 fine-tuning guide here. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty…

Open weights apache-2.0 transformers

Model · Image and text to text

Qwen3.8-Flash-Next-GGUF

Unsloth AI

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other