SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

endless-frontier_BigBang-v1-GGUF

by Bartowski bartowski/endless-frontier_BigBang-v1-GGUF

Using llama.cpp release b10262 for quantization. Don't know which to choose? Grab Q4KM (21.86GB) - usually a good mix of size and performance.

Parameters
Context
Weights585.2 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads984.4k

Model Card

By Bartowski, published under apache-2.0, revision 3ecf951aa564.

Using llama.cpp release b10262 for quantization. Don't know which to choose? Grab Q4KM (21.86GB) - usually a good mix of size and performance. Download instructions available here First, make sure you have the Hugging Face CLI installed: The files marked true in the Split column above are stored as multiple parts in a folder. To download all the parts to a local folder, run: You can either specify a new local-dir (endless-frontierBigBang-v1-bf16) or download them all in place (./) These quants run with llama.cpp - installable in one line via llama.app: llama-server includes a built-in chat web UI, served at http://localhost:8080 by default. These quants were made with llama.cpp release…

Read Bartowski's full model card

Llamacpp imatrix Quantizations of BigBang-v1 by endless-frontier

Using llama.cpp release b10262 for quantization.

Original model: https://huggingface.co/endless-frontier/BigBang-v1

Model details: - Parameter count: 36B - Input support: text, image (with mmproj file) - details - MTP: yes - details - imatrix: yes - details

How to run

Prompt format

<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
<think>

Don't know which to choose? Grab Q4_K_M (21.86GB) - usually a good mix of size and performance. Download instructions available here

Available files:

Filename Quant type File Size Split Description
endless-frontier_BigBang-v1-bf16.gguf bf16 71.07GB true Full BF16 weights.
endless-frontier_BigBang-v1-Q8_0.gguf Q8_0 37.81GB false Extremely high quality, generally unneeded but max available quant.
endless-frontier_BigBang-v1-Q6_K_L.gguf Q6_K_L 30.77GB false Uses Q8_0 for embed and output weights. Very high quality, near perfect, recommended.
endless-frontier_BigBang-v1-Q6_K.gguf Q6_K 30.53GB false Very high quality, near perfect, recommended.
endless-frontier_BigBang-v1-Q5_K_L.gguf Q5_K_L 25.81GB false Uses Q8_0 for embed and output weights. High quality, recommended.
endless-frontier_BigBang-v1-Q5_K_M.gguf Q5_K_M 25.49GB false High quality, recommended.
endless-frontier_BigBang-v1-Q5_K_S.gguf Q5_K_S 24.63GB false High quality, recommended.
endless-frontier_BigBang-v1-Q4_1.gguf Q4_1 22.45GB false Legacy format, similar performance to Q4_K_S but with improved tokens/watt on Apple silicon.
endless-frontier_BigBang-v1-Q4_K_L.gguf Q4_K_L 22.24GB false Uses Q8_0 for embed and output weights. Good quality, recommended.
endless-frontier_BigBang-v1-Q4_K_M.gguf Q4_K_M 21.86GB false Good quality, default size for most use cases, recommended.
endless-frontier_BigBang-v1-Q4_K_S.gguf Q4_K_S 21.07GB false Slightly lower quality with more space savings, recommended.
endless-frontier_BigBang-v1-Q4_0.gguf Q4_0 20.42GB false Legacy format, kept for compatibility with older tools.
endless-frontier_BigBang-v1-IQ4_NL.gguf IQ4_NL 20.33GB false Similar to IQ4_XS, but slightly larger.
endless-frontier_BigBang-v1-IQ4_XS.gguf IQ4_XS 19.28GB false Decent quality, smaller than Q4_K_S with similar performance, recommended.
endless-frontier_BigBang-v1-Q3_K_XL.gguf Q3_K_XL 17.80GB false Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability.
endless-frontier_BigBang-v1-IQ3_M.gguf IQ3_M 17.37GB false Medium-low quality, new method with decent performance comparable to Q3_K_M.
endless-frontier_BigBang-v1-Q3_K_L.gguf Q3_K_L 17.36GB false Lower quality but usable, good for low RAM availability.
endless-frontier_BigBang-v1-Q3_K_M.gguf Q3_K_M 16.70GB false Low quality.
endless-frontier_BigBang-v1-IQ3_XS.gguf IQ3_XS 16.69GB false Lower quality, new method with decent performance, slightly better than Q3_K_S.
endless-frontier_BigBang-v1-Q3_K_S.gguf Q3_K_S 15.98GB false Low quality, not recommended.
endless-frontier_BigBang-v1-IQ3_XXS.gguf IQ3_XXS 15.34GB false Lower quality, new method with decent performance, comparable to Q3 quants.
endless-frontier_BigBang-v1-Q2_K_L.gguf Q2_K_L 13.58GB false Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable.
endless-frontier_BigBang-v1-Q2_K.gguf Q2_K 13.09GB false Very low quality but surprisingly usable.
endless-frontier_BigBang-v1-IQ2_M.gguf IQ2_M 12.54GB false Relatively low quality, uses SOTA techniques to be surprisingly usable.
endless-frontier_BigBang-v1-IQ2_S.gguf IQ2_S 11.49GB false Low quality, uses SOTA techniques to be usable.
endless-frontier_BigBang-v1-IQ2_XS.gguf IQ2_XS 11.27GB false Low quality, uses SOTA techniques to be usable.
endless-frontier_BigBang-v1-IQ2_XXS.gguf IQ2_XXS 10.26GB false Very low quality, uses SOTA techniques to be usable.

Download a specific file:

hf download bartowski/endless-frontier_BigBang-v1-GGUF --include "endless-frontier_BigBang-v1-Q4_K_M.gguf" --local-dir ./

Downloading using the Hugging Face CLI

Click to view download instructions First, make sure you have the Hugging Face CLI installed:
pip install -U "huggingface_hub[cli]"
Download a specific file:
hf download bartowski/endless-frontier_BigBang-v1-GGUF --include "endless-frontier_BigBang-v1-Q4_K_M.gguf" --local-dir ./
The files marked `true` in the Split column above are stored as multiple parts in a folder. To download all the parts to a local folder, run:
hf download bartowski/endless-frontier_BigBang-v1-GGUF --include "endless-frontier_BigBang-v1-bf16/*" --local-dir ./
You can either specify a new local-dir (endless-frontier_BigBang-v1-bf16) or download them all in place (./)

How to run

These quants run with llama.cpp - installable in one line via llama.app:

curl -LsSf https://llama.app/install.sh | sh
llama-server -hf bartowski/endless-frontier_BigBang-v1-GGUF:Q4_K_M

llama-server includes a built-in chat web UI, served at http://localhost:8080 by default.

These quants were made with llama.cpp release b10262 - if this model's architecture is newly supported, you'll need that release or newer to run them.

They also work in: LM Studio · koboldcpp · ramalama · Jan AI · Text Generation Web UI · LoLLMs · Atomic Chat

Multimodal

This model supports image input. Alongside the quants, this repo includes the multimodal projector files mmproj-endless-frontier_BigBang-v1-f16.gguf and mmproj-endless-frontier_BigBang-v1-bf16.gguf, which pair with any quant above.

llama.cpp downloads the mmproj automatically when using -hf as shown above; if you're loading files manually, pass it with --mmproj.

MTP

This model has MTP (Multi-Token Prediction) layers, and they are included in these quants

MTP layers act as a built-in draft model, letting llama.cpp run speculative decoding for faster generation. To use them, add the following flag to your llama.cpp command:

--spec-type draft-mtp

Note: the MTP layers are stored at Q4_0 in the imatrix quants (except for the Q8_0 quant), since imatrix calibration does not exercise them. Q4_0 is chosen for its speed which massively benefits MTP performance.

imatrix

All quants made using imatrix option with dataset from here. The imatrix is available here: endless-frontier_BigBang-v1-imatrix.gguf.

Embed/output weights

Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.

ARM/AVX information

llama.cpp automatically "repacks" weights into an interleaved layout at load time for faster inference on ARM and AVX machines - details in this PR. This once required downloading special Q4_0_4_4/4_8/8_8 files; those are long gone. Online repacking now covers Q4_0, IQ4_NL, and most K-quants, so no special quant choice is needed for CPU inference.

Which file should I choose?

Click here for details An older (early 2024) but still useful write-up with charts comparing quant performances is provided by Artefact2 [here](https://gist.github.com/Artefact2/b5f810600771265fc1e39442288e8ec9) The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have. If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM. If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total. Hugging Face can also do this math for you: add your hardware in your [Local Apps settings](https://huggingface.co/settings/local-apps) and the model page will show which files fit. Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'. If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M. If you want to get more into the weeds, you can check out this extremely useful feature chart: [llama.cpp feature matrix](https://github.com/ggml-org/llama.cpp/wiki/Feature-matrix) But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size. These I-quants can also be used on CPU, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.

Credits

Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.

Thank you ZeroWw for the inspiration to experiment with embed/output.

Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

Identity and Version

Repository
bartowski/endless-frontier_BigBang-v1-GGUF
Publisher
Bartowski
Task
Image and text to text
Modality
Image and text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en
Revision
3ecf951aa564e7ef610ddb25f8196d03a80b80d3
First published
2026-08-07
Last updated
2026-08-07

Files and Weights

33 files, 585.2 GB in total. The weights are 31 files totalling 585.2 GB in gguf.

Weights31 files · 585.2 GB
Documentation1 file · 14.0 KB
Repository1 file · 4.0 KB
Every file
FileTypeSizeSHA-256
endless-frontier_BigBang-v1-IQ2_M.ggufWeights12.5 GB 5a799b51d6c2
endless-frontier_BigBang-v1-IQ2_S.ggufWeights11.5 GB 1dac5c04bf92
endless-frontier_BigBang-v1-IQ2_XS.ggufWeights11.3 GB 4df8d242670f
endless-frontier_BigBang-v1-IQ2_XXS.ggufWeights10.3 GB eaa55422a5c8
endless-frontier_BigBang-v1-IQ3_M.ggufWeights17.4 GB de81371bc893
endless-frontier_BigBang-v1-IQ3_XS.ggufWeights16.7 GB 357da439d789
endless-frontier_BigBang-v1-IQ3_XXS.ggufWeights15.3 GB 0052b0ae138d
endless-frontier_BigBang-v1-IQ4_NL.ggufWeights20.3 GB e073faa4ecc1
endless-frontier_BigBang-v1-IQ4_XS.ggufWeights19.3 GB 1e7e3b952d1f
endless-frontier_BigBang-v1-Q2_K.ggufWeights13.1 GB 17b5e7c04cad
endless-frontier_BigBang-v1-Q2_K_L.ggufWeights13.6 GB 5cac2ec21dee
endless-frontier_BigBang-v1-Q3_K_L.ggufWeights17.4 GB dbd5bd2ff034
endless-frontier_BigBang-v1-Q3_K_M.ggufWeights16.7 GB dc00761483e0
endless-frontier_BigBang-v1-Q3_K_S.ggufWeights16.0 GB 7471830e0839
endless-frontier_BigBang-v1-Q3_K_XL.ggufWeights17.8 GB 5df96087837f
endless-frontier_BigBang-v1-Q4_0.ggufWeights20.4 GB 6d9e17ac3705
endless-frontier_BigBang-v1-Q4_1.ggufWeights22.4 GB 9978522e8dd5
endless-frontier_BigBang-v1-Q4_K_L.ggufWeights22.2 GB 6f8a57951afd
endless-frontier_BigBang-v1-Q4_K_M.ggufWeights21.9 GB ecd382669651
endless-frontier_BigBang-v1-Q4_K_S.ggufWeights21.1 GB a80c909dd9ff
endless-frontier_BigBang-v1-Q5_K_L.ggufWeights25.8 GB 0a37061d89ea
endless-frontier_BigBang-v1-Q5_K_M.ggufWeights25.5 GB ccb848bc501e
endless-frontier_BigBang-v1-Q5_K_S.ggufWeights24.6 GB f9f02de5006c
endless-frontier_BigBang-v1-Q6_K.ggufWeights30.5 GB ca76cc0cec8e
endless-frontier_BigBang-v1-Q6_K_L.ggufWeights30.8 GB cc6fbae1144c
endless-frontier_BigBang-v1-Q8_0.ggufWeights37.8 GB e798934a94a0
endless-frontier_BigBang-v1-bf16/endless-frontier_BigBang-v1-bf16-00001-of-00002.ggufWeights39.9 GB e8ae2a1ec599
endless-frontier_BigBang-v1-bf16/endless-frontier_BigBang-v1-bf16-00002-of-00002.ggufWeights31.1 GB 83601dc51faa
endless-frontier_BigBang-v1-imatrix.ggufWeights192.2 MB 05eb8672e7ff
mmproj-endless-frontier_BigBang-v1-bf16.ggufWeights902.8 MB 2f5c1945c3bb
mmproj-endless-frontier_BigBang-v1-f16.ggufWeights899.3 MB 5a6a19c9ea91
README.mdDocumentation14.0 KB
.gitattributesRepository4.0 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
585.2 GB
Download from Bartowski

Released by Bartowski through its official repository on Hugging Face. Read the license.

Built From

  • Derived from endless-frontier/BigBang-v1
  • Quantized from endless-frontier/BigBang-v1

Memory Requirements

PrecisionWeights in memory
As published585.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About endless-frontier_BigBang-v1-GGUF

Can I use endless-frontier_BigBang-v1-GGUF commercially?

Yes. endless-frontier_BigBang-v1-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Image and text to text

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

Michał Piszczek

I built this quant because the ready-made FP4 file answered the wrong question. It was fast, but on my short WikiText-2 control it scored 6.4949 PPL. Plain Q40 scored 6.3798. The first higher-quality hybrid went too far the other way: good perplexity, 34.19 tok/s, and no comfortable room for 256K plus vision. This is the build that survived both gates. It is a 17.1 GB, 5.01 BPW mixed-precision GGUF of Qwen/Qwen3.8-27B. It keeps large, tolerant matrices in native NVFP4 and spends more bits on selected attention, Gated DeltaNet, and late FFN tensors. The trained MTP layer remains embedded in the same GGUF. This is not a fine-tune. I built the private calibration workload from 5,472 messages…

Open weights apache-2.0

Model · Image and text to text

Huihui-Qwen3.8-27B-abliterated-GGUF

Huihui.ai

This is an uncensored version of Qwen/Qwen3.8-27B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. The newly added Huihui-Qwen3.8-27B-abliterated-GSQ-RCO series come from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. Only layers 23 to 51 have been ablated, while the other layers remain unablated. It may come with a small disclaimer warning. The size after conversion may differ from the original GGUF. The newly added Huihui-Qwen3.8-27B-abliterated-UD series come from unsloth/Qwen3.8-27B-GGUF. Only layers 18 to 51 have been ablated(Previously…

Open weights apache-2.0 transformers

Qwen3.8-27B uncensored by HauhauCS 0/465 Refusals. This is the Aggressive variant: direct answers, no refusal behavior, and minimal preamble on hard prompts. Every text GGUF preserves Qwen3.8's native NextN head, and this release adds HauhauCS FastMTP: a specific acceleration sidecar qualified across the complete quant lineup at maximum native context. Vision is included through the separate BF16 projector. No changes to datasets or intended capabilities. This release preserves Qwen3.8-27B's text, reasoning, agentic, image, and video capabilities while applying the HauhauCS Aggressive uncensoring profile. Pick Aggressive when you specifically want the model to get to the answer without…

Open weights apache-2.0

Model · Image and text to text

Gemma-4-E4B-Uncensored-HauhauCS-Aggressive

HauhauCS

Gemma 4 E4B-IT uncensored by HauhauCS. 0/465 Refusals\ No changes to datasets or capabilities. Fully functional, 100% of what the original authors intended - just without the refusals. These are meant to be the best lossless uncensored models out there. Stronger uncensoring — model is fully unlocked and won't refuse prompts. May occasionally append short disclaimers (baked into base model training, not refusals) but full content is always generated. For a more conservative uncensor that keeps some safety guardrails, check the Balanced variant when it's available. All quants generated with importance matrix (imatrix) for optimal quality preservation on abliterated weights. KP ("Perfect")…

Open weights gemma

Model · Image and text to text

Qwen3.5-9B-GGUF

Unsloth AI

You can now also fine-tune the model locally with Unsloth. - Read our Qwen3.5 fine-tuning guide here. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty…

Open weights apache-2.0 transformers

Model · Image and text to text

Qwen3.8-Flash-Next-GGUF

Unsloth AI

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other