SAVRN
Search Contact SAVRN

Open-weight model · Text to image

Z-Image-Turbo-GGUF

by Unsloth AI unsloth/Z-Image-Turbo-GGUF

Welcome to the official repository for the Z-Image(造相)project! Z-Image is a powerful and highly efficient image generation model with 6B parameters.

Parameters
Context
Weights90.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads285.3k

Model Card

By Unsloth AI, published under apache-2.0, revision 6c80814333b7.

Welcome to the official repository for the Z-Image(造相)project! Z-Image is a powerful and highly efficient image generation model with 6B parameters. Currently there are three variants: - Z-Image-Turbo – A distilled version of Z-Image that matches or exceeds leading competitors with only 8 NFEs (Number of Function Evaluations). It offers sub-second inference latency on enterprise-grade H800 GPUs and fits comfortably within 16G VRAM consumer devices. It excels in photorealistic image generation, bilingual text rendering (English & Chinese), and robust instruction adherence. - Z-Image-Base – The non-distilled foundation model. By releasing this checkpoint, we aim to unlock the full potential…

Read Unsloth AI's full model card

[!NOTE] This is a GGUF quantized version of Z-Image-Turbo. unsloth/Z-Image-Turbo-GGUF uses Unsloth Dynamic 2.0 methodology for SOTA performance. Important layers are upcasted to higher precision.

Samples


- Image
An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

[![Official Site](https://img.shields.io/badge/Official%20Site-333399.svg?logo=homepage)](https://tongyi-mai.github.io/Z-Image-blog/)  [![GitHub](https://img.shields.io/badge/GitHub-Z--Image-181717?logo=github&logoColor=white)](https://github.com/Tongyi-MAI/Z-Image)  [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Checkpoint-Z--Image--Turbo-yellow)](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo)  [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Online_Demo-Z--Image--Turbo-blue)](https://huggingface.co/spaces/Tongyi-MAI/Z-Image-Turbo)  [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Mobile_Demo-Z--Image--Turbo-red)](https://huggingface.co/spaces/akhaliq/Z-Image-Turbo)  [![ModelScope Model](https://img.shields.io/badge/%20Checkpoint-Z--Image--Turbo-624aff)](https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo)  [![ModelScope Space](https://img.shields.io/badge/%20Online_Demo-Z--Image--Turbo-17c7a7)](https://www.modelscope.cn/aigc/imageGeneration?tab=advanced&versionId=469191&modelType=Checkpoint&sdVersion=Z_IMAGE_TURBO&modelUrl=modelscope%253A%252F%252FTongyi-MAI%252FZ-Image-Turbo%253Frevision%253Dmaster%7D%7BOnline)  [![Art Gallery PDF](https://img.shields.io/badge/%F0%9F%96%BC%20Art_Gallery-PDF-ff69b4)](assets/Z-Image-Gallery.pdf)  [![Web Art Gallery](https://img.shields.io/badge/%F0%9F%8C%90%20Web_Art_Gallery-online-00bfff)](https://modelscope.cn/studios/Tongyi-MAI/Z-Image-Gallery/summary)  Welcome to the official repository for the Z-Image(造相)project!

Z-Image

Z-Image is a powerful and highly efficient image generation model with 6B parameters. Currently there are three variants:

  • Z-Image-Turbo – A distilled version of Z-Image that matches or exceeds leading competitors with only 8 NFEs (Number of Function Evaluations). It offers sub-second inference latency on enterprise-grade H800 GPUs and fits comfortably within 16G VRAM consumer devices. It excels in photorealistic image generation, bilingual text rendering (English & Chinese), and robust instruction adherence.

  • Z-Image-Base – The non-distilled foundation model. By releasing this checkpoint, we aim to unlock the full potential for community-driven fine-tuning and custom development.

  • Z-Image-Edit – A variant fine-tuned on Z-Image specifically for image editing tasks. It supports creative image-to-image generation with impressive instruction-following capabilities, allowing for precise edits based on natural language prompts.

Model Zoo

Model Hugging Face ModelScope
Z-Image-Turbo

Z-Image-Base To be released To be released
Z-Image-Edit To be released To be released

Showcase

Photorealistic Quality: Z-Image-Turbo delivers strong photorealistic image generation while maintaining excellent aesthetic quality.

Accurate Bilingual Text Rendering: Z-Image-Turbo excels at accurately rendering complex Chinese and English text.

Prompt Enhancing & Reasoning: Prompt Enhancer empowers the model with reasoning capabilities, enabling it to transcend surface-level descriptions and tap into underlying world knowledge.

Creative Image Editing: Z-Image-Edit shows a strong understanding of bilingual editing instructions, enabling imaginative and flexible image transformations.

Model Architecture

We adopt a Scalable Single-Stream DiT (S3-DiT) architecture. In this setup, text, visual semantic tokens, and image VAE tokens are concatenated at the sequence level to serve as a unified input stream, maximizing parameter efficiency compared to dual-stream approaches.

Performance

According to the Elo-based Human Preference Evaluation (on Alibaba AI Arena), Z-Image-Turbo shows highly competitive performance against other leading models, while achieving state-of-the-art results among open-source models.


Click to view the full leaderboard

Quick Start

Install the latest version of diffusers, use the following command:

Click here for details for why you need to install diffusers from source We have submitted two pull requests ([#12703](https://github.com/huggingface/diffusers/pull/12703) and [#12715](https://github.com/huggingface/diffusers/pull/12715)) to the diffusers repository to add support for Z-Image. Both PRs have been merged into the latest official diffusers release. Therefore, you need to install diffusers from source for the latest features and Z-Image support.
pip install git+https://github.com/huggingface/diffusers
import torch
from diffusers import ZImagePipeline

# 1. Load the pipeline
# Use bfloat16 for optimal performance on supported GPUs
pipe = ZImagePipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
torch_dtype=torch.bfloat16,
low_cpu_mem_usage=False,
)
pipe.to("cuda")

# [Optional] Attention Backend
# Diffusers uses SDPA by default. Switch to Flash Attention for better efficiency if supported:
# pipe.transformer.set_attention_backend("flash") # Enable Flash-Attention-2
# pipe.transformer.set_attention_backend("_flash_3") # Enable Flash-Attention-3

# [Optional] Model Compilation
# Compiling the DiT model accelerates inference, but the first run will take longer to compile.
# pipe.transformer.compile()

# [Optional] CPU Offloading
# Enable CPU offloading for memory-constrained devices.
# pipe.enable_model_cpu_offload()

prompt = "Young Chinese woman in red Hanfu, intricate embroidery. Impeccable makeup, red floral forehead pattern. Elaborate high bun, golden phoenix headdress, red flowers, beads. Holds round folding fan with lady, trees, bird. Neon lightning-bolt lamp (), bright yellow glow, above extended left palm. Soft-lit outdoor night background, silhouetted tiered pagoda (西安大雁塔), blurred colorful distant lights."

# 2. Generate Image
image = pipe(
prompt=prompt,
height=1024,
width=1024,
num_inference_steps=9, # This actually results in 8 DiT forwards
guidance_scale=0.0, # Guidance should be 0 for the Turbo models
generator=torch.Generator("cuda").manual_seed(42),
).images[0]

image.save("example.png")

Decoupled-DMD: The Acceleration Magic Behind Z-Image

Decoupled-DMD is the core few-step distillation algorithm that empowers the 8-step Z-Image model.

Our core insight in Decoupled-DMD is that the success of existing DMD (Distributaion Matching Distillation) methods is the result of two independent, collaborating mechanisms:

  • CFG Augmentation (CA): The primary enginedriving the distillation process, a factor largely overlooked in previous work.
  • Distribution Matching (DM): Acts more as a regularizer, ensuring the stability and quality of the generated output.

By recognizing and decoupling these two mechanisms, we were able to study and optimize them in isolation. This ultimately motivated us to develop an improved distillation process that significantly enhances the performance of few-step generation.

DMDR: Fusing DMD with Reinforcement Learning

Building upon the strong foundation of Decoupled-DMD, our 8-step Z-Image model has already demonstrated exceptional capabilities. To achieve further improvements in terms of semantic alignment, aesthetic quality, and structural coherence—while producing images with richer high-frequency details—we present DMDR.

Our core insight behind DMDR is that Reinforcement Learning (RL) and Distribution Matching Distillation (DMD) can be synergistically integrated during the post-training of few-step models. We demonstrate that:

  • RL Unlocks the Performance of DMD
  • DMD Effectively Regularizes RL

Download

pip install -U huggingface_hub
HF_XET_HIGH_PERFORMANCE=1 hf download Tongyi-MAI/Z-Image-Turbo

Citation

If you find our work useful in your research, please consider citing:

@article{team2025zimage,
  title={Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer},
  author={Z-Image Team},
  journal={arXiv preprint arXiv:2511.22699},
  year={2025}
}

@article{liu2025decoupled,
  title={Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield},
  author={Dongyang Liu and Peng Gao and David Liu and Ruoyi Du and Zhen Li and Qilong Wu and Xin Jin and Sihan Cao and Shifeng Zhang and Hongsheng Li and Steven Hoi},
  journal={arXiv preprint arXiv:2511.22677},
  year={2025}
}

@article{jiang2025distribution,
  title={Distribution Matching Distillation Meets Reinforcement Learning},
  author={Jiang, Dengyang and Liu, Dongyang and Wang, Zanyi and Wu, Qilong and Jin, Xin and Liu, David and Li, Zhen and Wang, Mengmeng and Gao, Peng and Yang, Harry},
  journal={arXiv preprint arXiv:2511.13649},
  year={2025}
}

Identity and Version

Repository
unsloth/Z-Image-Turbo-GGUF
Publisher
Unsloth AI
Task
Text to image
Modality
Image
Library
ggml
Parameters
Not stated by the source
Languages
en
Revision
6c80814333b7b6a70a2e5b469a7c6437ce65de0f
First published
2025-12-21
Last updated
2026-01-08

Files and Weights

32 files, 90.4 GB in total. The weights are 15 files totalling 90.3 GB in gguf.

Weights15 files · 90.3 GB
Documentation1 file · 15.7 KB
Other15 files · 57.7 MB
Repository1 file · 3.4 KB
Every file
FileTypeSizeSHA-256
z-image-turbo-BF16.ggufWeights12.3 GB 9ccc4fb8b845
z-image-turbo-F16.ggufWeights12.3 GB 996ed5b13fc1
z-image-turbo-Q2_K.ggufWeights3.6 GB b0c126881789
z-image-turbo-Q3_K_M.ggufWeights4.2 GB 7070b605165c
z-image-turbo-Q3_K_S.ggufWeights4.0 GB 918f233b1697
z-image-turbo-Q4_0.ggufWeights4.6 GB 302b9cf7e7dd
z-image-turbo-Q4_1.ggufWeights4.9 GB 63c9a92831a0
z-image-turbo-Q4_K_M.ggufWeights5.0 GB e6494f87de6a
z-image-turbo-Q4_K_S.ggufWeights4.7 GB d9a1e25e0751
z-image-turbo-Q5_0.ggufWeights5.3 GB 6a83bb03ea9f
z-image-turbo-Q5_1.ggufWeights5.5 GB e84236eaed78
z-image-turbo-Q5_K_M.ggufWeights5.6 GB 847e33003ae6
z-image-turbo-Q5_K_S.ggufWeights5.2 GB 3d476eadb1a5
z-image-turbo-Q6_K.ggufWeights5.9 GB fc137d87b49e
z-image-turbo-Q8_0.ggufWeights7.2 GB f163d60b0eb4
README.mdDocumentation15.7 KB
assets/DMDR.webpOther173.0 KB 2e6f3053b98d
assets/Z-Image-Gallery.pdfOther15.8 MB 6f9895b3246d
assets/architecture.webpOther422.4 KB 261af62ecc7e
assets/decoupled-dmd.webpOther152.1 KB 4568ca559b99
assets/leaderboard.pngOther2.0 MB e9fd4aa185bb
assets/leaderboard.webpOther63.8 KB
assets/reasoning.pngOther7.7 MB 96c16b2c8d8d
assets/showcase.jpgOther6.4 MB f6ee74e066e0
assets/showcase_editing.pngOther4.7 MB 7d720c3157fd
assets/showcase_realistic.pngOther6.3 MB 697e6f6857f6
assets/showcase_rendering.pngOther7.6 MB 3556dd66be22
assets/sloth_gogh.pngOther2.0 MB db85b8fbb9b2
assets/sloth_mall.pngOther1.5 MB b89f85c9bfae
assets/sloth_sign.pngOther1.2 MB 8564ceb7d329
assets/sloth_wes.pngOther1.6 MB 193b98511009
.gitattributesRepository3.4 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
90.3 GB
Download from Unsloth AI

Released by Unsloth AI through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published90.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Z-Image-Turbo-GGUF

Can I use Z-Image-Turbo-GGUF commercially?

Yes. Z-Image-Turbo-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to image

Juggernaut-XL-v9

RunDiffusion

The SDXL ecosystem is the single most mature corner of open image generation, and v9 is its most refined photorealism checkpoint. Choose Juggernaut XL v9 when you want: - Photorealism that holds up under scrutiny — skin texture, micro-contrast, and natural lighting that translates from concept to print. - Reasonable hardware — runs comfortably on 8 GB of VRAM, unlike newer DiT-based models that demand 16+ GB. - The full SDXL toolbox — drop-in compatibility with the thousands of SDXL ControlNets, IP-Adapter variants, AnimateDiff, regional prompting tools, and LoRAs already in your workflow. - Battle-tested reliability — 26+ months in production, used in agencies, studios, and shipping…

Open weights creativeml-openrail-m diffusers

Model · Text to image

Qwen-Image-Lightning

Lightx2v

Please refer to Qwen-Image-Lightning github to learn how to use the models. make sure to install diffusers from main (pip install git+https://github.com/huggingface/diffusers.git)

Open weights apache-2.0 diffusers

Model · Text to image

Z-Image-Lora

nphSi

+ Always use full LoRa name with "vrtlxxxx" trigger in prompt like "Alba Baptista (vrtlalbabaptista) in a swimming pool". "Woman" or "1girl" will NOT work due to my way i do captions. + Add the gender to the prompt for confusing names like "Alex Jones". + Remove the name when internal model knowledge is bad or censored or is confusing to model like "Sandy Cheeks" or "Kate Middleton". + When using a Lora with multiple triggers (vrtlxx,vrtlyy) do not use the real character name but only trigger or a combination of it. "vrtlMain" always combines all trigger-words. Angourie Rice, January Jones, Julianna Guill, Ursula Corbero, Judith Rakers, Alina Merkau, Kiernan Shipka, Leslie Bibb, Marie…

Open weights apache-2.0 diffusers

Model · Text to image

Flux2-Klein-9B-True-V2

Wikee Yang

Recommend smthemex/ComfyUIUniBlockSwap plugin for LOW VRAM users, it only requires 4-6GB of VRAM to run the full bf16 model. 您可以尝试一下全新的 ComfyUI GGUF 模型加载插件 smthemex/ComfyUIDifGGUF,它能适配更多的 GGUF 文件格式,并且将很快集成低显存(4-8GB)显卡的 GGUF 模型块卸载管理能力。 Recommend to try the new GGUF ComfyUI loader plugin, smthemex/ComfyUIDifGGUF, it compatible with more GGUF format, and will add low VRAM (4-8GB) management for GGUF soon. 1. 图像的真实感和质感进一步改善,基本接近香蕉(Nano Banana)的水平。 2. 提示词遵循和还原能力,参数适配性和LoRA兼容性进一步改善,。 The V2 version of this model has undergone a full fine-tuning based on FLUX.2-Klein-9B-True-V1. Compared to the V1 version, it has some improvements and enhancements as below: 1. The realism and texture of images…

Open weights other diffusers

Model · Text to image

FLUX.1-dev-gguf

City

This is a direct GGUF conversion of black-forest-labs/FLUX.1-dev As this is a quantized model not a finetune, all the same restrictions/original license terms still apply. The model files can be used with the ComfyUI-GGUF custom node. Place model files in ComfyUI/models/unet - see the GitHub readme for further install instructions. Please refer to this chart for a basic overview of quantization types.

Open weights other gguf

Model · Text to image

Pony_Diffusion_V6_XL

Narontaka

Pony Diffusion V6 is a versatile SDXL finetune capable of producing stunning SFW and NSFW visuals of various anthro, feral, or humanoids species and their interactions based on simple natural language prompts. CHECK "ABOUT THIS VERSION" ON THE RIGHT IF YOU ARE NOT ON "V6" FOR IMPORTANT INFORMATION. Please join our Discord Server to support development of new versions of this model and get access to free SD bot and check out more examples of this model capabilities on our prompt sharing website or follow the author on Twitter. Important information Make sure you load this model with clip skip 2 (or -2 in some software), otherwise you will be getting low quality blobs. This model supports a…

Open weights cdla-permissive-2.0 diffusers