SAVRN
Search Contact SAVRN

Open-weight model · Text to image

Z-Image-Turbo

by Tongyi-MAI Tongyi-MAI/Z-Image-Turbo

Welcome to the official repository for the Z-Image(造相)project! Z-Image is a powerful and highly efficient image generation model family with 6B parameters.

Parameters6.2B
Context
Weights32.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads674.1k

Runs On

What it takes to serve Z-Image-Turbo (6.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 12.3 GB 14.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 6.2 GB 7.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 3.1 GB 3.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on Z-Image-Turbo

The download is 32.9 GB across 32 files, but the 16-bit working set is 12.3 GB of weights and 14.8 GB of memory, so plan disk and GPU memory separately for Z-Image-Turbo. Tongyi-MAI distilled it from the 6.2 billion parameter Z-Image to produce a picture in 8 function evaluations, positioned for photorealistic output and English and Chinese text rendering. On the 192 GB MI300X at $1.85 an hour, 14.8 GB leaves most of the card free; 8-bit needs 7.4 GB and 4-bit 3.7 GB.

Apache 2.0 permits commercial image generation, modification and redistribution, with the license and any NOTICE file kept, plus an express patent grant. Two checks: the weights were updated January 30, 2026, two months after the November 25, 2025 release, so pin the revision you tested, and three papers describe the model, arXiv:2511.22699, arXiv:2511.22677 and arXiv:2511.13649, to read before relying on the 8-step claim.

Model Card

By Tongyi-MAI, published under apache-2.0, revision f332072aa78b.

- Image
An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

[![Official Site](https://img.shields.io/badge/Official%20Site-333399.svg?logo=homepage)](https://tongyi-mai.github.io/Z-Image-blog/)  [![GitHub](https://img.shields.io/badge/GitHub-Z--Image-181717?logo=github&logoColor=white)](https://github.com/Tongyi-MAI/Z-Image)  [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Checkpoint-Z--Image--Turbo-yellow)](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo)  [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Online_Demo-Z--Image--Turbo-blue)](https://huggingface.co/spaces/Tongyi-MAI/Z-Image-Turbo)  [![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Mobile_Demo-Z--Image--Turbo-red)](https://huggingface.co/spaces/akhaliq/Z-Image-Turbo)  [![ModelScope Model](https://img.shields.io/badge/%20Checkpoint-Z--Image--Turbo-624aff)](https://www.modelscope.cn/models/Tongyi-MAI/Z-Image-Turbo)  [![ModelScope Space](https://img.shields.io/badge/%20Online_Demo-Z--Image--Turbo-17c7a7)](https://www.modelscope.cn/aigc/imageGeneration?tab=advanced&versionId=469191&modelType=Checkpoint&sdVersion=Z_IMAGE_TURBO&modelUrl=modelscope%3A%2F%2FTongyi-MAI%2FZ-Image-Turbo%3Frevision%3Dmaster)  [![Art Gallery PDF](https://img.shields.io/badge/%F0%9F%96%BC%20Art_Gallery-PDF-ff69b4)](assets/Z-Image-Gallery.pdf)  [![Web Art Gallery](https://img.shields.io/badge/%F0%9F%8C%90%20Web_Art_Gallery-online-00bfff)](https://modelscope.cn/studios/Tongyi-MAI/Z-Image-Gallery/summary)  Welcome to the official repository for the Z-Image(造相)project!

Z-Image

Z-Image is a powerful and highly efficient image generation model family with 6B parameters. Currently there are four variants:

Read the full model card (1,041 words)

Identity and Version

Repository
Tongyi-MAI/Z-Image-Turbo
Publisher
Tongyi-MAI
Task
Text to image
Modality
Image
Library
diffusers
Parameters
6.2B parameters
Languages
en
Revision
f332072aa78be7aecdf3ee76d5c247082da564a6
First published
2025-11-25
Last updated
2026-01-30

Files and Weights

32 files, 32.9 GB in total. The weights are 7 files totalling 32.8 GB in safetensors.

Weights7 files · 32.8 GB
Configuration8 files · 84.7 KB
Tokenizer4 files · 15.9 MB
Documentation1 file · 13.7 KB
Other11 files · 51.3 MB
Repository1 file · 2.2 KB
Every file
FileTypeSizeSHA-256
text_encoder/model-00001-of-00003.safetensorsWeights4.0 GB 328a91d31223
text_encoder/model-00002-of-00003.safetensorsWeights4.0 GB 6cd087b31630
text_encoder/model-00003-of-00003.safetensorsWeights99.6 MB 7ca841ee75b9
transformer/diffusion_pytorch_model-00001-of-00003.safetensorsWeights10.0 GB 95facd593e25
transformer/diffusion_pytorch_model-00002-of-00003.safetensorsWeights10.0 GB a4bbe43ee184
transformer/diffusion_pytorch_model-00003-of-00003.safetensorsWeights4.7 GB aba4e37a590e
vae/diffusion_pytorch_model.safetensorsWeights167.7 MB f5b59a268515
model_index.jsonConfiguration467 B
scheduler/scheduler_config.jsonConfiguration173 B
text_encoder/config.jsonConfiguration726 B
text_encoder/generation_config.jsonConfiguration239 B
text_encoder/model.safetensors.index.jsonConfiguration32.8 KB
transformer/config.jsonConfiguration473 B
transformer/diffusion_pytorch_model.safetensors.index.jsonConfiguration49.0 KB
vae/config.jsonConfiguration805 B
README.mdDocumentation13.7 KB
assets/DMDR.webpOther173.0 KB 2e6f3053b98d
assets/Z-Image-Gallery.pdfOther15.8 MB 6f9895b3246d
assets/architecture.webpOther422.4 KB 261af62ecc7e
assets/decoupled-dmd.webpOther152.1 KB 4568ca559b99
assets/leaderboard.pngOther2.0 MB e9fd4aa185bb
assets/leaderboard.webpOther63.8 KB
assets/reasoning.pngOther7.7 MB 96c16b2c8d8d
assets/showcase.jpgOther6.4 MB f6ee74e066e0
assets/showcase_editing.pngOther4.7 MB 7d720c3157fd
assets/showcase_realistic.pngOther6.3 MB 697e6f6857f6
assets/showcase_rendering.pngOther7.6 MB 3556dd66be22
.gitattributesRepository2.2 KB
tokenizer/merges.txtTokenizer1.7 MB
tokenizer/tokenizer.jsonTokenizer11.4 MB aeb13307a71a
tokenizer/tokenizer_config.jsonTokenizer9.7 KB
tokenizer/vocab.jsonTokenizer2.8 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
32.8 GB
Download from Tongyi-MAI

Released by Tongyi-MAI through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2511.13649
  • Described by arXiv:2511.22677
  • Described by arXiv:2511.22699

Memory Requirements

PrecisionWeights in memory
As published32.8 GB
16-bit12.3 GB
8-bit6.2 GB
4-bit3.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Z-Image-Turbo

How much GPU memory does Z-Image-Turbo need?

About 14.8 GB at 16-bit and 3.7 GB at 4-bit: the weights (6.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Z-Image-Turbo on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Z-Image-Turbo commercially?

Yes. Z-Image-Turbo is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to image

stable-diffusion-3.5-large

Stability AI

Stable Diffusion 3.5 Large is a Multimodal Diffusion Transformer (MMDiT) text-to-image model that features improved performance in image quality, typography, complex prompt understanding, and resource-efficiency. Please note: This model is released under the Stability Community License. Visit Stability AI to learn or contact us for commercial licensing details. - For individuals and organizations with annual revenue above $1M: please contact us to get an Enterprise License. For local or self-hosted use, we recommend ComfyUI for node-based UI inference, or diffusers or GitHub for programmatic use. - Text Encoders: This model was trained on a wide variety of data, including synthetic data and…

Access requested at publisher other 8.1B parameters diffusers

basemodel: stabilityai/stable-diffusion-xl-base-1.0 - stable-diffusion-xl - stable-diffusion-xl-diffusers - text-to-image - diffusers - inpainting SD-XL Inpainting 0.1 is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input, with the extra capability of inpainting the pictures by using a mask. The SD-XL Inpainting 0.1 was initialized with the stable-diffusion-xl-base-1.0 weights. The model is trained for 40k steps at resolution 1024x1024 and 5% dropping of the text-conditioning to improve classifier-free classifier-free guidance sampling. For inpainting, the UNet has 5 additional input channels (4 for the encoded masked-image and 1 for the…

Open weights openrail++ 2.6B parameters diffusers

SDXL consists of an ensemble of experts pipeline for latent diffusion: In a first step, the base model is used to generate (noisy) latents, which are then further processed with a refinement model (available here: https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0/) specialized for the final denoising steps. Note that the base model can be used as a standalone module. Alternatively, we can use a two-stage pipeline as follows: First, the base model is used to generate latents of the desired output size. In the second step, we use a specialized high-resolution model and apply a technique called SDEdit (https://arxiv.org/abs/2108.01073, also known as "img2img") to the latents…

Open weights openrail++ 2.6B parameters diffusers

Model · Text to image

sdxl-turbo

Stability AI

SDXL-Turbo is a fast generative text-to-image model that can synthesize photorealistic images from a text prompt in a single network evaluation. A real-time demo is available here: http://clipdrop.co/stable-diffusion-turbo Please note: For commercial use, please refer to https://stability.ai/license. SDXL-Turbo is a distilled version of SDXL 1.0, trained for real-time synthesis. SDXL-Turbo is based on a novel training method called Adversarial Diffusion Distillation (ADD) (see the technical report), which allows sampling large-scale foundational image diffusion models in 1 to 4 steps at high image quality. This approach uses score distillation to leverage large-scale off-the-shelf image…

Open weights other 2.6B parameters diffusers

Model · Text to image

animagine-xl-4.0

Cagliostro Labs

Animagine XL 4.0, also stylized as Anim4gine, is the ultimate anime-themed finetuned SDXL model and the latest installment of Animagine XL series. Despite being a continuation, the model was retrained from Stable Diffusion XL 1.0 with a massive dataset of 8.4M diverse anime-style images from various sources with the knowledge cut-off of January 7th 2025 and finetuned for approximately 2650 GPU hours. Similar to the previous version, this model was trained using tag ordering method for the identity and style training. With the release of Animagine XL 4.0 Opt (Optimized), the model has been further refined with an additional dataset, improving stability, anatomy accuracy, noise reduction…

Open weights openrail++ 2.6B parameters diffusers

This repository contains a model that generates highly aesthetic images of resolution 1024x1024, as well as portrait and landscape aspect ratios. You can use the model with Hugging Face Diffusers. Playground v2.5 is a diffusion-based text-to-image generative model, and a successor to Playground v2. Playground v2.5 is the state-of-the-art open-source model in aesthetic quality. Our user studies demonstrate that our model outperforms SDXL, Playground v2, PixArt-α, DALL-E 3, and Midjourney 5.2. For details on the development and training of our model, please refer to our blog post and technical report. Install diffusers >= 0.27.0 and the relevant dependencies. - The pipeline uses the…

Open weights other 2.6B parameters diffusers