SAVRN
Search Contact SAVRN

Open-weight model · Text to image

controlnet-union-sdxl-1.0

by Qi xinsir/controlnet-union-sdxl-1.0

Use bucket training like novelai, can generate high resolutions images of any aspect ratio - Use large amount of high quality data(over 10000000 images), the dataset covers a diversity of situation - Use re-captioned prompt like DALLE.3, use CogVLM to…

Parameters1.3B
Context
Weights5.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads118.7k

Runs On

What it takes to serve controlnet-union-sdxl-1.0 (1.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 2.5 GB 3.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.3 GB 1.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.6 GB 0.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Qi, published under apache-2.0, revision 801a4a3fa3d4.

Use bucket training like novelai, can generate high resolutions images of any aspect ratio - Use large amount of high quality data(over 10000000 images), the dataset covers a diversity of situation - Use re-captioned prompt like DALLE.3, use CogVLM to generate detailed description, good prompt following ability - Use many useful tricks during training. Including but not limited to date augmentation, mutiple loss, multi resolution - Use almost the same parameter compared with original ControlNet. No obvious increase in network parameter or computation. - Support 10+ control conditions, no obvious performance drop on any single condition compared with training independently - Support multi…

Read Qi's full model card

ControlNet++: All-in-one ControlNet for image generations and editing!

ProMax Model has released!! 12 control + 5 advanced editing, just try it!!!

Network Arichitecture

Advantages about the model

  • Use bucket training like novelai, can generate high resolutions images of any aspect ratio
  • Use large amount of high quality data(over 10000000 images), the dataset covers a diversity of situation
  • Use re-captioned prompt like DALLE.3, use CogVLM to generate detailed description, good prompt following ability
  • Use many useful tricks during training. Including but not limited to date augmentation, mutiple loss, multi resolution
  • Use almost the same parameter compared with original ControlNet. No obvious increase in network parameter or computation.
  • Support 10+ control conditions, no obvious performance drop on any single condition compared with training independently
  • Support multi condition generation, condition fusion is learned during training. No need to set hyperparameter or design prompts.
  • Compatible with other opensource SDXL models, such as BluePencilXL, CounterfeitXL. Compatible with other Lora models.

We design a new architecture that can support 10+ control types in condition text-to-image generation and can generate high resolution images visually comparable with midjourney. The network is based on the original ControlNet architecture, we propose two new modules to: 1 Extend the original ControlNet to support different image conditions using the same network parameter. 2 Support multiple conditions input without increasing computation offload, which is especially important for designers who want to edit image in detail, different conditions use the same condition encoder, without adding extra computations or parameters. We do thoroughly experiments on SDXL and achieve superior performance both in control ability and aesthetic score. We release the method and the model to the open source community to make everyone can enjoy it.

Inference scripts and more details can found: https://github.com/xinsir6/ControlNetPlus/tree/main

If you find it useful, please give me a star, thank you very much

SDXL ProMax version has been released!!!,Enjoy it!!!

I am sorry that because of the project's revenue and expenditure are difficult to balance, the GPU resources are assigned to other projects that are more likely to be profitable, the SD3 trainging is stopped until I find enough GPU supprt, I will try my best to find GPUs to continue training. If this brings you inconvenience, I sincerely apologize for that. I want to thank everyone who likes this project, your support is what keeps me going

Note: we put the promax model with a promax suffix in the same huggingface model repo, detailed instructions will be added later.

Advanced editing features in Promax Model

Tile Deblur

Tile variation

Tile Super Resolution

Following example show from 1M resolution --> 9M resolution

Image Inpainting

Image Outpainting

Visual Examples

Openpose

Depth

Canny

Lineart

AnimeLineart

Mlsd

Scribble

Hed

Pidi(Softedge)

Teed

Segment

Normal

Multi Control Visual Examples

Openpose + Canny

Openpose + Depth

Openpose + Scribble

Openpose + Normal

Openpose + Segment

Identity and Version

Repository
xinsir/controlnet-union-sdxl-1.0
Publisher
Qi
Task
Text to image
Modality
Image
Library
diffusers
Parameters
1.3B parameters
Languages
Not stated by the source
Revision
801a4a3fa3d4c936f4feea95b98607bc6726f80c
First published
2024-07-07
Last updated
2024-07-30

Files and Weights

126 files, 5.1 GB in total. The weights are 2 files totalling 5.0 GB in safetensors.

Weights2 files · 5.0 GB
Configuration2 files · 2.5 KB
Documentation1 file · 9.9 KB
Other120 files · 44.2 MB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
diffusion_pytorch_model.safetensorsWeights2.5 GB a9e13fd61f31
diffusion_pytorch_model_promax.safetensorsWeights2.5 GB 9fae2e50cb43
config.jsonConfiguration1.2 KB
config_promax.jsonConfiguration1.3 KB
README.mdDocumentation9.9 KB
images/000000_pose_concat.webpOther247.7 KB
images/000001_openpose_scribble_concat.webpOther316.9 KB
images/000001_pose_concat.webpOther292.4 KB
images/000002_openpose_scribble_concat.webpOther216.1 KB
images/000002_pose_concat.webpOther212.6 KB
images/000003_openpose_scribble_concat.webpOther397.8 KB
images/000003_pose_concat.webpOther387.8 KB
images/000004_openpose_scribble_concat.webpOther211.6 KB
images/000004_pose_concat.webpOther735.3 KB
images/000005_depth_concat.webpOther184.7 KB
images/000005_openpose_scribble_concat.webpOther243.9 KB
images/000006_depth_concat.webpOther215.0 KB
images/000006_openpose_scribble_concat.webpOther197.7 KB
images/000007_depth_concat.webpOther171.9 KB
images/000007_openpose_canny_concat.webpOther420.7 KB
images/000008_depth_concat.webpOther738.5 KB
images/000008_openpose_canny_concat.webpOther436.1 KB
images/000009_depth_concat.webpOther367.1 KB
images/000009_openpose_canny_concat.webpOther172.5 KB
images/000010_canny_concat.webpOther301.7 KB
images/000010_openpose_canny_concat.webpOther292.3 KB
images/000011_canny_concat.webpOther274.3 KB
images/000011_openpose_canny_concat.webpOther649.4 KB
images/000012_canny_concat.webpOther555.8 KB
images/000012_openpose_canny_concat.webpOther673.0 KB
images/000013_canny_concat.webpOther194.5 KB
images/000013_openpose_depth_concat.webpOther147.0 KB
images/000014_canny_concat.webpOther205.4 KB
images/000014_openpose_depth_concat.webpOther162.8 KB
images/000015_lineart_concat.webpOther298.4 KB
images/000015_openpose_depth_concat.webpOther115.9 KB
images/000016_lineart_concat.webpOther306.3 KB
images/000016_openpose_depth_concat.webpOther230.2 KB
images/000017_lineart_concat.webpOther323.5 KB
images/000017_openpose_depth_concat.webpOther252.6 KB
images/000018_lineart_concat.webpOther159.6 KB
images/000018_openpose_depth_concat.webpOther140.7 KB
images/000019_lineart_concat.webpOther124.1 KB
images/000019_openpose_normal_concat.webpOther309.9 KB
images/000020_anime_lineart_concat.webpOther415.7 KB
images/000020_openpose_normal_concat.webpOther341.2 KB
images/000021_anime_lineart_concat.webpOther331.9 KB
images/000021_openpose_normal_concat.webpOther134.9 KB
images/000022_anime_lineart_concat.webpOther509.5 KB
images/000022_openpose_normal_concat.webpOther146.9 KB
images/000023_anime_lineart_concat.webpOther247.1 KB
images/000023_openpose_normal_concat.webpOther294.8 KB
images/000024_anime_lineart_concat.webpOther405.8 KB
images/000024_openpose_normal_concat.webpOther93.3 KB
images/000025_mlsd_concat.webpOther599.4 KB
images/000025_openpose_sam_concat.webpOther327.1 KB
images/000026_mlsd_concat.webpOther547.5 KB
images/000026_openpose_sam_concat.webpOther346.4 KB
images/000027_mlsd_concat.webpOther1.1 MB ec2719db68c6
images/000027_openpose_sam_concat.webpOther322.1 KB
images/000028_mlsd_concat.webpOther558.8 KB
images/000028_openpose_sam_concat.webpOther190.2 KB
images/000029_mlsd_concat.webpOther630.0 KB
images/000029_openpose_sam_concat.webpOther358.2 KB
images/000030_openpose_sam_concat.webpOther240.7 KB
images/000030_scribble_concat.webpOther350.6 KB
images/000031_scribble_concat.webpOther82.2 KB
images/000032_scribble_concat.webpOther485.4 KB
images/000033_scribble_concat.webpOther544.2 KB
images/000034_scribble_concat.webpOther346.6 KB
images/000035_hed_concat.webpOther431.6 KB
images/000036_hed_concat.webpOther173.7 KB
images/000037_hed_concat.webpOther380.9 KB
images/000038_hed_concat.webpOther703.3 KB
images/000039_hed_concat.webpOther300.4 KB
images/000040_softedge_concat.webpOther154.7 KB
images/000041_softedge_concat.webpOther248.4 KB
images/000042_softedge_concat.webpOther235.4 KB
images/000043_softedge_concat.webpOther461.6 KB
images/000044_softedge_concat.webpOther250.7 KB
images/000045_ted_concat.webpOther196.0 KB
images/000046_ted_concat.webpOther223.9 KB
images/000047_ted_concat.webpOther201.6 KB
images/000048_ted_concat.webpOther225.8 KB
images/000049_ted_concat.webpOther679.4 KB
images/000050_seg_concat.webpOther175.5 KB
images/000051_seg_concat.webpOther279.2 KB
images/000052_seg_concat.webpOther600.3 KB
images/000053_seg_concat.webpOther587.2 KB
images/000054_seg_concat.webpOther209.7 KB
images/000055_normal_concat.webpOther246.7 KB
images/000056_normal_concat.webpOther332.0 KB
images/000057_normal_concat.webpOther198.8 KB
images/000058_normal_concat.webpOther628.6 KB
images/000059_normal_concat.webpOther155.7 KB
images/100000_tile_blur_concat.webpOther296.8 KB
images/100001_tile_blur_concat.webpOther395.8 KB
images/100002_tile_blur_concat.webpOther311.9 KB
images/100003_tile_blur_concat.webpOther524.2 KB
images/100004_tile_blur_concat.webpOther178.6 KB
images/100005_tile_blur_concat.webpOther469.3 KB
images/100006_tile_var_concat.webpOther234.0 KB
images/100007_tile_var_concat.webpOther314.7 KB
images/100008_tile_var_concat.webpOther600.6 KB
images/100009_tile_var_concat.webpOther486.5 KB
images/100010_tile_var_concat.webpOther931.6 KB
images/100011_tile_var_concat.webpOther204.1 KB
images/100012_outpainting_concat.webpOther629.2 KB
images/100013_outpainting_concat.webpOther230.2 KB
images/100014_outpainting_concat.webpOther275.6 KB
images/100015_outpainting_concat.webpOther461.3 KB
images/100016_outpainting_concat.webpOther201.2 KB
images/100017_outpainting_concat.webpOther136.0 KB
images/100018_inpainting_concat.webpOther120.4 KB
images/100019_inpainting_concat.webpOther230.9 KB
images/100020_inpainting_concat.webpOther519.6 KB
images/100021_inpainting_concat.webpOther187.3 KB
images/100022_inpainting_concat.webpOther137.9 KB
images/100023_inpainting_concat.webpOther305.2 KB
images/ControlNet++.pngOther224.3 KB
images/masonry.webpOther4.1 MB a2da28443ec2
images/tile_super1.webpOther114.7 KB
images/tile_super1_9upscale.webpOther897.8 KB
images/tile_super2.webpOther47.8 KB
images/tile_super2_9upscale.webpOther368.0 KB
.gitattributesRepository1.6 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
5.0 GB
Download from Qi

Released by Qi through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published5.0 GB
16-bit2.5 GB
8-bit1.3 GB
4-bit0.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About controlnet-union-sdxl-1.0

How much GPU memory does controlnet-union-sdxl-1.0 need?

About 3 GB at 16-bit and 0.8 GB at 4-bit: the weights (1.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run controlnet-union-sdxl-1.0 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use controlnet-union-sdxl-1.0 commercially?

Yes. controlnet-union-sdxl-1.0 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to image

controlnet-openpose-sdxl-1.0

Qi

thanks feiyuuu for report the problem. When using the default pose line the performance may be unstable, this is because the pose label use more thick line in training to have a better look. This difference can be fix by using the following method: Find the util.py in controlnetaux python package, usually the path is like: /your anaconda3 path/envs/your env name/lib/python3.8/site-packages/controlnetaux/openpose/util.py Replace the drawbodypose function with the following code: Use the code below to get started with the model. HumanArt [https://github.com/IDEA-Research/HumanArt], select 2000 images with ground truth pose annotations to generate images and calculate mAP. We are the SOTA…

Open weights apache-2.0 1.3B parameters diffusers

Model · Text to image

sd-turbo

Stability AI

SD-Turbo is a fast generative text-to-image model that can synthesize photorealistic images from a text prompt in a single network evaluation. We release SD-Turbo as a research artifact, and to study small, distilled text-to-image models. For increased quality and prompt understanding, we recommend SDXL-Turbo. Please note: For commercial use, please refer to https://stability.ai/license. SD-Turbo is a distilled version of Stable Diffusion 2.1, trained for real-time synthesis. SD-Turbo is based on a novel training method called Adversarial Diffusion Distillation (ADD) (see the technical report), which allows sampling large-scale foundational image diffusion models in 1 to 4 steps at high…

Open weights 866M parameters diffusers

Model · Text to image

LCM_Dreamshaper_v7

Simian Luo

Distilled from Dreamshaper v7 fine-tune of Stable-Diffusion v1-5 with only 4,000 training iterations (~32 A100 GPU Hours). By distilling classifier-free guidance into the model's input, LCM can generate high-quality images in very short inference time. We compare the inference time at the setting of 768 x 768 resolution, CFG scale w=8, batchsize=4, using a A800 GPU. You can try out Latency Consistency Models directly on: To run the model yourself, you can leverage the Diffusers library: 1. Install the library: 2. Run the model: For more information, please have a look at the official docs: https://huggingface.co/docs/diffusers/api/pipelines/latentconsistencymodels#latent-consistency-models…

Open weights mit 860M parameters diffusers

Model · Text to image

stable-diffusion-v1-5

SD v1.5

Modifications to the original model card are in red or green Stable Diffusion is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input. For more information about how Stable Diffusion functions, please have a look at 's Stable Diffusion blog. The Stable-Diffusion-v1-5 checkpoint was initialized with the weights of the Stable-Diffusion-v1-2 checkpoint and subsequently fine-tuned on 595k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling. You can use this both with the Diffusers library and RunwayML GitHub repository ( now deprecated ), ComfyUI…

Open weights creativeml-openrail-m 860M parameters diffusers

Model · Text to image

dreamshaper-7

Lykon

lykon/dreamshaper-7 is a Stable Diffusion model that has been fine-tuned on runwayml/stable-diffusion-v1-5. For more general information on how to run text-to-image models with Diffusers, see the docs. - Version 8 focuses on improving what V7 started. Might be harder to do photorealism compared to realism focused models, as it might be hard to do anime compared to anime focused models, but it can do both pretty well if you're skilled enough. Check the examples! - Version 7 improves lora support, NSFW and realism. If you're interested in "absolute" realism, try AbsoluteReality. - Version 6 adds more lora support and more style in general. It should also be better at generating directly at…

Open weights creativeml-openrail-m 860M parameters diffusers

Model · Text to image

stable-diffusion-v1-4

CompVis

Stable Diffusion is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input. For more information about how Stable Diffusion functions, please have a look at 's Stable Diffusion with Diffusers blog. The Stable-Diffusion-v1-4 checkpoint was initialized with the weights of the Stable-Diffusion-v1-2 checkpoint and subsequently fine-tuned on 225k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling. This weights here are intended to be used with the Diffusers library. If you are looking for the weights to be loaded into the CompVis Stable Diffusion codebase…

Open weights creativeml-openrail-m 860M parameters diffusers