thanks feiyuuu for report the problem. When using the default pose line the performance may be unstable, this is because the pose label use more thick line in training to have a better look. This difference can be fix by using the following method: Find the util.py in controlnetaux python package, usually the path is like: /your anaconda3 path/envs/your env name/lib/python3.8/site-packages/controlnetaux/openpose/util.py Replace the drawbodypose function with the following code: Use the code below to get started with the model. HumanArt [https://github.com/IDEA-Research/HumanArt], select 2000 images with ground truth pose annotations to generate images and calculate mAP. We are the SOTA…
Use bucket training like novelai, can generate high resolutions images of any aspect ratio - Use large amount of high quality data(over 10000000 images), the dataset covers a diversity of situation - Use re-captioned prompt like DALLE.3, use CogVLM to…
Runs On
What it takes to serve controlnet-union-sdxl-1.0 (1.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 2.5 GB | 3.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 1.3 GB | 1.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.6 GB | 0.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
By Qi, published under apache-2.0, revision 801a4a3fa3d4.
Use bucket training like novelai, can generate high resolutions images of any aspect ratio - Use large amount of high quality data(over 10000000 images), the dataset covers a diversity of situation - Use re-captioned prompt like DALLE.3, use CogVLM to generate detailed description, good prompt following ability - Use many useful tricks during training. Including but not limited to date augmentation, mutiple loss, multi resolution - Use almost the same parameter compared with original ControlNet. No obvious increase in network parameter or computation. - Support 10+ control conditions, no obvious performance drop on any single condition compared with training independently - Support multi…
Read Qi's full model card
ControlNet++: All-in-one ControlNet for image generations and editing!
ProMax Model has released!! 12 control + 5 advanced editing, just try it!!!
Network Arichitecture
Advantages about the model
- Use bucket training like novelai, can generate high resolutions images of any aspect ratio
- Use large amount of high quality data(over 10000000 images), the dataset covers a diversity of situation
- Use re-captioned prompt like DALLE.3, use CogVLM to generate detailed description, good prompt following ability
- Use many useful tricks during training. Including but not limited to date augmentation, mutiple loss, multi resolution
- Use almost the same parameter compared with original ControlNet. No obvious increase in network parameter or computation.
- Support 10+ control conditions, no obvious performance drop on any single condition compared with training independently
- Support multi condition generation, condition fusion is learned during training. No need to set hyperparameter or design prompts.
- Compatible with other opensource SDXL models, such as BluePencilXL, CounterfeitXL. Compatible with other Lora models.
We design a new architecture that can support 10+ control types in condition text-to-image generation and can generate high resolution images visually comparable with midjourney. The network is based on the original ControlNet architecture, we propose two new modules to: 1 Extend the original ControlNet to support different image conditions using the same network parameter. 2 Support multiple conditions input without increasing computation offload, which is especially important for designers who want to edit image in detail, different conditions use the same condition encoder, without adding extra computations or parameters. We do thoroughly experiments on SDXL and achieve superior performance both in control ability and aesthetic score. We release the method and the model to the open source community to make everyone can enjoy it.
Inference scripts and more details can found: https://github.com/xinsir6/ControlNetPlus/tree/main
If you find it useful, please give me a star, thank you very much
SDXL ProMax version has been released!!!,Enjoy it!!!
I am sorry that because of the project's revenue and expenditure are difficult to balance, the GPU resources are assigned to other projects that are more likely to be profitable, the SD3 trainging is stopped until I find enough GPU supprt, I will try my best to find GPUs to continue training. If this brings you inconvenience, I sincerely apologize for that. I want to thank everyone who likes this project, your support is what keeps me going
Note: we put the promax model with a promax suffix in the same huggingface model repo, detailed instructions will be added later.
Advanced editing features in Promax Model
Tile Deblur
Tile variation
Tile Super Resolution
Following example show from 1M resolution --> 9M resolution
Image Inpainting
Image Outpainting
Visual Examples
Openpose
Depth
Canny
Lineart
AnimeLineart
Mlsd
Scribble
Hed
Pidi(Softedge)
Teed
Segment
Normal
Multi Control Visual Examples
Openpose + Canny
Openpose + Depth
Openpose + Scribble
Openpose + Normal
Openpose + Segment
Identity and Version
- Repository
- xinsir/controlnet-union-sdxl-1.0
- Publisher
- Qi
- Task
- Text to image
- Modality
- Image
- Library
- diffusers
- Parameters
- 1.3B parameters
- Languages
- Not stated by the source
- Revision
- 801a4a3fa3d4c936f4feea95b98607bc6726f80c
- First published
- 2024-07-07
- Last updated
- 2024-07-30
Files and Weights
126 files, 5.1 GB in total. The weights are 2 files totalling 5.0 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| diffusion_pytorch_model.safetensors | Weights | 2.5 GB | a9e13fd61f31 |
| diffusion_pytorch_model_promax.safetensors | Weights | 2.5 GB | 9fae2e50cb43 |
| config.json | Configuration | 1.2 KB | — |
| config_promax.json | Configuration | 1.3 KB | — |
| README.md | Documentation | 9.9 KB | — |
| images/000000_pose_concat.webp | Other | 247.7 KB | — |
| images/000001_openpose_scribble_concat.webp | Other | 316.9 KB | — |
| images/000001_pose_concat.webp | Other | 292.4 KB | — |
| images/000002_openpose_scribble_concat.webp | Other | 216.1 KB | — |
| images/000002_pose_concat.webp | Other | 212.6 KB | — |
| images/000003_openpose_scribble_concat.webp | Other | 397.8 KB | — |
| images/000003_pose_concat.webp | Other | 387.8 KB | — |
| images/000004_openpose_scribble_concat.webp | Other | 211.6 KB | — |
| images/000004_pose_concat.webp | Other | 735.3 KB | — |
| images/000005_depth_concat.webp | Other | 184.7 KB | — |
| images/000005_openpose_scribble_concat.webp | Other | 243.9 KB | — |
| images/000006_depth_concat.webp | Other | 215.0 KB | — |
| images/000006_openpose_scribble_concat.webp | Other | 197.7 KB | — |
| images/000007_depth_concat.webp | Other | 171.9 KB | — |
| images/000007_openpose_canny_concat.webp | Other | 420.7 KB | — |
| images/000008_depth_concat.webp | Other | 738.5 KB | — |
| images/000008_openpose_canny_concat.webp | Other | 436.1 KB | — |
| images/000009_depth_concat.webp | Other | 367.1 KB | — |
| images/000009_openpose_canny_concat.webp | Other | 172.5 KB | — |
| images/000010_canny_concat.webp | Other | 301.7 KB | — |
| images/000010_openpose_canny_concat.webp | Other | 292.3 KB | — |
| images/000011_canny_concat.webp | Other | 274.3 KB | — |
| images/000011_openpose_canny_concat.webp | Other | 649.4 KB | — |
| images/000012_canny_concat.webp | Other | 555.8 KB | — |
| images/000012_openpose_canny_concat.webp | Other | 673.0 KB | — |
| images/000013_canny_concat.webp | Other | 194.5 KB | — |
| images/000013_openpose_depth_concat.webp | Other | 147.0 KB | — |
| images/000014_canny_concat.webp | Other | 205.4 KB | — |
| images/000014_openpose_depth_concat.webp | Other | 162.8 KB | — |
| images/000015_lineart_concat.webp | Other | 298.4 KB | — |
| images/000015_openpose_depth_concat.webp | Other | 115.9 KB | — |
| images/000016_lineart_concat.webp | Other | 306.3 KB | — |
| images/000016_openpose_depth_concat.webp | Other | 230.2 KB | — |
| images/000017_lineart_concat.webp | Other | 323.5 KB | — |
| images/000017_openpose_depth_concat.webp | Other | 252.6 KB | — |
| images/000018_lineart_concat.webp | Other | 159.6 KB | — |
| images/000018_openpose_depth_concat.webp | Other | 140.7 KB | — |
| images/000019_lineart_concat.webp | Other | 124.1 KB | — |
| images/000019_openpose_normal_concat.webp | Other | 309.9 KB | — |
| images/000020_anime_lineart_concat.webp | Other | 415.7 KB | — |
| images/000020_openpose_normal_concat.webp | Other | 341.2 KB | — |
| images/000021_anime_lineart_concat.webp | Other | 331.9 KB | — |
| images/000021_openpose_normal_concat.webp | Other | 134.9 KB | — |
| images/000022_anime_lineart_concat.webp | Other | 509.5 KB | — |
| images/000022_openpose_normal_concat.webp | Other | 146.9 KB | — |
| images/000023_anime_lineart_concat.webp | Other | 247.1 KB | — |
| images/000023_openpose_normal_concat.webp | Other | 294.8 KB | — |
| images/000024_anime_lineart_concat.webp | Other | 405.8 KB | — |
| images/000024_openpose_normal_concat.webp | Other | 93.3 KB | — |
| images/000025_mlsd_concat.webp | Other | 599.4 KB | — |
| images/000025_openpose_sam_concat.webp | Other | 327.1 KB | — |
| images/000026_mlsd_concat.webp | Other | 547.5 KB | — |
| images/000026_openpose_sam_concat.webp | Other | 346.4 KB | — |
| images/000027_mlsd_concat.webp | Other | 1.1 MB | ec2719db68c6 |
| images/000027_openpose_sam_concat.webp | Other | 322.1 KB | — |
| images/000028_mlsd_concat.webp | Other | 558.8 KB | — |
| images/000028_openpose_sam_concat.webp | Other | 190.2 KB | — |
| images/000029_mlsd_concat.webp | Other | 630.0 KB | — |
| images/000029_openpose_sam_concat.webp | Other | 358.2 KB | — |
| images/000030_openpose_sam_concat.webp | Other | 240.7 KB | — |
| images/000030_scribble_concat.webp | Other | 350.6 KB | — |
| images/000031_scribble_concat.webp | Other | 82.2 KB | — |
| images/000032_scribble_concat.webp | Other | 485.4 KB | — |
| images/000033_scribble_concat.webp | Other | 544.2 KB | — |
| images/000034_scribble_concat.webp | Other | 346.6 KB | — |
| images/000035_hed_concat.webp | Other | 431.6 KB | — |
| images/000036_hed_concat.webp | Other | 173.7 KB | — |
| images/000037_hed_concat.webp | Other | 380.9 KB | — |
| images/000038_hed_concat.webp | Other | 703.3 KB | — |
| images/000039_hed_concat.webp | Other | 300.4 KB | — |
| images/000040_softedge_concat.webp | Other | 154.7 KB | — |
| images/000041_softedge_concat.webp | Other | 248.4 KB | — |
| images/000042_softedge_concat.webp | Other | 235.4 KB | — |
| images/000043_softedge_concat.webp | Other | 461.6 KB | — |
| images/000044_softedge_concat.webp | Other | 250.7 KB | — |
| images/000045_ted_concat.webp | Other | 196.0 KB | — |
| images/000046_ted_concat.webp | Other | 223.9 KB | — |
| images/000047_ted_concat.webp | Other | 201.6 KB | — |
| images/000048_ted_concat.webp | Other | 225.8 KB | — |
| images/000049_ted_concat.webp | Other | 679.4 KB | — |
| images/000050_seg_concat.webp | Other | 175.5 KB | — |
| images/000051_seg_concat.webp | Other | 279.2 KB | — |
| images/000052_seg_concat.webp | Other | 600.3 KB | — |
| images/000053_seg_concat.webp | Other | 587.2 KB | — |
| images/000054_seg_concat.webp | Other | 209.7 KB | — |
| images/000055_normal_concat.webp | Other | 246.7 KB | — |
| images/000056_normal_concat.webp | Other | 332.0 KB | — |
| images/000057_normal_concat.webp | Other | 198.8 KB | — |
| images/000058_normal_concat.webp | Other | 628.6 KB | — |
| images/000059_normal_concat.webp | Other | 155.7 KB | — |
| images/100000_tile_blur_concat.webp | Other | 296.8 KB | — |
| images/100001_tile_blur_concat.webp | Other | 395.8 KB | — |
| images/100002_tile_blur_concat.webp | Other | 311.9 KB | — |
| images/100003_tile_blur_concat.webp | Other | 524.2 KB | — |
| images/100004_tile_blur_concat.webp | Other | 178.6 KB | — |
| images/100005_tile_blur_concat.webp | Other | 469.3 KB | — |
| images/100006_tile_var_concat.webp | Other | 234.0 KB | — |
| images/100007_tile_var_concat.webp | Other | 314.7 KB | — |
| images/100008_tile_var_concat.webp | Other | 600.6 KB | — |
| images/100009_tile_var_concat.webp | Other | 486.5 KB | — |
| images/100010_tile_var_concat.webp | Other | 931.6 KB | — |
| images/100011_tile_var_concat.webp | Other | 204.1 KB | — |
| images/100012_outpainting_concat.webp | Other | 629.2 KB | — |
| images/100013_outpainting_concat.webp | Other | 230.2 KB | — |
| images/100014_outpainting_concat.webp | Other | 275.6 KB | — |
| images/100015_outpainting_concat.webp | Other | 461.3 KB | — |
| images/100016_outpainting_concat.webp | Other | 201.2 KB | — |
| images/100017_outpainting_concat.webp | Other | 136.0 KB | — |
| images/100018_inpainting_concat.webp | Other | 120.4 KB | — |
| images/100019_inpainting_concat.webp | Other | 230.9 KB | — |
| images/100020_inpainting_concat.webp | Other | 519.6 KB | — |
| images/100021_inpainting_concat.webp | Other | 187.3 KB | — |
| images/100022_inpainting_concat.webp | Other | 137.9 KB | — |
| images/100023_inpainting_concat.webp | Other | 305.2 KB | — |
| images/ControlNet++.png | Other | 224.3 KB | — |
| images/masonry.webp | Other | 4.1 MB | a2da28443ec2 |
| images/tile_super1.webp | Other | 114.7 KB | — |
| images/tile_super1_9upscale.webp | Other | 897.8 KB | — |
| images/tile_super2.webp | Other | 47.8 KB | — |
| images/tile_super2_9upscale.webp | Other | 368.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 5.0 GB
Released by Qi through its official repository on Hugging Face. Read the license.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 5.0 GB |
| 16-bit | 2.5 GB |
| 8-bit | 1.3 GB |
| 4-bit | 0.6 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About controlnet-union-sdxl-1.0
How much GPU memory does controlnet-union-sdxl-1.0 need?
About 3 GB at 16-bit and 0.8 GB at 4-bit: the weights (1.3B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run controlnet-union-sdxl-1.0 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use controlnet-union-sdxl-1.0 commercially?
Yes. controlnet-union-sdxl-1.0 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Similar Models
SD-Turbo is a fast generative text-to-image model that can synthesize photorealistic images from a text prompt in a single network evaluation. We release SD-Turbo as a research artifact, and to study small, distilled text-to-image models. For increased quality and prompt understanding, we recommend SDXL-Turbo. Please note: For commercial use, please refer to https://stability.ai/license. SD-Turbo is a distilled version of Stable Diffusion 2.1, trained for real-time synthesis. SD-Turbo is based on a novel training method called Adversarial Diffusion Distillation (ADD) (see the technical report), which allows sampling large-scale foundational image diffusion models in 1 to 4 steps at high…
Distilled from Dreamshaper v7 fine-tune of Stable-Diffusion v1-5 with only 4,000 training iterations (~32 A100 GPU Hours). By distilling classifier-free guidance into the model's input, LCM can generate high-quality images in very short inference time. We compare the inference time at the setting of 768 x 768 resolution, CFG scale w=8, batchsize=4, using a A800 GPU. You can try out Latency Consistency Models directly on: To run the model yourself, you can leverage the Diffusers library: 1. Install the library: 2. Run the model: For more information, please have a look at the official docs: https://huggingface.co/docs/diffusers/api/pipelines/latentconsistencymodels#latent-consistency-models…
Modifications to the original model card are in red or green Stable Diffusion is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input. For more information about how Stable Diffusion functions, please have a look at 's Stable Diffusion blog. The Stable-Diffusion-v1-5 checkpoint was initialized with the weights of the Stable-Diffusion-v1-2 checkpoint and subsequently fine-tuned on 595k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling. You can use this both with the Diffusers library and RunwayML GitHub repository ( now deprecated ), ComfyUI…
lykon/dreamshaper-7 is a Stable Diffusion model that has been fine-tuned on runwayml/stable-diffusion-v1-5. For more general information on how to run text-to-image models with Diffusers, see the docs. - Version 8 focuses on improving what V7 started. Might be harder to do photorealism compared to realism focused models, as it might be hard to do anime compared to anime focused models, but it can do both pretty well if you're skilled enough. Check the examples! - Version 7 improves lora support, NSFW and realism. If you're interested in "absolute" realism, try AbsoluteReality. - Version 6 adds more lora support and more style in general. It should also be better at generating directly at…
Stable Diffusion is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input. For more information about how Stable Diffusion functions, please have a look at 's Stable Diffusion with Diffusers blog. The Stable-Diffusion-v1-4 checkpoint was initialized with the weights of the Stable-Diffusion-v1-2 checkpoint and subsequently fine-tuned on 225k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling. This weights here are intended to be used with the Diffusers library. If you are looking for the weights to be loaded into the CompVis Stable Diffusion codebase…