SAVRN
Search Contact SAVRN

Open-weight model · Text to image

controlnet-openpose-sdxl-1.0

by Qi xinsir/controlnet-openpose-sdxl-1.0

thanks feiyuuu for report the problem. When using the default pose line the performance may be unstable, this is because the pose label use more thick line in training to have a better look.

Parameters1.3B
Context
Weights5.0 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads109.6k

Runs On

What it takes to serve controlnet-openpose-sdxl-1.0 (1.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 2.5 GB 3.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.3 GB 1.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.6 GB 0.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By Qi, published under apache-2.0, revision 23f966cd5cfd.

State of the art ControlNet-openpose-sdxl-1.0 model, below are the result for midjourney and anime, just for show

controlnet-openpose-sdxl-1.0

  • Developed by: xinsir
  • Model type: ControlNet_SDXL
  • License: apache-2.0
  • Finetuned from model [optional]: stabilityai/stable-diffusion-xl-base-1.0

Model Sources [optional]

- Paper [optional]: https://arxiv.org/abs/2302.05543

Examples

Replace the default draw pose function to get better result

thanks feiyuuu for report the problem. When using the default pose line the performance may be unstable, this is because the pose label use more thick line in training to have a better look. This difference can be fix by using the following method:

Find the util.py in controlnet_aux python package, usually the path is like: /your anaconda3 path/envs/your env name/lib/python3.8/site-packages/controlnet_aux/open_pose/util.py Replace the draw_bodypose function with the following code:

Read the full model card (767 words)

Identity and Version

Repository
xinsir/controlnet-openpose-sdxl-1.0
Publisher
Qi
Task
Text to image
Modality
Image
Library
diffusers
Parameters
1.3B parameters
Languages
Not stated by the source
Revision
23f966cd5cfdd3f7729c903e243d87152162d2b7
First published
2024-05-13
Last updated
2024-07-09

Files and Weights

27 files, 5.0 GB in total. The weights are 2 files totalling 5.0 GB in safetensors.

Weights2 files · 5.0 GB
Configuration1 file · 1.2 KB
Documentation1 file · 7.5 KB
Other22 files · 8.7 MB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
diffusion_pytorch_model.safetensorsWeights2.5 GB b8524e557a7d
diffusion_pytorch_model_twins.safetensorsWeights2.5 GB 54a2afb1bd21
config.jsonConfiguration1.2 KB
README.mdDocumentation7.5 KB
000001_scribble_concat.webpOther324.8 KB
000003_scribble_concat.webpOther290.2 KB
000005_scribble_concat.webpOther358.6 KB
000008_scribble_concat.webpOther278.1 KB
000010_scribble_concat.webpOther152.7 KB
000015_scribble_concat.webpOther196.4 KB
000024_scribble_concat.webpOther149.9 KB
000028_scribble_concat.webpOther354.7 KB
000030_scribble_concat.webpOther207.7 KB
000031_scribble_concat.webpOther235.1 KB
000042_scribble_concat.webpOther183.8 KB
000044_scribble_concat.webpOther112.6 KB
000047_scribble_concat.webpOther287.3 KB
000048_scribble_concat.webpOther273.0 KB
000083_scribble_concat.webpOther269.4 KB
000101_scribble_concat.webpOther213.3 KB
000127_scribble_concat.webpOther230.5 KB
000128_scribble_concat.webpOther232.6 KB
000155_scribble_concat.webpOther183.1 KB
000180_scribble_concat.webpOther122.7 KB
masonry0.webpOther2.0 MB 699840380289
masonry_real.webpOther2.0 MB e1467a55d77d
.gitattributesRepository1.6 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
5.0 GB
Download from Qi

Released by Qi through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2302.05543

Memory Requirements

PrecisionWeights in memory
As published5.0 GB
16-bit2.5 GB
8-bit1.3 GB
4-bit0.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About controlnet-openpose-sdxl-1.0

How much GPU memory does controlnet-openpose-sdxl-1.0 need?

About 3 GB at 16-bit and 0.8 GB at 4-bit: the weights (1.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run controlnet-openpose-sdxl-1.0 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use controlnet-openpose-sdxl-1.0 commercially?

Yes. controlnet-openpose-sdxl-1.0 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text to image

controlnet-union-sdxl-1.0

Qi

Use bucket training like novelai, can generate high resolutions images of any aspect ratio - Use large amount of high quality data(over 10000000 images), the dataset covers a diversity of situation - Use re-captioned prompt like DALLE.3, use CogVLM to generate detailed description, good prompt following ability - Use many useful tricks during training. Including but not limited to date augmentation, mutiple loss, multi resolution - Use almost the same parameter compared with original ControlNet. No obvious increase in network parameter or computation. - Support 10+ control conditions, no obvious performance drop on any single condition compared with training independently - Support multi…

Open weights apache-2.0 1.3B parameters diffusers

Model · Text to image

sd-turbo

Stability AI

SD-Turbo is a fast generative text-to-image model that can synthesize photorealistic images from a text prompt in a single network evaluation. We release SD-Turbo as a research artifact, and to study small, distilled text-to-image models. For increased quality and prompt understanding, we recommend SDXL-Turbo. Please note: For commercial use, please refer to https://stability.ai/license. SD-Turbo is a distilled version of Stable Diffusion 2.1, trained for real-time synthesis. SD-Turbo is based on a novel training method called Adversarial Diffusion Distillation (ADD) (see the technical report), which allows sampling large-scale foundational image diffusion models in 1 to 4 steps at high…

Open weights 866M parameters diffusers

Model · Text to image

LCM_Dreamshaper_v7

Simian Luo

Distilled from Dreamshaper v7 fine-tune of Stable-Diffusion v1-5 with only 4,000 training iterations (~32 A100 GPU Hours). By distilling classifier-free guidance into the model's input, LCM can generate high-quality images in very short inference time. We compare the inference time at the setting of 768 x 768 resolution, CFG scale w=8, batchsize=4, using a A800 GPU. You can try out Latency Consistency Models directly on: To run the model yourself, you can leverage the Diffusers library: 1. Install the library: 2. Run the model: For more information, please have a look at the official docs: https://huggingface.co/docs/diffusers/api/pipelines/latentconsistencymodels#latent-consistency-models…

Open weights mit 860M parameters diffusers

Model · Text to image

stable-diffusion-v1-5

SD v1.5

Modifications to the original model card are in red or green Stable Diffusion is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input. For more information about how Stable Diffusion functions, please have a look at 's Stable Diffusion blog. The Stable-Diffusion-v1-5 checkpoint was initialized with the weights of the Stable-Diffusion-v1-2 checkpoint and subsequently fine-tuned on 595k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling. You can use this both with the Diffusers library and RunwayML GitHub repository ( now deprecated ), ComfyUI…

Open weights creativeml-openrail-m 860M parameters diffusers

Model · Text to image

dreamshaper-7

Lykon

lykon/dreamshaper-7 is a Stable Diffusion model that has been fine-tuned on runwayml/stable-diffusion-v1-5. For more general information on how to run text-to-image models with Diffusers, see the docs. - Version 8 focuses on improving what V7 started. Might be harder to do photorealism compared to realism focused models, as it might be hard to do anime compared to anime focused models, but it can do both pretty well if you're skilled enough. Check the examples! - Version 7 improves lora support, NSFW and realism. If you're interested in "absolute" realism, try AbsoluteReality. - Version 6 adds more lora support and more style in general. It should also be better at generating directly at…

Open weights creativeml-openrail-m 860M parameters diffusers

Model · Text to image

stable-diffusion-v1-4

CompVis

Stable Diffusion is a latent text-to-image diffusion model capable of generating photo-realistic images given any text input. For more information about how Stable Diffusion functions, please have a look at 's Stable Diffusion with Diffusers blog. The Stable-Diffusion-v1-4 checkpoint was initialized with the weights of the Stable-Diffusion-v1-2 checkpoint and subsequently fine-tuned on 225k steps at resolution 512x512 on "laion-aesthetics v2 5+" and 10% dropping of the text-conditioning to improve classifier-free guidance sampling. This weights here are intended to be used with the Diffusers library. If you are looking for the weights to be loaded into the CompVis Stable Diffusion codebase…

Open weights creativeml-openrail-m 860M parameters diffusers