SDXL

Explore smarter pose control for pixel art characters with AI-driven tools

Controlling character pose in SDXL with a reference image is a delicate balancing act, and this user's struggle with duplicated limbs is a familiar one.

4 min readMachine Learning

The question posed by this user is a familiar one for anyone who has tried to bend Stable Diffusion to their will, especially at the demanding resolution of 128x128 pixel art. They are doing everything right on paper: preprocessing the reference, using IP-Adapter for identity, and ControlNet for pose. Yet the model still betrays them by duplicating limbs or ignoring the rig. This is not a failure of effort, but a fundamental misunderstanding of how these conditioning signals interact. ControlNet and IP-Adapter are not two dials that work independently; they are competing forces in a high-dimensional space, and the model is often caught in the crossfire.

The core issue is that IP-Adapter is incredibly aggressive at encoding appearance, and that appearance often includes the spatial layout of the original image. When your reference character has an arm at their side, the adapter sees that as part of the identity. When ControlNet then asks for an arm raised, the model faces a contradiction: it has a strong signal for "this person, as seen" and a strong signal for "this new pose." The result is a compromise that looks like a horror show. Our take is that you need to stop thinking about this as a conflict to be resolved with strength percentages and start thinking about it as a data problem. The reason your multiple reference images (front, rear, left, right) are not solving this is because you are still giving the model a 2D photograph of a 3D object and asking it to infer the underlying skeleton.

This is where the broader context of Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges becomes relevant. In production systems, you rarely rely on a single model to do everything. You build a pipeline. For your case, the most practical solution is to decouple the pose from the image entirely. Instead of using a reference image for the pose, generate a clean, stick-figure skeleton or a 3D mesh render that has no texture, no shading, and no character identity. Feed that into ControlNet. Then, use IP-Adapter only on a cropped, tight shot of the character's face or a specific detailed part of the costume, not the full body. By removing the full-body spatial context from the IP-Adapter input, you force it to focus on style and texture rather than geometry. This is a hack, but it is a hack that works because it respects the inductive biases of the models.

We would also suggest looking at the Forrester Function article not for the math, but for the lesson in optimization: when you have multiple variables interacting non-linearly, grid-searching a few values is not enough. You need to change the objective. In practice, this means you should not be using a single ControlNet pass. Instead, generate the pose map at a much higher resolution, say 512x512, and then downscale it to 128x128 just before the ControlNet injection. This gives the model cleaner edges to work with. Also, consider inpainting the face after the fact. Generate the pose, then use the reference image to inpaint the facial details in a second pass. This is more expensive, but it is far more reliable than trying to get a single generation to do everything perfectly.

The takeaway here is direct: your current approach is not broken, it is just imprecise. The model is not stubborn; it is just following conflicting instructions. Stop asking it to do everything at once. Separate the geometry from the texture, give each one a clean, unambiguous input, and you will find that the duplicated limbs disappear. The open question is whether you can invest the time to build this two-stage pipeline, or whether you will continue to fight the model's nature. We would bet that the moment you stop trying to control the pose through the reference image and instead control it through a pure skeleton, your consistency will jump. That is the specific detail to watch for in your next experiment.

From Machine Learning

Hi, I’m working on generating ~128×128 pixel art and trying to generate different poses of the same character.

Start with a reference image and preprocess it into cleaner/more pixel-art-like data (often removing transparency or setting up fixed number of pallets or descaling) Use IP-Adapter for the character/reference appearance. Use ControlNet pose/rig conditioning to control the target pose. I’m also experimenting with multiple references (front, rear, left, right), with pose/rig and depth annotations. For the target pose, I provide a separate pose reference through the conditioning pipeline.

Read the original at Machine Learning