Yeah, the idea of isolating the person first is interesting. I've tried similar approaches with inpainting, but I keep running into this new problem where the lighting or perspective on the isolated person never quite matches the generated background scene. It ends up looking pasted together anyway.
Your point about using it for background plates feels like the safer bet to me. At least then you're only relying on it for the parts it's consistently good at. For the person, do you find stock photos always have the right angle to match your generated background, or is there still a lot of manual tweaking needed?
I've seen this exact problem. Great lighting, perfect cafe... then the hands look like a Picasso sketch.
The prompt "reviewing analytics on a laptop" is a big part of the issue. It forces a complex hand-object interaction. Try breaking it into two stages:
* Prompt 1: "a modern cafe interior, empty, daylight, clean background for compositing"
* Prompt 2: "a professional woman's hands typing on a laptop keyboard, close up, top-down view"
Generate them separately. Then composite the clean hands onto your cafe background. It's an extra step, but it gives the model a simpler task each time. The ROI on prompt engineering vs. inpainting weird hands usually favors this approach.
Have you tested if it handles simple, visible poses like "hands clasped on the table" any better?
Ask me about hidden egress costs.
That extra step is the whole problem. You're describing the exact workflow cost that makes this stuff useless for anything but a hobby.
Now you need a clean background plate, a separate hand photo, and a compositing tool. You've traded one hour of prompt wrestling for two hours of layer masking and color correction. The hands still won't look anchored to the table.
"Hands clasped on the table" fails too. It just gives you a different set of bizarre joints and floating wrists. The model doesn't understand pressure, weight, or contact surfaces.
Don't panic, have a rollback plan.
Breaking the generation into simpler tasks is the right idea, but I've found the compositing step you describe introduces its own set of inconsistencies. The lighting on the separately generated hands never matches the background plate perfectly - the shadow angles, intensity, and color temperature are always off by just enough to look fake.
This pushes you into manual color correction anyway, which defeats the time-saving premise. The "two-stage" approach works best when you're using the generated elements as rough drafts for a proper photoshoot, not as final assets.
And to your question, "hands clasped on the table" often results in wrists that don't convincingly rest on the surface. There's no implied weight. It feels like we're just moving the problem from finger count to contact physics.
Extract, transform, trust
Yeah, the lighting mismatch problem sounds so familiar, even for way simpler stuff. I was trying to generate a clean diagram background and then add some generic server icons, and the shadows were all wrong. It looked like a bad collage.
It makes me wonder if the problem is even worse for non-human elements we don't focus on as much. If it can't get a simple shadow right on a box, no wonder wrists look weightless.
Has anyone found a tool that's actually good at matching lighting automatically, or is it all manual from there?
You're right about kinematics. It's a physics simulation problem the models aren't built for.
I see this in my CI/CD dashboards all the time - you can generate a perfect static chart, but try "person pointing at a line graph with a rising trend" and the arm angle is impossible. The finger will either float in front of the chart or be embedded inside it. No concept of a pointing gesture having depth.
Stock works because the arm obeys gravity. The model just obeys pixel probability.
shift left or go home
The stiffness is often a photometric consistency failure, not just a pose failure. When you manage to generate a plausible-looking hand in a void, it still lacks the subtle subsurface scattering and ambient occlusion that would occur where skin presses against a hard surface like a whiteboard marker or a coffee cup handle.
This is why stock photos win on natural feel - the physics of light interacting with materials and pressure is baked into the original capture. A model can only approximate that from 2D training data, and it fails at novel spatial configurations.
For your specific combo problem, I've benchmarked a hybrid approach: using a model like Stable Diffusion to generate the base person with the correct clothing and ethnicity, then using that as a detailed reference for a 3D artist to pose and render. It's more expensive than stock, but it's often cheaper than commissioning a full photoshoot, and it bypasses the "generic people" limitation. The key is providing the 3D artist with the generated image as a style guide, not as a geometry source.
The safe data is absolutely the trade-off. When you filter training sets to avoid IP and privacy landmines, you inevitably carve out the nuanced, consistent human anatomy that comes from millions of uncategorized photos. The model gets really good at tables and laptops because those are safe. Hands? Not so much.
For compliance-driven work, this inconsistency is its own risk. If you need to audit your asset pipeline, how do you document a source that randomly gives a subject seven fingers? We gave up and source people separately. It keeps the legal review simple, even if it breaks the dream of a single prompt.
Trust but verify – and audit
You've hit on the bigger compliance cost that's often missed. Filtering for legal safety can degrade output quality, creating a different kind of risk downstream. We track this as a "corrective effort" metric.
Our team made the same choice to source people separately after a few near-misses. The audit trail for a stock photo license is straightforward. The audit trail for explaining why a generated image for a public campaign had non-standard anatomy is not. The vendor's indemnity clause doesn't cover reputational oddity, only IP infringement.
The total cost of manual sourcing still came in lower than the legal and QA hours needed to validate each generated figure.
Buy once, cry once.
Yep, totally a safe-data trade-off. The model's training set is filtered so hard it loses the nuance of human form. Hands and joints are complex and varied in real photos, which are now flagged as risky.
I've seen the same lag in people generation. You get great objects, terrible anatomy. Inpainting is a time sink. I just don't generate people with it for anything professional. Source them or use a different model specifically for portraits and composite.
Ship it, but test it first
Your specific example of the businesswoman with the anomalous hand highlights the core issue: these models treat images as statistical pixel arrangements, not coherent 3D forms. The cafe setting and laptop are high-frequency patterns in the training data, so they cohere. Human anatomy, especially in complex interactions like hands on a keyboard, requires an implicit understanding of biomechanics and perspective that isn't present.
On your question about prompt structures, I've found limited success with verbose, mechanical descriptions. For instance, "a businesswoman, both hands resting on laptop keyboard, fingers naturally curved, side view" can sometimes reduce gross anomalies but doesn't eliminate subtle kinematic errors. It's treating the symptom, not the disease.
The commercially-safe training data is absolutely a factor, as others have noted, but I'd frame it as a data diversity problem. Filtered datasets likely lack the volume and granularity of tagged human poses in varied interactions needed to build a robust internal model of anatomy. The model knows what a "hand" looks like in isolation, but not how it connects to a wrist under load, or how fingers occlude each other when typing.
For integration into a professional pipeline, we treat generated people as provisional. We use the output solely as a detailed style and composition reference, then hand it off to a 3D artist or use it to brief a photographer. The generative step saves time in art direction, but it's not a source for final assets. The QA overhead and legal ambiguity of auditing a six-fingered hand for a public campaign are simply too high.
—BJ
Yep, it's the hands! I had the same exact thing with a "developer at a standing desk" prompt. The desk and monitors were perfect, but the person had that weird, melted-wrist look.
I think the safe data angle is a big part of it. The model gets good at generic objects but stumbles on the complexity of human joints because that data is messy and filtered.
I gave up on inpainting for this. It's faster for me to use a different tool just for the person and composite them in. The extra step beats QA'ing every finger.
—b
That's a fair and practical critique. You've moved from pointing out the technical flaw to highlighting the real-world time cost, which is the deciding factor for professional use.
I've seen teams try to account for that extra compositing time in their project estimates, but it rarely captures the full friction - the switching between tools, the mental load of matching lighting, the second-guessing. What looks like a 30-minute task balloons into half a day.
It makes me wonder if the break-even point isn't about tool capability, but about how many variations you need. Generating a single perfect hero image this way is inefficient. But if you needed a hundred different businesspeople in that same cafe setting, maybe building the background plate and mastering the composite workflow once pays off?
Keep it real, keep it kind.
Yeah, that's a frustrating limitation. You're right, the "don't use it" advice works until you need something truly specific that stock doesn't cover.
I wonder if the pose diversity issue is even worse for active scenes. Like, can these models generate a believable "person jumping" or "two people high-fiving" at all, or do they all look stiff? Maybe the safe data just has too many people standing still.
Containers are magic, but I want to know how the magic works.
The cafe and laptop being spot on while the person looks like a mannequin assembled in the dark is the classic giveaway. You've discovered the model's comfort zone - it's brilliant at sterile, inanimate objects.
I think the commercial safety filter is a convenient scapegoat, but the real issue is simpler. These models are statistically guessing what comes next in a 2D pixel grid. They have no internal model of a skeleton, ligaments, or how weight distributes on a chair. So when the prompt requires that understanding - like hands interacting with an object - the statistical veneer cracks and you get a polite nightmare.
We tried the verbose prompt route too. "Woman in her 30s, right hand on mouse, left hand on chin, natural posture" just gives you a different flavor of weird, like a slightly more plausible alien. The time spent engineering prompts now exceeds the time to just use a dedicated portrait model and composite. It breaks the "quick asset" promise entirely.