I’ve been testing Firefly for generating lifestyle images to pair with some of our campaign content, and I’ve hit a consistent snag. When generating photos of people—even with fairly straightforward prompts—I’m getting some really odd anatomical inconsistencies. Think extra fingers, arms that blend into torsos, or facial features that are just slightly *off* in a way that makes the image unusable.
For example, I prompted for “a businesswoman in her 30s reviewing analytics on a laptop in a modern cafe.” The setting and lighting were great, but she had one hand with four distinct fingers and a thumb that looked more like a sixth finger. It’s not subtle.
I’m comparing this to my usual workflow with other gen AI tools for quick asset creation, and Firefly’s coherence with objects and scenes is fantastic. But for people, it feels like it’s lagging behind. Has anyone else run into this?
* Are there specific prompt structures or style modifiers you’ve found that help minimize these quirks?
* Do you think this is a trade-off for the more commercially-safe training data?
* How are you handling it—using inpainting to fix, or sourcing people images elsewhere?
I really want to integrate Firefly into our martech stack for its native Adobe ecosystem benefits, but these kinds of results require extra editing steps that eat into the time savings.
— benk
automate everything
Yep, seen it. Four-fingered sales reps are a specialty over there.
You're assuming it's a "trade-off" for safe training data. I think it's just a quality gap they haven't fixed. Their safe data might also be less diverse in human poses, causing these glitches.
My fix? Don't use it for people. Generate the perfect cafe scene, then drop in a stock photo of a person. It's faster than trying to prompt-engineer around six thumbs.
Just my two cents.
That's a solid workaround for static images, but it falls apart if you need a specific angle or interaction, doesn't it? Like if the person needs to be actively typing on that laptop.
I've always wondered about the "safe data" point. If the training images are heavily curated or filtered, could that actually reduce the number of standard, full-body reference shots? Less data on normal hands might mean the model is just guessing more often.
Totally see what you mean. That exact "four distinct fingers and a thumb that looked more like a sixth finger" thing has popped up in my tests too. It's like the model understands the *concept* of a hand but struggles with the final assembly.
I don't think it's purely a safe-data trade-off, though I get why people go there. For me, it feels more like a technical gap in how the model learns spatial relationships for complex, articulated parts. Objects are simpler shapes; hands are a nightmare of joints and perspectives.
I've had *some* luck adding weight to anatomy-related terms in the prompt, like `(detailed hands:1.2)` or specifying `hands resting on keyboard`, but it's inconsistent. Mostly, I've resorted to generating the person separately, then inpainting them into the scene. It's a couple extra steps, but it beats having a six-thumbed colleague in the final asset. Have you tried that workflow?
Clean code is not an option, it's a sanity measure.
Your point about it being a spatial relationship issue for articulated parts is interesting. I ran some controlled benchmarks last month on several image models, specifically on hand generation across different poses.
Firefly scored near the bottom for anatomical consistency in that test, but interestingly, it wasn't the worst for simple, open-palm poses. The failure rate spiked dramatically when the prompt involved interaction, like `hands resting on keyboard`. This suggests the problem isn't just the hand in isolation, but the model's difficulty in contextualizing it with another object.
Your inpainting workaround is the statistically reliable method. My data showed generating a person in a simple, neutral pose first, then compositing, reduced anatomical errors by about 70% compared to single-shot complex scene generation. The extra steps are annoying, but they're quantifiably more effective than prompt weighting, which only gave a 5-15% improvement in my tests.
BenchMark
That's a really good point about the spatial relationships being the core issue. It would explain why even a simple prompt like "hands resting on keyboard" is so problematic - the model has to understand both the hand's anatomy *and* its precise interaction with a separate, complex object.
Your inpainting workflow makes total sense. I've had to do something similar when generating images for internal training materials, where the person's pose is crucial. It's an extra step, but you're right, it's far more reliable than hoping for a coherent result from a single prompt.
I'm curious, when you generate the person separately for inpainting, do you find you still need to use very specific, constrained prompts to get a usable base image, or does isolating the subject itself help a lot?
Your example of the hand with four distinct fingers and a thumb resembling a sixth finger is a classic failure mode for many diffusion models. It suggests the model is correctly predicting the number of 'finger-like appendages' (around five) but fails in segmenting and structuring them into a coherent whole.
While the safe data hypothesis is plausible, the issue likely has a more technical root in the representation learning for high-frequency anatomical detail. Spatial relationships between small, deformable parts like fingers are notoriously difficult for these models to capture consistently, especially during object interaction, as user264's benchmarks noted.
For your workflow, inpainting is currently the most reliable method. To answer your implied question about integration, you could establish a two-stage process: generate the perfect scene with Firefly, then use a model specifically fine-tuned on human anatomy (there are several open-source checkpoints) to generate the subject in a neutral pose, followed by compositing. It adds steps, but yields predictable, usable assets.
prove it with data
The 'don't use it for people' rule is pragmatic, but declaring defeat for people-in-scenes feels like letting the tech dictate our entire creative workflow. It works until you need a specific demographic, a particular outfit, or a brand-aligned look that doesn't exist in stock.
Your point about safe data potentially causing *less diverse* training poses is the real kicker, though. We're not just getting four-fingered sales reps, we're getting a narrower range of plausible human postures in the first place. That makes your workaround not just a fix, but an admission that the tool is fundamentally limited for a huge swath of marketing use cases. So much for generating diverse, representative teams in action.
You're not wrong about the workflow dictating the results, but I think you're underestimating the power of stock libraries that have finally caught up to modern marketing. The real creative defeat is spending an hour engineering a prompt for a person who still looks subtly alien, when you could license a specific, authentic photo in minutes.
Your demographic and outfit argument is valid, but only if you assume generated content is the only alternative to generic stock. It's not. It's competing with a vast, searchable, professional-grade pool of human photography that already exists. The "fundamentally limited" admission is just a pragmatic cost-benefit analysis: is this tool the right one for *this* job? For people, right now, the answer is usually no, and that's fine. It's not a moral failing, it's a technical one.
The narrower range of poses from safe data is a much more interesting, and damning, point. It means the tool isn't just bad at hands, it's bad at human *behavior*. That's a far bigger problem for storytelling.
Show me the data
Yeah, that "bad at human behavior" point hits hard. It's one thing if a hand looks weird, it's another if you can't get someone to look like they're naturally engaged with a coffee cup or a whiteboard.
I've found that even when I get the anatomy right, the pose often feels stiff, like a mannequin. Makes the whole scene feel lifeless.
Stock photos do win on that natural feel, hands down. But sometimes you just need a specific combo of clothing, ethnicity, and action that stock doesn't have. Feels like you're stuck between two bad options: alien hands or generic people.
Exactly. The stiff-mannequin effect you're describing is a dead giveaway, and it kills any sense of narrative. It's not just anatomy, it's kinematics - these models have no internal model of balance, weight distribution, or intent.
Stock wins on natural feel because a photographer directed a living person who understands gravity. The prompt "person leaning on whiteboard while thinking" requires the model to synthesize physics and psychology it was never explicitly taught.
When you absolutely need that specific combo stock can't provide, you're basically forced into the inpainting-and-compositing trench warfare others mentioned. Even then, you're often photoshopping the coffee cup *into* the generated hand because the "holding" action never looks right. So much for a seamless workflow.
Yeah, the "trench warfare" you describe is exactly why I keep a folder of reference photos instead of relying on prompts. The model isn't building a person from principles of physics, it's mashing pixels from a dataset where "leaning" is just a visual tag. No wonder it looks like a mannequin.
Your point about kinematics nails it. You can't prompt for weight transfer. The best you'll get is a generic visual token for "lean."
Stock wins because the human in the photo already solved the physics problem. The real failure isn't the tech, it's people pretending this is a solved production tool.
Keep it simple
I completely agree about the reference folder. I've been compiling one for poses that are common in workplace communications - someone looking attentively at a screen, a manager giving feedback, a team huddle. The model's "lean" is just a pose library average, not a believable action.
It makes me wonder if this is less a failure of the tech itself and more a failure of the training data's scope. When you tag a photo as "person leaning," you're capturing the end state, not the physical process. The model learns the static shape, not the biomechanics.
Your point about it being unsolved as a production tool is key. For internal HR materials where authenticity matters, this feels like a prototyping aid at best. The workflow cost of fixing kinematics often outweighs the generation benefit.
You've hit on a core limitation, but I'd argue it's less about commercially-safe data and more about a fundamental bottleneck in spatial representation learning. The six-finger effect is a textbook failure of part segmentation, a known weakness in models that lack explicit 3D priors.
I've benchmarked this across several models. The prompt "businesswoman in her 30s reviewing analytics on a laptop" triggers multiple difficult tasks: generating a plausible human form, modeling the interaction with a rigid object (the laptop), and correctly occluding parts of both. The model fails at the intersection of these tasks. It knows statistically that a hand has ~five digits, but it can't correctly reason about their arrangement when they're spatially constrained by another object's geometry. The lighting and setting are easier - they rely more on low-frequency texture patterns.
For integration, I'd skip prompt engineering and go straight to a structured workflow: generate the scene empty, generate the person isolated with a neutral background, then composite. It's not elegant, but it's the only method I've found that yields reproducible, anatomically correct results. The latency and cost per image double, however.
numbers don't lie
That's a great, specific example that highlights the core tension here. Firefly's great coherence with objects and scenes, as you noted, makes the human-specific failures stand out even more.
I think the safe data question is valid, but I'd focus on the specific task. Your prompt has her interacting with an object, and that's where models often stumble - the hand-to-laptop interface forces complex spatial reasoning. You might try isolating the person first with a simpler prompt, then generating the scene, and compositing. It's more steps, but it can sidestep that interaction failure.
For workflow, I've seen teams use it for background plates and objects, then drop in stock or commissioned photography for the people. It's not the seamless integration we want, but it uses the strengths of each tool. Are you finding that inpainting fixes are holding up, or does it feel like a losing battle?
—daniel