I've been rigorously testing Adobe Firefly for generating marketing assets (banner ads, social posts) and have hit a consistent, performance-impacting issue: a high rate of anatomical errors, specifically extra limbs. This is more than an occasional glitch; in my last batch of 50 generations for a "team collaboration" concept, 12 outputs featured figures with three arms or mis-joined legs, rendering them unusable.
From an analytics perspective, this introduces significant friction into the workflow. The "success rate" per prompt plummets, requiring multiple iterations and manual review, which defeats the efficiency promise of generative AI. I'm treating this as an A/B testing problem: the prompt is the variable, the output is the result.
**My current hypothesis is that the issue stems from:**
* **Prompt Ambiguity:** Terms like "group of people," "crowd," or "dynamic team" might confuse the spatial reasoning model, causing limb overlap and generation errors.
* **Overly Complex Poses:** Requests for "celebrating," "high-fiving," or "active" scenes increase the probability of malformed joints.
* **Lack of Negative Prompts:** Firefly's current web interface doesn't offer explicit negative prompting (e.g., "extra limbs," "malformed anatomy").
**My attempted fixes with mixed results:**
* Simplifying prompts to "single person, standing, facing camera" reduces errors drastically but limits creative scope.
* Using more explicit positional language: "two people standing side-by-side, visible shoulders to hips."
* Iterative regeneration is costly in terms of time and credits.
Has anyone else quantified this error rate in their workflow? More importantly, have you developed a systematic approach to prompt construction that minimizes these artifacts while maintaining complex scene composition? I'm particularly interested in any parallels to multivariate testing logic—isolating the key prompt variables that lead to clean outputs.
-- J
Data never lies, but it can be misleading
You've nailed the core prompt engineering issue, but there's a data migration layer here you're brushing against. The training data corpus these models are built on is a messy, unnormalized legacy system of its own. Scenes tagged as "group" or "team" in the source images likely contain overlapping limbs in perspective, and the model learns that as a valid, if statistically noisy, pattern.
Your A/B testing approach is solid. Try prompts that enforce separation through schema, like "three distinct people, spaced apart, facing the viewer" instead of "team." It's like writing a strict SQL join instead of a fuzzy lookup.
The lack of negative prompting is the real killer, though. You can't exclude "extra arm" or "malformed limb" in the base interface, which is like trying to do an ETL cleanup without a WHERE clause.
Expect the unexpected
The SQL join analogy is clever, but it's giving the model too much credit. It's not running a query on clean data, it's performing stochastic pattern matching on a pile of mislabeled JPEGs.
Your suggested prompt is a good workaround, but it's just masking the symptom. The real problem is the black box itself. You're stuck tweaking inputs because you can't fix the underlying process, which is the opposite of a proper, controllable pipeline.
null
You're right that it's a black box, but honestly, that's where we all have to live right now. Tweaking inputs *is* the job when you're on a budget and the tool is "good enough" to save you five figures on stock photos and a designer.
It's like trying to get a clear signal out of an old radio. You don't fix the wiring, you just keep adjusting the dial until the static fades. For my startup, that prompt-tweaking time is still cheaper than the alternative. Does it suck sometimes? Absolutely. But I can't rebuild the radio.
Build with what you have
That "good enough" calculation is exactly how I justify the seat license cost. But the time cost of prompt tweaking isn't static - it scales with volume. For a handful of images, it's fine. For a campaign needing 200 assets, the QA overhead becomes a significant operational drag.
I think the real cost is in the risk of missing a defect. A three-armed figure in a banner ad that goes live because the reviewer was fatigued is a reputational hit no budget can justify. The black box imposes a hidden QC tax.
Buyer beware, but with a spreadsheet.
Interesting approach, treating it like an A/B test. I've been trying to track success rates manually in a spreadsheet and it's messy.
You're right about the prompt ambiguity. I've found the same thing. Using simpler, more literal language like "two people standing side by side" helps, but it also gives you much less interesting images. It's a trade-off.
Have you thought about logging the actual prompt text and the defect type (like "extra arm") together? You could maybe spot patterns in what specific words trigger the errors.
Your hypothesis about prompt ambiguity is correct, but I'd frame the core problem as a vendor control deficiency. When you can't use negative prompts to exclude "extra limb," the vendor is essentially refusing to let you define an acceptable output boundary. This is a failure in their risk management framework.
From a security and compliance lens, particularly for SOC 2 or ISO 27001 controls around data processing integrity, this is problematic. You're being asked to integrate an unpredictable system into a production workflow. The "hidden QC tax" another user mentioned is a direct operational risk. For audits, we'd document this as a lack of input validation and output verification controls on the vendor's side.
Your A/B testing approach is the right mitigation, but document it as a compensating control. Log every prompt, the defect rate, and the time spent on manual review. That data becomes your evidence for a risk acceptance decision or, better, your leverage to demand Adobe provide proper negative prompting capabilities.
trust but verify