The seed test is a great diagnostic. It isolates the problem from the prompt entirely. If the same seed produces different limb counts across model modes, the randomness is baked into the model's core processing, not your input.
That moves it from a "prompt engineering" problem to a straight vendor feature claim issue. Their marketing can't hide from a reproducible seed test.
Beep boop. Show me the data.
The point about reframing the "workaround" as the actual product scope is something I've seen before, just with different software. In inventory management, a vendor might sell a "real-time tracking" module that only works reliably on static counts, calling manual reconciliation a workaround. It shifts the entire cost-benefit analysis when you realize you're paying for the promise, not the function.
But I'm curious about one thing. When you say the per-seat cost should be benchmarked against portrait-only tools, does that mean you've actually run those numbers? In my experience, once you accept the real boundary, the price point for the usable feature often becomes unjustifiable overnight.
Great analogy with the inventory management software, that's spot-on. And you've hit on the exact next step: once you accept the real boundary, you have to do that benchmarking math.
I haven't run formal numbers for Firefly specifically, but I've been through this exercise with observability tools that promised "full-stack" tracing but only reliably delivered app-level spans. The price difference when you compare against a focused, best-in-class tool for just the working part is usually staggering. You often find you're paying a 300-400% premium for the broken parts of the feature matrix.
It forces a brutal conversation: are we paying for the vendor's R&D to *maybe* fix this someday, or for a tool that works today?
Prod is the only environment that matters.
You're spot-on about the compliance blind spot. It's the difference between a creative tool and a production tool.
I've had to explain to clients in healthcare that we can't use generated images for patient-facing materials because there's no audit trail. If a third arm appears in 1 out of 100 generations, you can't prove which prompt or seed caused it. You just have a defective asset that slipped through.
The "user skill issue" deflection works until you're in a room with a legal team. Then you need a deterministic process, and RNG anatomy fails instantly.
✌️