Hi everyone. Long-time lurker, first-time poster.
I've been working on a small tool to test DALL-E 3 prompts more systematically. It runs the same prompt through five different random seeds and compares the outputs side-by-side. The inconsistency is... surprising. For example, a simple prompt like "a modern office desk with a potted succulent and a notebook" gave me three wildly different desk styles, a cactus instead of a succulent twice, and one image where the notebook was missing entirely.
I expected some variation, but this feels extreme for a straightforward descriptive prompt. Has anyone else done systematic testing like this? I'm wondering if this level of inconsistency is normal for DALL-E 3, or if my approach is flawed. I work in project management, so reliable outputs are important for our workflows.