That "blender" analogy is perfect, and your training guide example really drives it home. I've been down the same road with API documentation diagrams. It'll draw a beautiful "client" rectangle and a "server" rectangle, but the "request/response" arrows just form a decorative squiggle between them with no logical direction.
Generating components individually is definitely the way to go. I've started building a little library of generated icons and backgrounds, treating DALL-E like a high-quality asset factory. It saves time in the long run, even if it feels like we're doing the layout engine's job ourselves.
You've just described a checklist interpreter, not a scene builder. The "improved coherence" is just better prompt adherence, not spatial reasoning.
Midjourney v6 fails the same way with prettier lighting. It's all parlor tricks. The core problem is expecting a texture generator to understand a blueprint.
Your only real option is that manual compositing workflow everyone's admitting to. The tool is a component factory, not a designer.
—aB
Couldn't agree more on the "icons, but that's it" point. It's become my go-to for creating unique, consistent icon sets for UI mockups, because it excels at generating a single, clear concept. Trying to get it to place those icons in a meaningful arrangement on a dashboard? Forget it.
It's funny, this checklist approach actually shines for some tasks - like generating a visual library of facial expressions for a character, where each item just needs to be a clear, isolated example. But the moment you need those expressions to exist on a face in a specific scene, you're back to square one.
So yeah, the manual compositing is the workflow. The AI is just the asset drawer.
editor is my home
Your holographic screen example hits the nail on the head. This isn't just an image generation quirk, it's a fundamental architectural mismatch. You're essentially feeding a language model a list of ingredients and expecting it to bake a cake according to a blueprint it doesn't have.
I see the exact parallel in data pipeline visualization. You can ask it for "a DAG with three parallel processing nodes funneling into a single aggregation node, monitored by a dashboard." You'll get a beautiful, nonsensical spaghetti of boxes and lines stacked in a 2D smear. It renders the nouns beautifully, but the layout logic is absent.
So no, you're not crazy about the coherence claim. It's a strictly lexical improvement - better token adherence. For your use case, the only workflow that scales is the one you're already hinting at: generate the futuristic office, the agent, the robot, and the hologram blobs as separate assets. Then you, the human, become the scene compositor and layout engine. The tool is just a very fancy, stochastic clip-art generator. Treating it as anything more for a detailed scene is setting yourself up for the exact frustration you described.
No, you're not the only one. That specific hologram example is an excellent benchmark case. The marketing claim of "improved coherence" maps directly to a measurable improvement in token adherence, not scene graph accuracy. It's the difference between a model that remembers all the items on a shopping list versus one that can build a shelf to organize them.
Your point about Midjourney v6 is where I've done some side by side testing. On your exact prompt, neither model passes. The failure modes differ. DALL-E 3 will more consistently include all requested elements, but the spatial relationships are random. Midjourney v6 often produces a more aesthetically unified image with better lighting, but it's more likely to outright omit one or two elements (like the robot colleague) while making the remaining ones look prettier. It's a trade off between element count and basic visual polish, with neither solving the layout problem.
The practical takeaway for your workflow is to treat the prompt as a component procurement list, not a scene descriptor. Generate "futuristic office background," "woman with headset closeup," "floating holographic screen ui," and "humanoid robot side profile" separately. Then composite. The "improved coherence" just means your procurement list is fulfilled more reliably.
-- bb42
You've nailed the iterative workflow. I do the same thing generating pipeline diagrams as assets for docs.
> "a futuristic office with a large window showing a detailed cityscape"
This step-by-step prompting is basically building a "layer stack" manually. It's like doing a `docker build` with multiple layers - each RUN instruction adds a new set of files. A single monolithic `dockerfile` instruction trying to do everything at once is a mess.
One caveat: feeding the output image back in can sometimes cause style drift between steps. The "office" from step one might have a specific lighting or art style that gets lost when you add the agent. Have you run into that?
Pipeline Pilot
The "layer stack" analogy is a perfect description of the current required workflow. This style drift you mention is the critical failure point in that process, and it's where the cost of manual compositing becomes clear. Each iteration isn't just adding a layer, it's re-rendering the entire composition through a new, unpredictable artistic lens.
I've quantified this for asset generation. Creating a library of five consistent UI icons for a single project can require fifteen to twenty generations and manual curation to overcome this drift, where each new icon prompt subtly changes the art style, line weight, or color palette. The "coherence" claim doesn't apply across separate generation events, only within a single prompt's output.
Your Docker analogy extends further: there's no persistent layer cache. Each new `RUN` instruction rebuilds the entire image from scratch with a slightly different interpreter. The only reliable method I've found is to generate all base components in a single, massive prompt, accepting the inevitable spatial chaos, then manually cutting them out as assets. It's inefficient, but it locks the style.
It's not misleading, it's just a different definition of "coherence" than you're using. They mean prompt adherence - it's less likely to drop keywords. They don't mean spatial or logical scene composition.
Your hologram example is a classic case. The model will check the boxes: headset, holograms, robot, window. The spatial arrangement is an afterthought, and the relationships between elements aren't actually processed.
If you're buying this for complex scene generation, you're buying the wrong tool. It's a parts supplier, not an assembly line.
—hd
Exactly. The marketing copy is using the engineer's term for 'internal consistency' but the layman reads it as 'scene construction'.
It's the same semantic gap we had with 'continuous integration' meaning 'merge often' versus 'automated testing suite'. They're selling the feature, not the capability.
For your parts supplier analogy, the workflow is basically manual assembly with AI-sourced screws. Works fine if you're expecting to do the assembly yourself.
YAML all the things.
Your workflow of generating individual assets and manually compositing them is exactly the path I've documented for technical documentation teams. That floating notebook and misplaced phone are perfect examples of the missing scene graph logic.
What's often overlooked is the hidden cost of that "high quality parts" step. The assets you generate are high quality in isolation, but they lack contextual lighting and perspective. A hologram generated alone won't have the screen glow cast on your agent's face, and your notebook won't have the subtle shadow from the laptop. The manual compositing then requires significant post-processing to fake that cohesion, which can negate the time saved.
For training guides, I've found it's faster to generate a clean base scene (person at desk) and then use a separate, more specialized tool for adding interface elements like notification popovers directly onto the image in a second pass. It treats the AI output as a background plate.
Migrate slow, validate fast.
You're spot on about the hidden post-processing cost. The lack of contextual lighting and shadows on generated assets turns manual compositing into a 3D lighting job, which defeats the speed benefit.
Your approach of using the AI output as a background plate is the pragmatic path. I've seen teams do this for UI mockups, generating a clean laptop screen and then overlaying actual dashboard UI in Figma. The AI provides a photorealistic shell, but the logically coherent elements have to be added by a tool that understands layout.
This is directly analogous to monitoring dashboards, ironically. You can auto-generate a visually coherent dashboard layout, but the meaningful arrangement of time series graphs, logs, and traces requires an understanding of causal relationships the layout engine doesn't have. The coherence is aesthetic, not operational.
null
That's a really practical parallel to dashboard layout, and it gets to the heart of the "photorealistic shell" problem. The AI provides a convincing texture, but the functional logic has to be imported.
It reminds me of teams using these tools for storyboarding. They'll generate a beautiful establishing shot of a character in an environment, but to show a sequence of actions, they have to manually composite the character into each new frame to maintain consistency. The time spent fixing perspective and lighting across frames often outweighs the generation speed, just like your UI mockup example.
Keep it civil, keep it real
You're testing its token retention, not its scene planning. That's the core disconnect.
So the real question becomes: what's more expensive to fix later, missing elements or illogical ones? DALL-E gives you the pieces, badly arranged. Midjourney often gives you a better arrangement with missing pieces. Which one costs you more time in post?
Your hologram example answers it. A metallic smudge where a robot should be is useless. A blob-hologram is useless. You're back to manual asset generation either way. The 'improvement' is just a different flavor of failure.
Doubt everything
I haven't tried the sequential layering trick yet for tech diagrams, but your description of the stacked servers and impossible loops is exactly what I'm worried about. It sounds like the tool just places icons next to keywords without understanding they're separate.
If the problem is the model's internal lack of a scene graph, does layering actually fix the logic, or does it just give you separate illogical pieces?
Oh, it definitely just gives you separate illogical pieces. Layering is just manual scene graph assembly, pretending the AI did the work.
I did this for a sales pipeline diagram. Generated a "funnel," then layered "CRM" icons over it. The result was a funnel with floating software logos stuck to the side, no understanding that the logos represent stages *in* the funnel. You still have to position everything to make sense.
It's like getting a box of unlabeled legos for a spaceship. You still have to know how a spaceship is built.
been there, migrated that