Skip to content
Notifications
Clear all

Am I the only one who finds the 'improved coherence' claim misleading for complex scenes?

41 Posts
39 Users
0 Reactions
70 Views
(@carlam)
Reputable Member
Joined: 3 months ago
Posts: 234
Topic starter   [#27097]

Okay, I've been putting DALL-E 3 through its paces for a few weeks now, mostly for creating marketing visuals and some concept art. The 'dramatically improved coherence' is the headline feature everyone's talking about, right?

But when I push it beyond a simple subject-on-background image, things get... weird. I'm talking about scenes with 3-4 specific elements interacting in a specific environment. For example, I tried: *"a customer service agent with a headset, looking at three floating holographic screens showing different graphs, in a futuristic office with a large window showing a cityscape, with a robot colleague passing by in the background."*

What I got was a mess. The agent might have a headset, but the holograms are just blobs or are superimposed on her face. The robot colleague is often just a metallic smudge near her shoulder. The window might be there, but the cityscape is a painterly blur that doesn't match the office's perspective. It *feels* like the prompt is understood, but the composition falls apart.

So my question is: how does this "improved coherence" actually compare to something like Midjourney v6 for the same type of complex, multi-element scene? Is DALL-E 3's strength more about single-subject fidelity and text rendering, rather than true scene composition? I'd love to see side-by-side benchmarks from anyone who's tested both on similar detailed prompts.

What's everyone else's experience? Are you getting great results with complex scenes, or are you also having to simplify your prompts way down to get something usable?


Benchmarking my way to better decisions


   
Quote
(@cloud_cost_owen)
Reputable Member
Joined: 6 months ago
Posts: 181
 

Oh totally. It's that classic "overprompting" wall. You get the feeling it parsed every word, then just... assigned them a layer at random.

I've had better luck breaking it into a sequence. Like:
1. "a futuristic office with a large window showing a detailed cityscape"
2. Take that output, feed it back in with: "in the office from the previous image, add a customer service agent with a headset looking at three floating holographic screens"
It's a hassle, but it layers the scene better than a single monolithic prompt.

For a *truly* complex single-shot image, Midjourney v6 still feels more like it's building a cohesive 3D space in its head. DALL-E 3 is amazing for prompt adherence and readability, but spatial reasoning across many elements isn't its strong suit yet.



   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

You've hit on the fundamental architectural trade-off. DALL-E 3's "improved coherence" refers primarily to its language model's superior text understanding and prompt fidelity, not to spatial or compositional coherence within the image plane.

It's a prioritization of semantic token alignment over 3D scene synthesis. Your prompt is perfectly understood as a list of objects and attributes, but the system lacks a persistent internal model of the scene to place those objects in correct relative scale, occlusion, and perspective. The holograms become blobs because "floating in front of" is a spatial relationship that its diffusion process struggles to enact consistently against other strong visual tokens like "face."

Comparing it to Midjourney v6 for this specific task is apt. Midjourney often sacrifices strict prompt adherence for stronger internal compositional harmony, giving it an edge in building that "cohesive 3D space" from a complex prompt. For your use case, the sequential layering method mentioned by the other user is effectively a manual workaround to impose the compositional coherence the model lacks in a single pass.


—BJ


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

Yeah, spot on. Your futuristic office example is exactly the problem I ran into. The "coherence" is about the language, not the physical space. It'll nail every noun you listed, but then the scene composition gets jumbled.

I've had the same results trying to generate a simple CI/CD pipeline diagram with multiple services and arrows. It renders all the icons, but the logical flow and layout are a complete mess.

For complex scenes, I'm using it as a component generator now, like user370 said, then compositing the pieces myself in an editor. It's a bummer for a single-shot concept.


Automate everything.


   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Right? The marketing really hammers on "coherence" but they're talking about a different thing entirely. It's not the image that's coherent, it's just the prompt understanding. It'll read your list perfectly, but it's like giving perfect instructions to a chef who then throws everything in a blender.

Your example hits home. For our team's training guides, I tried getting a simple scene of a person watching a tutorial on a laptop, with a phone showing a notification and a notebook open next to them. Got all the objects, but the notebook was floating mid-air and the phone was somehow behind the person's head. The spatial logic just isn't there yet for packing multiple elements into a realistic layout.

So for complex scenes, I've stopped expecting a final image. I use DALL-E 3 to generate clean, individual assets (your agent, a nice hologram blob, a robot silhouette) and composite them myself. It's an extra step, but the parts are high quality. For a single, cohesive shot, you're right to look elsewhere.


Trust the trial period.


   
ReplyQuote
(@emmae)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Oh wow, this is super helpful to read. I'm just starting to use DALL-E 3 for mock-ups in Salesforce and ran into the same wall with what I thought was a simple request.

I tried something like "a sales manager pointing at a dashboard on a monitor, with a customer 360 view on a tablet and a stack of reports on the desk." I got all the items, sure, but the manager's hand was going through the monitor and the tablet was just... stuck to the side of his head? It looked ridiculous.

So when you say the 'coherence' is about language, not space, that clicks. It understood my nouns perfectly but had no idea how a desk works. Is the layering trick the only real workaround right now, or are there specific prompt words that help a tiny bit?



   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Yeah, that's exactly it. It's like it's doing perfect keyword spotting but failing at the spatial arrangement. Your example with the holograms becoming blobs or overlapping the face is so familiar.

I've tried something similar for generating diagrams of a multi-service architecture. It would draw perfect little server and database icons, but they'd be stacked on top of each other or connected with arrows going in impossible loops. The "coherence" is totally about text matching, not about building a logical, physical layout.

I'm curious, have you tried using the sequential layering trick that was mentioned? I wonder if that works better for tech scenes or if the problem is just as bad.


Learning by breaking


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

Your CI/CD pipeline example hits the nail on the head. The tool understands the words "service," "arrow," and "pipeline" but has zero concept of directed acyclic graphs or logical flow. It's just placing tokens.

Using it as a component generator is smart, and honestly, that's where it's useful for professional work. The real bummer is that the layering trick starts to fall apart when you need those components to interact spatially in the final composite. You still have to do the layout yourself.

For diagrams, I've had slightly better luck feeding it a rough mermaid.js output as a text prompt, but even then it just treats it as a texture. The spatial logic just isn't in the model.



   
ReplyQuote
(@hannahj)
Reputable Member
Joined: 3 months ago
Posts: 290
 

You're absolutely right about the mermaid.js output being treated as a texture. It highlights that the model is interpreting visual patterns, not structural semantics.

I've had similar results trying to use it for system architecture by feeding it Graphviz DOT language descriptions. It renders boxes and lines, but the layout bears no relation to the graph's topology. The 'rankdir' attribute is completely ignored.

This reinforces that for any logical or spatial structure, we're still firmly in the component assembly phase. The tool excels at generating a convincing icon of a database, but it cannot arrange three of them in a meaningful cluster.


Data is the new oil – but only if refined


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 3 months ago
Posts: 270
 

You're right about the component assembly being the only real workaround right now. I've tried the sequential layering trick for a similar tech marketing scene - a conference booth with a demo station and a banner - and it's a mixed bag. It helps isolate objects, but you still hit a wall when you need them to interact.

For your diagram problem, I found that generating icons separately - like a clean server icon, a separate database icon - and then manually placing them in a slide deck was faster than fighting for a coherent layout from a single prompt. It does feel like admitting defeat, though.

I'm wondering if the "blobbing" and impossible arrows happen because the model is trained on countless images of objects, but far fewer that diagram complex, abstract relationships in a clean way. Does that track with what you've seen?



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Yep, you've nailed the exact frustration. Your hologram example is a perfect illustration of where the "coherence" claim meets reality. It's great at parsing every word in your list, but it has no internal blueprint for where those items should actually go in 3D space.

For your comparison question, Midjourney v6 can sometimes feel like it has a slightly better gut feeling for composition and depth in these packed scenes, but honestly, it's a different flavor of the same fundamental problem. It might make the robot colleague look cooler, but the holograms could still be weirdly fused with the agent's desk. The core issue is that these models aren't building a scene layout, they're painting tokens that match your words.

My workaround for marketing scenes has been to generate the office and the agent separately, then the holograms as transparent PNGs, and composite them myself. It's an extra step, but it's the only way to get that spatial logic right.


ship it


   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Exactly. It interprets prompts as a checklist, not a spatial layout. I use it for icons, but that's it.

For realistic compositing, your manual method is the only reliable one right now. It's not a workaround, it's the actual workflow.


Metrics don't lie.


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Yep, you're hitting the core limitation. The "improved coherence" is purely linguistic. It parses your entire sentence flawlessly, then fails to construct a scene graph from it.

For your exact hologram scene example, I'd argue Midjourney v6 wouldn't solve it either, just fail differently. It might produce a more pleasing aesthetic blur, but the spatial relationships would still be nonsense. The root issue is neither model has a true physics or layout engine; they're generating plausible pixel arrangements for text tokens.

The real comparison isn't DALL-E 3 vs. Midjourney for this use case. It's whether you're better off using either as a glorified stock photo generator for individual components, which you then composite manually. For any professional marketing visual requiring specific object interaction, that's still the only workflow that works reliably. These tools are component factories, not scene assemblers.



   
ReplyQuote
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

You're not alone. That hologram example is a perfect case study. The "improved coherence" means it won't forget the robot colleague or the headset, but it has zero ability to place them correctly in 3D space.

I've seen the same thing trying to generate pipeline diagrams. It draws perfect little server icons and arrows, but the arrows point in circles and the servers are stacked on top of each other. The model has no internal concept of layout.

For your use case, generating components separately is the only practical workflow. Get your futuristic office, your agent, your robot, and your hologram screens as individual images, then composite them yourself. It's more manual, but it's the only way to get the spatial relationships right. Treating it as a component generator is a defeat, but it's the only thing that works reliably.


Build once, deploy everywhere


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Exactly. The "coherent" marketing visuals are a trap for anyone trying to build a scene with actual objects. It's not a scene generator, it's a texture generator that sometimes arranges those textures into familiar shapes.

Your pipeline diagram example is perfect. It's the same core problem, just abstract instead of physical. It can render the *idea* of a pipeline, but not the *structure*.

So yeah, the manual compositing workflow you described isn't a workaround, it's the job now. The tool just provides the raw materials. Anyone expecting it to do the layout is setting themselves up for frustration.


CRM is a necessary evil


   
ReplyQuote
Page 1 / 3