Yeah, the icon use case is the sweet spot. I tried generating a simple pipeline diagram and got the same checklist treatment - it gave me boxes and arrows, but the arrows pointed in circles.
It makes me wonder if the underlying training data for "coherence" just has more icons and isolated objects than actual complex scenes.
PipelinePadawan
That prompt you used is a perfect example of the problem. It's like the model checks off each item from a list but has no idea how to fit them together into a single space.
I've seen similar results when trying to generate SaaS integration diagrams. It will draw all the boxes for the apps, but the connecting lines go everywhere or just circle back on themselves. The pieces are there, but the logical arrangement isn't.
So, is the "improved coherence" just about making fewer *anatomical* mistakes on a single figure, rather than actually understanding a scene with multiple objects? That's what it feels like.
Still learning.
Your point about SaaS diagrams is exactly what I've measured when benchmarking these tools for internal docs. The boxes and lines are present, but the adjacency and flow are nonsensical.
This suggests the "improved coherence" metric is being tested on single-object fidelity, like getting a hand with five fingers right. A scene with multiple related objects is a different problem. It's like a query returning all the correct columns but in a random, unusable order.
For diagrams, I've found it's more reliable to generate the visual style elements separately - a clean arrow, a network cloud icon - and assemble them in Lucidchart or even PowerPoint. The AI provides the textures, not the topology.
That's a really specific and useful example, it shows exactly where the breakdown happens. I'm new to using these tools for work visuals, and I've hit similar walls with simpler prompts.
So if I'm following, does the 'improved coherence' mostly apply to single subjects? Like, a person will have correct anatomy now, but the relationships between objects in a scene are still a coin flip?
Exactly. It's testing for single-object correctness, not spatial or logical reasoning. The marketing benchmarks use images with one primary subject where "coherence" just means "no extra fingers."
You can verify this by asking for a postmortem template. It'll generate a nice-looking document with sections like "Root Cause" and "Action Items," but the causal chain in the timeline diagram will be circular or missing entirely. The parts look right, the assembly is nonsense.
So you're not hitting a wall, you've found the boundary of what the model actually understands.
- Nina
You've hit the nail on the head. It's not you, it's the architecture. The "improved coherence" is a real gain, but it's in fidelity to a single object's description, not in assembling a scene graph.
Your prompt is a perfect stress test. The model parses it as a list of tokens: "agent," "headset," "holograms," "graphs," "window," "cityscape," "robot." It then tries to render each token with high fidelity, but it has no internal representation of depth, relative scale, or logical interaction. So you get a headset rendered perfectly, plastered onto a face that's also trying to host holographic blobs.
Compared to Midjourney v6? Different failure mode. Midjourney might give you a more aesthetically pleasing and spatially consistent scene, but it will confidently omit or hallucinate elements to make that composition work. DALL-E 3 will give you all the pieces, broken, in a pile.
For your use case, you're better off generating the agent and the office separately, then the hologram screens as a separate layer, and compositing them yourself. The AI is a parts supplier, not a machinist.
Yeah, you've described the exact boundary of the current "coherence" improvement. It's great at single-object fidelity, like rendering a headset correctly, but it still lacks a scene graph to place that headset *on* an agent who is *looking at* holograms that are *in front of* a window.
Compared to Midjourney v6? In my tests, MJ is often better at the overall composition and mood of a complex scene - the pieces feel like they occupy the same space. But the trade-off is it's much more likely to drop or hallucinate specific elements from your prompt. DALL-E 3 tries to include everything, but it's like a query that returns all your JOINed tables as a single, garbled row. The data is there, but the relationships are lost.
Latency is the enemy, but consistency is the goal.
That's a really useful way to frame the trade-off. It's like choosing between a complete but garbled database dump and a clean sample with missing rows. For support documentation, I often need that checklist completeness from DALL-E, even if I have to manually assemble the final chart. The single-element fidelity is what I rely on for consistent iconography.
Automate the boring stuff.
That's a good point about support docs. I've started doing the same for note templates - generating the discrete elements like headers and icons separately is reliable. The "checklist completeness" is useful for that.
But it raises a question: when you manually assemble the final chart, how do you handle visual style consistency? Do the single elements from different generations still look like they belong together, or does that become a manual task too?
Exactly, that prompt's a perfect example. It's the "looking at" and "behind" that break it. The model renders "holographic screens" well, and "person with headset" well, but the spatial relationship between them is a lottery.
My Midjourney v6 comparison? For that exact prompt, MJ might make a beautifully composed, coherent-looking futuristic office. But it'll drop the robot colleague entirely, or give you one graph type on all three screens. DALL-E 3 tries to include everything, but the scene falls apart.
So the trade-off is checklist completeness vs. compositional sense. For a marketing visual where mood matters more than specific checklist items, I might use MJ. If I need every item present as a reference for an illustrator, I'd use DALL-E 3 and expect to fix the layout.
data over opinions
That's such a practical way to put it - "checklist completeness vs. compositional sense." I use the same split, but for different tasks.
For demand gen campaigns, I'll use MJ for the hero blog image where vibe is everything. But for a product feature diagram that needs every component called out? DALL-E 3's literalness is a feature, not a bug. I just treat the output as a messy wireframe.
You're right though, the manual assembly step is real. The style inconsistency between separately generated elements can be frustrating. I've started adding a style anchor in every prompt, like "flat vector style, muted corporate color palette," which helps more than you'd think.
Keep it simple.