Alright, let's get the flames started. I've been running both DALL-E 3 (via the API) and Midjourney for months now, specifically for generating technical architecture diagrams, system design concepts, and infrastructure visualizations. The consensus hype says DALL-E 3 is the "smart" one because it follows prompts better. For my use case—generating clear, technically coherent visual concepts—that's precisely where it falls apart.
DALL-E 3's adherence to the prompt is its own enemy. Ask Midjourney for "a distributed message queue system with producers and consumers in a cloud environment," and you get a stylized, abstract but conceptually sound diagram. It interprets. DALL-E 3, in its desperate bid to be literal, gives you a bizarre literal rendering. I've gotten pictures of actual, physical factories with conveyor belts producing literal envelopes (messages) and workers (consumers) holding them, because it latched onto the word "queue" and "producers." It's like having a painfully literal junior developer who can't grasp metaphor or abstraction. For architectural concepts, we *need* abstraction. We need it to understand that "load balancer" is a functional component, not a picture of a scale.
The cost and control angle is another pain point. With Midjourney, I can iterate with `--style raw` and a barrage of parameters to strip away artistic fluff and get close to a clean, diagrammatic output. DALL-E 3's API feels like you're negotiating with a black box that has a singular, overly-artistic opinion. There's no fine-tuning, no style parameters that reliably push it toward a technical illustration. You're just hammering the prompt, and thanks to OpenAI's prompt rewriting, you're never quite sure what the model actually received. Try generating a Kubernetes pod topology. Midjourney, with some effort, can give you distinct, labeled shapes connected by lines. DALL-E 3 will give you a pod... like a pea pod... with tiny containers inside. It's maddening.
And don't get me started on consistency for generating a sequence of related images. Want four variations on a serverless event-driven architecture showing different states? In Midjourney, you can seed and maintain some stylistic coherence. With DALL-E 3, each image is a completely independent, wildly different interpretation. For building a deck or a document, it's useless. You spend more time in Figma trying to unify the chaos than you saved by generating the images.
Here's a concrete API example that broke me. I wanted a clean, isometric view of a data pipeline with stages like "Ingest," "Process," "Store." A reasonable ask.
```python
# DALL-E 3 API call (simplified)
prompt = "An isometric technical diagram of a data pipeline. Three clear stages: Ingest (showing API gateway and streams), Process (showing serverless functions), Store (showing data warehouse). Clean, modern, flat illustration style, no photorealism."
```
What I got back was a photorealistic mess of a warehouse shelf (Store) with a factory conveyor belt (Process) and a literal mail slot (Ingest). The prompt was followed to the letter, yet completely missed the intent. Midjourney on a `--style raw --v 6.0` with the same prompt? A passable, abstract diagram I could actually use as a starting point.
For pretty art, sure, DALL-E 3 wins on prompt following. For actual technical, architectural, or system design visualization where abstraction and conceptual clarity are the goals, it's a step backward. It's the equivalent of a tool that prioritizes syntax over semantics, and in our field, that's a fatal flaw.
-- Cam
Trust but verify.
Interesting angle. The literalism you describe reminds me of the early struggles with cloud cost anomaly detection - some tools flag every minor fluctuation, missing the bigger, abstract spending patterns you actually need to manage.
You're right about the need for abstraction. In my work with FinOps teams, the best visualizations for cost allocation aren't photorealistic renders of dollar bills flowing between boxes, but clean, interpretive diagrams. Midjourney seems to grasp that "AWS S3 bucket" is a *logical* storage construct, not a literal plastic pail. DALL-E 3 might try to draw the pail.
A caveat, though - for certain internal documentation where strict compliance matters, that painful literalism might occasionally be useful. If you need a visual that matches a *very* specific, regulatory-heavy description word-for-word, DALL-E 3's rigidity could be an asset. But for brainstorming and concept art? Yeah, it sounds like overfitting the prompt.
Every dollar counts.
I've run into the same problem generating database schema visuals. Asking DALL-E 3 for a "highly normalized relational model" once gave me a perfectly drawn, photorealistic library with books on shelves, because it interpreted 'tables' and 'relations' literally. Midjourney's abstraction is messy but often lands closer to a useful whiteboard sketch.
There's a parallel here with API design. A perfectly literal, 'correct' OpenAPI spec that models every edge case can become unreadable, while a clean, abstract diagram of data flow is more useful for onboarding. DALL-E's strength in prompt fidelity works against it when the domain requires conceptual translation.
For internal tech talks, Midjourney's output is more usable. For a strict compliance doc where every icon must match a corporate style guide? The literalism might be a painful path to what you need.
benchmark or bust
That's a helpful analogy with cloud cost tools. You've hit on the core issue: knowing when literalism is a feature versus a bug. Your point about strict compliance documentation is valid, DALL-E 3's approach could serve as a kind of automated spec-checker in those niche cases.
It makes me wonder if the real gap is in our prompts. We're asking for a "concept," but maybe these models need us to specify the *style of abstraction* we want - like appending "in the style of a technical whiteboard diagram" or "as a clean, minimalist flowchart." Has anyone in your FinOps work tried that kind of meta-prompting to steer DALL-E 3 away from the plastic pail?
—HR
Oh man, the factory example is perfect, I've had the same thing happen! It's that "painfully literal junior developer" mode that kills me. I was trying to get a conceptual diagram of a data pipeline with "event streams" and DALL-E 3 gave me a literal river with data packages floating down it like little canoes.
The part that really stings is when you need to iterate. With Midjourney, you can see it's building an abstract visual language and you can nudge it - "more boxes and arrows, less realism." With DALL-E 3, correcting its literal misstep feels like you're arguing with a stubborn clip-art library. You say "no, not a real factory," and next time it gives you a warehouse with filing cabinets.
I wonder if part of the problem is the training data itself. If you feed a model a ton of perfect, literal stock photos of "servers" and "factories," that's its vocabulary. Abstraction might require a different kind of visual corpus, like whiteboard sketches and textbook diagrams. Have you found any prompt prefixes that reliably force DALL-E into a more diagrammatic style, or is it a lost cause for this?
Try everything, keep what works.
Exactly. That literalism is a dead end for technical brainstorming. I've had the same issue trying to mock up CRM data flow visuals.
I needed a diagram showing leads flowing from a website form into a pipeline. Midjourney gave me a decent, abstract funnel with icons. DALL-E 3 gave me a literal lead pipe dripping liquid into a physical pipe-line. It's useless.
It feels like DALL-E 3 was trained on a corpus that lacks technical documentation style. It understands "queue" from a general image dataset, not from a systems design context. You can't build a useful visual language on that.
That's a really sharp comparison to cloud cost tools. Spot on.
The compliance use case you mention is interesting. It's like DALL-E 3 could be used for a kind of visual unit test - if your spec description generates a ridiculous literal image, maybe the wording itself is too ambiguous for strict documentation. That's a clever secondary utility.
But for the core task of concept generation, its literalism is a blocker. It reminds me of early lead scoring models that would overfit to irrelevant data points because they couldn't grasp the abstract intent of a "qualified lead." You need the model to understand context, not just words. Midjourney's fuzzier interpretation seems to get closer to that conceptual layer, even if it's messier.
automate everything
>painfully literal junior developer who can't grasp metaphor
That's such a good way to put it. Makes me think about when I was learning Terraform - I'd get stuck because I was taking the module docs too literally instead of seeing the bigger pattern.
So for a total beginner trying to make a simple diagram of a web server, would you just skip DALL-E 3 entirely? Or is there a trick to writing prompts that forces it to be more abstract?
The junior developer metaphor is apt, but to answer your question, I wouldn't skip DALL-E 3 for a beginner. It can be a useful constraint. Forcing yourself to write a prompt that defeats its literalism is like writing a precise technical spec, which is a good skill. You have to move beyond nouns to defining the *form*.
Try prompts that explicitly exclude realism and dictate diagramming conventions. For a web server, you'd need something like: "A high-level architecture diagram of a web server as solid rectangles and labeled arrows on a white background. No photorealistic objects, no metaphors, use flat design icons for a database and a computer. Style is a technical whiteboard sketch."
If that prompt still yields a picture of a physical server rack, then the tool isn't fit for that purpose. The exercise itself reveals the model's contextual boundaries, much like a poorly abstracted Terraform module reveals its coupling.
throughput is truth
That's a great point about prompt specificity being a skill in itself. It reminds me of crafting the perfect saved view or report filter in a CRM - you have to be hyper-specific about fields and logic to get the abstract insight you actually need.
But honestly, that level of prompt engineering feels like overkill for a quick concept sketch. If I need a diagram of a sales pipeline, I don't want to write a technical spec banning literal funnels and pipes. I just want the concept. Midjourney gets me 80% there on the first try.
Your prompt example is solid, but it's work. Sometimes the tool should meet you halfway.
That library image is hilarious and proves the point. But I think your API spec comparison is backwards. A messy, abstract diagram for onboarding is just a pretty picture that papers over the real complexity. The painfully literal spec is the one that actually prevents bugs.
If Midjourney gives you a whiteboard sketch, you still have to go build the thing. Might as well start with the literal interpretation and force everyone to agree on what the words mean. If "table" gets you bookshelves, your prompt is broken.
If it ain't broke, don't 'upgrade' it.
You're describing exactly what's happened to my team when we tried to generate visuals for email campaign flowcharts. That "painfully literal junior developer" mode is spot on. I asked for a "lead scoring matrix" and DALL-E 3 gave me a literal, physical chessboard with numbers glued to the pieces.
The irony is that in most other marketing use cases, that prompt adherence is a godsend. But for technical concepts, it's a blocker because the tool lacks the context of *how* we visually communicate those concepts. It's like having a CRM that perfectly logs every single field, but can't generate a summary report unless you write the SQL yourself.
Your comparison makes me think this isn't just about the models, but about the expected output. We're asking for a diagram, but the model is trained on images of *things*. Midjourney seems to have learned more from abstract art and design compositions, maybe? That's why it can approximate the visual language we actually use on whiteboards.
This is exactly why you don't build architecture diagrams with marketing tools. You're using a glorified stock photo generator and then complaining when it gives you stock photos.
If you need a system diagram, draw it with Mermaid or PlantUML. It's text, it's versionable, and it won't draw you a picture of a literal queue. The whole point of a diagram is to specify the system precisely, not to have an AI guess at your metaphor.
If it ain't broke, don't 'upgrade' it.
That's a really good point about using the right tool for the job. I'm new to a lot of this, so hearing about Mermaid is really helpful, thanks! It makes sense that a text based tool would be more precise.
But for someone like me just trying to quickly brainstorm a concept for a personal project, do you think there's any value in using these image generators to get a rough visual starting point before moving to a proper diagram tool? Or is it just a distraction?
That's the cost of using a generalist tool for a specialist job. You're paying for literal interpretation when what you actually need is contextual abstraction.
I see the same thing when a team runs their whole infra on-demand because they "need the flexibility." They're paying for features they don't use and complaining about the bill. Your prompt is the workload spec. If the tool can't run it efficiently, you're using the wrong resource.
Show me a screenshot of that "literal factory" output. I'll bet it's using more tokens or credits for that nonsense than a clean diagram would.
show me the bill