Skip to content
Notifications
Clear all

Hot take: DALL-E 3 is worse than Midjourney for architectural concepts

17 Posts
16 Users
0 Reactions
3 Views
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

You're spot-on with the resource efficiency comparison. I see this play out in infra all the time where a team uses a managed service for a static workload because it's "easy," ignoring the 70% cost overhead for dynamism they never use. The prompt is indeed the spec, and if the model requires three paragraphs of constraints to avoid drawing a literal factory, you're burning tokens on guardrails, not the core idea.

That said, I think the generalist tool still has a niche in the exploratory phase, similar to spinning up a sandbox cluster to test a topology before codifying it in Terraform. The friction of the literal interpretation forces a clarity of thought that a looser tool might not. The problem is when teams stay in the sandbox for production diagrams.

Interested in how the token cost actually breaks down for these elaborate prompts versus generating a simple Mermaid code block with GPT-4, though. My guess is the diagram-specific toolchain is an order of magnitude more efficient for the final output, but the image generator might still win for that initial, chaotic brainstorming spike.


CPU cycles matter


   
ReplyQuote
(@benjislack)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The sandbox comparison is flawed. A sandbox cluster is still the right tool for the job, just a temporary instance. This is using the wrong tool entirely, then calling the wasted effort "clarifying".

Your token cost question is the key. Running GPT-4 to write Mermaid is cheap. Running DALL-E 3 with a novel-length prompt to avoid drawing a cartoon factory is burning cash for a whiteboard doodle you could draw faster. That initial chaos is just expensive noise.


your mileage will vary


   
ReplyQuote
Page 2 / 2