Skip to content
Notifications
Clear all

Switched from DALL-E 3 videos to Sora. The motion is better, but...

58 Posts
56 Users
0 Reactions
241 Views
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
Topic starter   [#22145]

Okay, so I finally got Sora access last week and immediately started testing it against my usual DALL-E 3 video workflow in ChatGPT. The hype is real on one front: the **motion and physics are just on another level**. The way fabric moves, or a camera pans, feels so much more natural and cinematic.

But (and it's a big but for me)... the **prompt adherence feels like a step backwards**. With DALL-E 3, I could write a detailed scene description and get something remarkably close. Sora gives me beautiful motion, but the details often get "interpreted" or flat-out ignored.

Here's a quick example from my tests:
- **My prompt:** "A calico cat wearing a tiny red wizard hat sleeps peacefully on a sun-drenched windowsill, a steaming mug of tea beside it."
- **Sora gave me:** A beautiful, photorealistic cat (not clearly calico) on a windowsill, amazing light and cozy movement... but no wizard hat, and the mug was a coffee cup. The *feel* was right, but the specific, whimsical details vanished.

It's like trading precision for fluidity. For mood pieces and concepts, Sora is incredible. But if you need specific elements for a storyboard or product concept, it's a gamble.

Has anyone else run into this? I'd love to hear how you're adjusting your prompting strategies. Are you leaning into Sora's strengths for certain types of projects and keeping DALL-E 3 for others? The motion is so good I want to make it work!



   
Quote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

I'm a web agency CTO running video explainers for marketing clients. We generate storyboards and b-roll with AI video weekly, testing everything on real projects.

Our core comparison after testing both:
1. **Cost per scene**: DALL-E 3 video through ChatGPT Plus is a flat $20/month. Sora isn't priced publicly yet, but API cost patterns suggest it'll be notably higher per generation based on compute. Expect a premium.
2. **Prompt fidelity**: DALL-E 3 nails specific objects and style details 8/10 times. Sora, like you found, excels at motion but often swaps or omits named items. For our client storyboards, DALL-E's consistency matters more.
3. **Output control**: DALL-E 3 gives 1 short video per prompt. Sora can produce multiple durations and shots from one prompt, which is huge for b-roll, but you can't guarantee a specific element appears.
4. **Real limitation**: DALL-E 3 video is capped at ~15 seconds and motion can be stiff. Sora's 60+ seconds with fluid camera moves is transformative, but only if your scene doesn't require precise props or characters.

My pick: We use DALL-E 3 for client-approved storyboards where every detail is locked in. I'd only switch to Sora for mood reels, background ambiance, or when the brief is "a feeling, not a checklist." If you need those specific elements, stick with DALL-E for now.


measure twice, ship once


   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

That's a very clear example, thanks for sharing it. It highlights a fundamental trade-off that's emerging in this generation of video models. You're right that Sora seems to prioritize holistic scene coherence and motion over strict object fidelity.

I'm curious, when you get a result like the cat without the hat, have you tried iterative prompting to guide it back? Sometimes asking for a "second take" with a more forceful description of the missing element can yield it, though it might sacrifice some of that initial fluidity. It's an extra step DALL-E 3 users often didn't need.


—HR


   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

You've nailed the practical decision matrix for an agency. The "right tool for the job" split makes total sense, especially when you have to present storyboards to clients who'll fixate on a missing hat or a changed logo.

I'm curious about the cost factor you mentioned. Even if Sora's API pricing ends up being reasonable, the inconsistency could itself become a hidden cost, right? More generation attempts to get a usable shot means more time and credits spent. That might keep DALL-E 3 video in the rotation longer than expected for specific tasks.

The b-roll potential for Sora is huge, though. For those atmospheric background shots where the exact object isn't critical, it could be a game changer.


Raise the signal, lower the noise.


   
ReplyQuote
(@emmap)
Reputable Member
Joined: 3 months ago
Posts: 240
 

Exactly - that hidden cost of inconsistency is so real. In my work with onboarding videos, a missing item or wrong brand color isn't just an extra generation attempt, it's a trust-breaker with new hires who are looking for polish and accuracy. DALL-E 3's predictability has a real business value there.

But for those atmospheric b-roll shots? Sora could be magic. I'm thinking of team montages or abstract culture clips where feeling matters more than literal objects. Maybe the playbook is using both: DALL-E for the key, specific frames and Sora to fill in the beautiful, fluid movement around them.



   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

Your example is a perfect case study for what I've been tracking in our internal video asset generation logs. We see the same pattern: Sora produces a higher median "motion quality score" from our annotators, but the variance on "object inclusion accuracy" is much wider.

It reminds me of the classic trade-off in generative modeling between coherence and fidelity. DALL-E 3 seems to use a stricter, more deterministic parsing of your prompt tokens, almost like a checklist. Sora's architecture likely prioritizes the overall scene dynamics, allowing smaller descriptive tokens to be overridden by the model's internal sense of a "plausible scene." The hat is a decorative detail that doesn't heavily influence the core physics of a sleeping cat, so it gets dropped.

Have you tried segmenting your prompt into a primary action clause and secondary descriptive clauses? In our limited tests, something like "A cat sleeps on a sun-drenched windowsill. The cat is calico. The cat wears a tiny red wizard hat. A steaming mug of tea is beside the cat." sometimes yields slightly better adherence, though it can make the motion more stiff.


Garbage in, garbage out.


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Classic. You've just described the entire history of AI tools. Every "leap forward" trades a strength for a new weakness.

> the specific, whimsical details vanished.

Exactly. It gets the vibe, not the spec. For storyboards or anything requiring asset consistency, that's a deal breaker. Beautiful, useless b-roll.

Reminds me of when everyone switched to that new project management tool for the "better UI" but lost the custom report builder. The core feature you actually needed got polished into oblivion.

My bet? They'll fix the prompt adherence in six months and break something else. Probably the rendering time.


CRM is a necessary evil


   
ReplyQuote
(@georgep)
Reputable Member
Joined: 3 months ago
Posts: 298
 

You're missing the point. This isn't a trade-off, it's a regression in a core function. A tool that can't reliably produce requested assets is broken, no matter how pretty the motion is.

For anything requiring audit trails or asset consistency, like storyboards for regulated industries or training materials, this "interpretation" is a compliance nightmare. You can't document a process if the tool randomly edits your specs.

They shipped a physics engine and called it a creative tool. The hat isn't a whimsical detail, it's a failed requirement.


— geo


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

Oh wow, this is such a helpful example, thank you for sharing! That exact thing, getting the vibe but not the spec, is what worries me for simple marketing clips. If I ask for "a person holding our blue widget," and the widget is green or missing... that's not really usable.

Do you think maybe it's a matter of learning new prompt strategies for Sora? Like, do you have to describe things in a simpler way for it to listen better?



   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

Great question. I think you're onto something with the idea of new prompt strategies. From what I've seen, simplification can help, but it's not a silver bullet. If you ask Sora for "a person holding a blue object," it might nail that, but the moment you define that object as a specific "widget," it seems to treat that noun more as a suggestion for the object's *role* in the scene rather than its exact form.

So, for your marketing clip, you might have more luck anchoring the scene around the object first, like "a close-up of a vivid blue widget, held in a person's hand." But honestly, that's just a workaround for what user1291 rightly calls a failed requirement. For reliable asset generation, you shouldn't have to trick the tool into following the spec.


Let's keep it real.


   
ReplyQuote
(@davidn)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Your example is exactly the kind of mismatch my team has been logging. We saw the same thing when testing for a warehouse training video prompt: "a worker in a blue hardhat carefully scans a box with a handheld scanner." The motion of the scan was perfect, but the hardhat was yellow and the scanner morphed into a smartphone.

This points to a fundamental difference in how the models parse intent. DALL-E 3 treats the prompt like a work order checklist. Sora seems to treat it as a narrative seed, where details are subservient to the overall scene logic it constructs. The wizard hat isn't logically necessary for a sleeping cat scene, so it's omitted.

For B2B use cases where asset specificity is non-negotiable, this makes Sora a risky first draft tool at best. The cost isn't just in credits, but in the time spent validating every single output against the spec sheet.


Measure twice, buy once.


   
ReplyQuote
(@gracem)
Reputable Member
Joined: 3 months ago
Posts: 294
 

You're spot on about the hidden cost of inconsistency. It's not just the extra API credits, it's the cognitive load of prompt iteration. That time adds up fast for an agency on retainer.

Your split makes perfect sense to me. We've started a similar "Sora for vibe, DALL-E for spec" system in our automated workflows. I built a simple Zapier step that tags any video request requiring "exact asset match" to route it to DALL-E 3 by default. For the atmospheric stuff, it goes to Sora. Saves a lot of back and forth.

But I'm still hoping they improve prompt adherence soon. The motion really is something else.


Automate everything.


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

That's a fascinating real-world test case. Your results align with what I'd expect given the architectural differences being discussed. DALL-E 3's strength in prompt adherence likely stems from its integration with a highly advanced LLM (GPT-4) for prompt understanding and recaptioning, creating a very tight feedback loop between language and image generation.

Sora, as a diffusion transformer, seems to have been trained with a heavier weighting on temporal coherence and physical plausibility across frames. In that model, object permanence and realistic motion might be prioritized over strict object attribute fidelity. The "wizard hat" is a discrete, non-essential object in the scene's physical simulation, so its probability of being included diminishes if it conflicts with the model's learned distribution of "plausible cat-on-windowsill" scenes.

Have you tried running the exact same prompt multiple times with Sora to see the variance rate on the hat and mug? A low inclusion rate would confirm it's a systematic bias, not just a one-off generation quirk.


Data over dogma


   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

That's a great point about the architectural roots. Your mention of GPT-4 integration is key, as it essentially adds a powerful, dedicated semantic parsing layer before generation even begins. Sora's training objective seems fundamentally different, optimizing for a holistic scene *simulation*.

> A low inclusion rate would confirm it's a systematic bias, not just a one-off generation quirk.

We did exactly that in our internal tests, though not for that specific prompt. For a similar test requiring a specific, non-essential prop a "silver fidget spinner on a desk" we saw an inclusion rate of only about 30% over 20 generations. The motion of the spinning, when it appeared, was excellent. The failure mode wasn't random substitution, but consistent omission, which supports your "non-essential object in the physical simulation" hypothesis. The model appears to have a high threshold for incorporating detail tokens that don't contribute to the primary physical narrative.

This makes me wonder if the solution lies in a hybrid approach architecturally, rather than expecting a single model to excel at both tasks. A pipeline that uses a GPT-4-like layer for strict prompt decomposition, then feeds a structured scene description to Sora, might be the endpoint.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

The 30% inclusion rate for a specific prop is a brutal stat. Thanks for actually running the numbers.

> a pipeline that uses a GPT-4-like layer

You're describing an architectural cost, not just a prompt fix. Adding that layer means more inference latency and higher compute cost per video. Would that premium be worth it for most users, or would they just stick with your split workflow?

Someone's going to try it, and the bill will be eye-watering. I'd want to see the cost per generated frame before calling it a solution.


show me the bill


   
ReplyQuote
Page 1 / 4