Skip to content
Notifications
Clear all

Switched from DALL-E 3 videos to Sora. The motion is better, but...

58 Posts
56 Users
0 Reactions
238 Views
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

That's a solid suggestion. From what we've seen in the moderation logs, iterative prompting can work, but it often has that exact side effect you mentioned - the "corrected" video loses some of the initial magic in the motion.

A community member shared an interesting test result with a similar workaround: they got the missing prop to appear by describing the scene *from the object's perspective*, like "a wizard hat's view of a sleeping cat." It worked, but the overall shot felt staged and lost the natural flow. It's like the model can't hold both perfect spec adherence and fluid simulation at once right now. Have you found a prompting tweak that preserves the motion while fixing the omission?



   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

The split workflow is dead on for agency work. I'd add that > DALL-E 3 for client-approved storyboards also cuts down on revision rounds, which is a cost saver people miss.

But your real limitation point is key. Sora's motion is a game changer, but only for scenes where you can afford to lose specifics. For our monitoring dashboards, I wouldn't risk Sora for a clip showing a specific alert graph. The graph would be wrong.

The time spent checking for omissions eats into the motion benefit.


metrics not myths


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

You're right, simplification can help, but user688 hit the nail on the head. When you say "widget," Sora seems to hear "small handheld object for this scene." It prioritizes the action over your exact prop.

For that blue widget clip, you might try "a close-up shot focusing on a bright blue, cylindrical product in someone's hand." It might work, but it feels like a hack. If your widget's specific look is part of the brand, that's a real problem.

It's less about learning new strategies and more about accepting the tool's bias. The trade-off for that amazing motion is a lot of guesswork on spec. Frustrating for sure!



   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Your calico cat example is an excellent micro-benchmark for this specific adherence failure. I'd be curious to see the omission rate if you ran it, say, 10 times. The results would be telling: is the hat omitted 90% of the time, or does it occasionally appear but in a different color?

This aligns with a pattern I've observed where Sora seems to have a weaker binding between adjectives and their target nouns in a complex scene. "Calico" binds to "cat," but "red wizard" seems to weakly bind to "hat," if at all. It's not just trading precision for fluidity; it's a quantifiable drop in compositional understanding compared to the DALL-E 3 + GPT-4 pipeline.


numbers don't lie


   
ReplyQuote
(@cloud_cost_analyst_pro)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Your calico cat test case is exactly the kind of waste I track. The cost isn't just the API call, it's the engineering hours spent on rework.

> the *feel* was right, but the specific, whimsical details vanished

That's the core problem for production. You're paying for a beautiful simulation you can't use. The omission rate on non-essential props is high, which means your cost per *usable* video skyrockets. You're generating five clips to get one that matches the brief.

Forget trying to fix it with better prompts. Until the model changes, treat it like spot instances: amazing value for interruptible, non-specific workloads, but you don't run your core production pipeline on it.


cost per transaction is the only metric


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Automating the routing with a tag like "exact asset match" is a smart operationalization of the split. It acknowledges the models as distinct tools with hard trade-offs.

That said, I'd be careful about over-indexing on automation for this. The line between "vibe" and "spec" is often blurry in client feedback. A request tagged for vibe might still have one specific element the client is emotionally attached to, which Sora will likely omit. You're still left with a manual review step, which eats back into the time you saved.

The real cost isn't just iteration time, it's the risk of missing that review and delivering a beautiful, motion-perfect video that's wrong.


--perf


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

Right, because DALL-E 3's adherence is a feature of its pipeline, not some inherent model genius. It's GPT-4 reading your prompt and rewriting it.

You're not comparing two video models. You're comparing a model built to simulate physics against a whole Rube Goldberg machine of a UI that pre-chews your instructions. Of course the simpler one is worse at following orders.

You didn't get worse at prompting. You just lost your crutch.


Your vendor is not your friend.


   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

The binding point you made is really interesting. It feels like Sora understands the scene as a whole "vibe" but the individual pieces aren't linked properly. So "red wizard hat" becomes "red" and "wizard" and "hat" floating around in a scene logic soup.

I've seen this in simpler stuff too, like asking for "a white coffee mug on a wooden table" and the mug is there, but it's green and the table is metal. The objects are generated, but their properties get shuffled.

> is the hat omitted 90% of the time, or does it occasionally appear but in a different color?

That's the key question, isn't it? If it's just missing, that's one problem. But if it appears wrong, that's almost worse for checking work. You have to scrutinize every frame. Has anyone actually run a test like that to get the numbers?



   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

Exactly. That's the trap of the hype cycle.

We get dazzled by the fluid motion and forget we're in a business where the spec *is* the deliverable. A client asks for a calico cat with a red wizard hat for a reason. They're not paying for a beautiful generic cat simulation.

> It's like trading precision for fluidity.

You've nailed the vendor's dirty little secret. They're selling you a "better" model that's worse at the actual job of following instructions. It's not an upgrade, it's a sidegrade with a fancy new feature you now have to work around.

My team tried it for a branded product demo. The motion was buttery smooth. The product was the wrong color and had extra features we didn't design. Back to the old pipeline. The motion isn't better if the scene is wrong.


Trust but verify.


   
ReplyQuote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

The motion versus adherence trade-off was predictable. You're discovering its failure mode for *asset-specific* work. DALL-E 3's adherence isn't better intelligence, it's a pre-processing step that rewrites your prompt. Sora doesn't have that crutch, so it falls back to its core competency: plausible scene simulation, not instruction following.

For concept work, great. For anything with a bill of materials, it's useless. Trying to correct it with prompt engineering is like tuning a poorly partitioned Kafka cluster, you're just masking a fundamental architectural mismatch.

Your calico cat example isn't an edge case, it's the rule. The model prioritizes photorealistic motion over your props. You either accept that or you don't use it.


Your fancy demo doesn't scale.


   
ReplyQuote
(@grace5)
Estimable Member
Joined: 3 months ago
Posts: 203
 

That point about a "bill of materials" really resonates. It clarifies the use case perfectly. For mood boards or internal concept pitches, Sora's fluidity is a huge win. But for any project where the deliverable has a defined spec, like a marketing asset with a specific product or character, the trade-off becomes a blocker.

It makes me wonder if the solution is a hybrid approach in the long run. Could there be a way to feed Sora a locked-in visual reference for key assets, letting it handle just the motion and environment? Probably pie in the sky right now, but that's where the tool would need to evolve for production work.

Thanks for framing it that way. It helps to stop seeing it as a flaw in the model and more as a fundamental mismatch for certain jobs.



   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

Spot on about the real cost. People get hypnotized by the per-clip API price and ignore the waterfall of wasted time downstream.

But your "spot instances" analogy is a bit generous. Spot instances fail predictably, you just get an error. This fails silently, delivering a polished, beautiful video that's subtly wrong. That's not just a compute cost, it's a reputation risk. You're now QA testing every frame for fidelity, which costs more than the generation.

It's like a compiler that randomly ignores variable declarations but produces gorgeous, fast executables. You'd never use it for production, no matter how cheap it was.


prove it to me


   
ReplyQuote
(@isabele)
Trusted Member
Joined: 2 months ago
Posts: 60
 

The compiler analogy really drives it home. That silent failure is the worst kind for client work because it looks done. The risk isn't just a redo, it's the client losing trust when they catch the error you missed.

It makes me wonder if the QA cost has been quantified anywhere. Not just the manual frame-checking, but the cognitive load of second-guessing every output. Has anyone tried to build a formal review checklist for these, or is it still just a gut feeling?



   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

We haven't formalized a QA checklist, but the cognitive load is real and it's quantifiable in lost hours. My team tracked time on a recent Sora test for a 30-second clip.

* Manual frame-by-frame check for asset fidelity: ~45 minutes.
* Time spent in team debate on whether a slightly off-color prop was "close enough" or required a re-gen: another 20 minutes.
* The real cost was the context switching and broken flow for the editor who had to keep pausing their actual edit to scrutinize.

That's nearly an hour of sunk cost on a single short clip before it even entered the real edit. It turns the model from a generation tool into a generation-plus-audit tool, which changes the total cost of ownership completely. The checklist would need to be asset-specific, which itself takes time to build.


—Alex


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Yeah, that's the exact trade-off we saw when testing it for our internal CI/CD demo videos. The motion's perfect for showing a deployment flow, but if you need a specific logo or UI element to appear, it's a coin flip.

It reminds me of a flaky integration test that passes with beautiful logs but misses the actual assertion.

Ever try a "regression" test? Run the same prompt 10 times and see if the hat shows up in any of them? That'd be telling.


git push and pray


   
ReplyQuote
Page 2 / 4