Skip to content
Notifications
Clear all

Thoughts on the new 'improved physics' claim? My test says no.

2 Posts
2 Users
0 Reactions
4 Views
(@ethanp)
Estimable Member
Joined: 1 week ago
Posts: 86
Topic starter   [#13960]

The recent update to Sora’s model card prominently highlights "improved physics and object permanence" as a key advancement. As someone who regularly stress-tests these systems for procedural consistency, I felt compelled to design a controlled evaluation. The claim, while exciting, warrants scrutiny beyond anecdotal evidence.

My methodology was straightforward: I generated a series of prompts centered on basic Newtonian mechanics and object interactions in simple scenes. For instance, "A marble rolls down a curved ramp, launches off a ledge, and arcs through the air before landing on a table." The critical frames were the launch trajectory and the impact. In multiple generations, the arc often defied a realistic parabolic path, with the marble appearing to drift or correct its course mid-air. More telling was the behavior upon landing; the impact frequently lacked a convincing transfer of momentum, with the marble sometimes appearing to stick or slide without the expected scatter or bounce.

Furthermore, object permanence tests—such as having an object pass behind a narrow occluder—still resulted in occasional morphing or dissolution. The model seems to prioritize frame-to-frame aesthetic coherence over strict physical adherence. This isn't to dismiss the improvement entirely; there is a subjective sense of "better" compared to previous versions. However, "improved" is not synonymous with "solved" or "reliable."

This leads to my central concern for this community: the terminology used in release notes. "Improved physics" sets a particular expectation that, in my testing, the model does not yet meet in a consistent, deterministic way. It risks user frustration when building workflows that depend on these properties. I believe our discussions here should focus on defining more granular, testable benchmarks for such claims—perhaps a shared suite of prompt templates we can all run to quantify progress.

I’m keen to hear if others have conducted similar systematic tests. Have you found specific scenarios where the physics *do* hold up reliably, or conversely, where they break down in novel ways? Comparing notes will help us all understand the practical, current boundaries of the tool.

— EthanP


Let's keep it constructive


   
Quote
(@ci_cd_crusader)
Reputable Member
Joined: 1 month ago
Posts: 139
 

Your test methodology is sound. I'd be curious to see if these physics inconsistencies correlate with the complexity of the scene's texturing or lighting. Sometimes the model allocates its parametric "budget" to visual fidelity at the expense of motion integrity.

From a systems perspective, this reminds me of testing a deployment pipeline where the logs claim success but the artifact checksums don't match. The high-level claim is made, but the underlying deterministic process isn't fully reliable.

Have you tried scripting the prompt generation and frame sampling? It would help isolate whether the errors are random or systemic to certain prompt structures.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote