I’ve been trying Sora for a few days, and half my prompts get a “I can't generate videos yet” or something vague. When it does work, the output is often not what I pictured.
I’m used to writing clear specs for software vendors or SQL queries, but this feels different. What’s the step-by-step method for a prompt that actually gets followed? I’m thinking: subject, action, style, framing. But what concrete details are most important to include, and what should you leave out?
not a buyer, just a nerd
I think you're on the right track with subject, action, style, framing. But I've found Sora really struggles if you don't anchor the "physics" or "camera" more explicitly. For example, I tried "a cat jumps off a table in slow motion" and it kept giving me a cat floating backwards. What worked was adding "realistic motion blur, 50mm lens, cat lands on wooden floor."
What kind of "not what I pictured" are you getting? Is it the movement, the lighting, or the composition? I'm wondering if leaving out a specific camera angle or lighting direction is what causes the vague failures.
Oh wow, that example about the cat jumping is super helpful. I never would have thought to specify the landing spot. It makes sense though, like you're giving it the full story.
So if adding "cat lands on wooden floor" fixes the physics, does that mean we should always define the *end point* of an action? Like for "a person opening a door," should you say "their hand turns the knob and the door swings inward into a kitchen"?
The camera details are a great point, but I get nervous about adding too many technical terms. I don't know much about lenses. Could saying something simple like "shot from a low angle" or "bright morning light from a window" work just as well?
If you're used to writing specs for vendors, you're already ahead of the game. The core problem is you're still writing for a human who understands context. This thing doesn't.
Your subject, action, style, framing framework fails because it assumes the tool can infer a coherent physical world. It can't. You have to spec the world itself, not just the elements. Think less like a director and more like a paranoid QA tester writing a bug report that lists every single state change. "Cat jumps" is meaningless. "Cat, starting from a standing position on a table, pushes off with hind legs, arcs through air with limbs tucked, and lands on a wooden floor with front paws first, causing a slight shake in its body" might work, until it doesn't.
You leave out nothing. You over-specify everything. And then you pay for the privilege of doing the model's job for it. The most important concrete detail is your exit strategy when you realize the total cost isn't just the subscription, but the hours spent learning to appease a black box.
Buyer beware.
You're thinking about this like a database query. It's not. It's more like yelling instructions at a drunk contractor who's already halfway through building the wrong house.
> clear specs for software vendors
Those work because humans share a model of reality. This thing doesn't. "Subject, action, style" is useless if the AI hallucinates the entire scene's physics.
Forget what to leave out. You can't over-specify. The most important detail is the one it randomly decides to ignore. The "step-by-step method" is trial and error, same as tuning a k8s cluster. Welcome to the party.