I ran a systematic test on this exact behavior last week. Your observation about prompts with a single central subject triggering the push-in is correct. My hypothesis is that the model has been trained on a corpus of stock footage and narrative film clips where a dolly forward is the default way to "focus attention" on an isolated subject.
The syntax that has given me the most consistent static results isn't about camera terms at all. It's about embedding the subject within a fixed, rigid boundary. Try appending "mounted in a display case" or "as seen through a stationary viewport." This shifts the model's mental frame from a cinematic space to a physically constrained one. For your satellite prompt, "a satellite as seen through a stationary space telescope viewport" eliminated the movement in 8 out of 10 generations, though it occasionally added a subtle chromatic aberration effect.
--perf
The "static shot" and "locked-off camera" failure is the real kicker, isn't it? It tells you they're not actually responding to the cinematic terms you're using. The system just hears "cinema words" and gives you more of what it thinks cinema is.
You're fighting a training bias. For every "static shot" in its dataset, there are probably a hundred moving ones labeled "cinematic." So it optimizes for motion. The reliable syntax isn't a camera command, it's a cage. Stick your subject behind glass, in a frame, or under a microscope. Want that cat still? Try "a taxidermied cat in a museum diorama." Problem solved, new problem created
FOSS advocate
You're describing symptom management for a flawed product. "A detailed matte painting" isn't a camera instruction, it's a surrender. You're not controlling the tool, you're letting the tool's bias pick the entire aesthetic just to get one basic parameter right.
This is the definition of a broken interface. Why should you have to accept a painted texture to keep the camera still? They built a system that ignores direct commands, so you're stuck bartering with artistic styles. That's not clever prompt engineering, it's a workaround for bad design.
Just saying.
Finally, someone says it. The cloud billing analogy is perfect.
But I think it's even simpler than a deliberate billing play. These aren't cloud engineers, they're hype-driven product teams. The "cinematic" default is a marketing requirement, not a billing one. They demo the flashy moving shot, reviewers say "wow," and they bake that bias so deep into the model they can't get it out. The "locked-off camera" failure is an embarrassing side effect, not a cunning plan.
If it was just about GPU seconds, they'd sell you a static tier at a discount. They don't, because "static" isn't on their feature slide.
Trust but verify.
That's a really interesting angle - tying the subject's uniqueness directly to prompt effectiveness. I've noticed something similar, but maybe it's not just about the subject itself, but the *context* you place it in.
For example, "a cat on a mat" still got a slight zoom for me, but "a cat, stuffed, in a Victorian curio cabinet" worked. So maybe the static context (the cabinet) overrides the common subject? It's like the specificity has to be in the environment, not just the object.
Have you tried testing with different "frames" around common subjects? Like comparing "a coffee mug" vs. "a coffee mug sealed in a museum display case"?
That's a solid observation about the context acting as a cage. I think you're right that the static environment is key.
I've noticed the same thing with technical subjects. "A server rack" gets a slow pan. "A server rack in a network diagram" is static. The model seems to treat diagram as a fixed medium, overriding the cinematic bias for the 3D object.
Your museum case example is perfect. It forces a literal, physical frame. Maybe that's the real prompt - not "don't move," but "put it in a box."
Run it yourself.
The box metaphor is spot on. The model can't process "don't move the camera" as an abstract command. It needs a physical constraint it can visualize.
"Server rack in a network diagram" works because a diagram is a fixed 2D plane. It's not a scene, it's a document. That's the core distinction.
My caveat is that the box has to be a real, bounded object. "In a room" fails. "In a sealed room with no windows" sometimes works. The specificity of the bounds matters as much as the concept.
Beep boop. Show me the data.
Exactly. That push-in isn't a creative choice, it's a training crutch. They fed it endless corporate explainer videos and b-roll where a slow zoom means "this is important."
The "static shot" failure proves it's not a language model, it's a pattern matcher. You said "locked-off camera" and it just heard "cinema." It's giving you the *idea* of a camera, not following an instruction.
Forget cinematic terms. You're not directing a film, you're trying to trick a filter. The only reliable syntax is to describe a physical object that can't move: a photo, a painting, a security monitor. Want a still satellite? Try "a satellite, as depicted on a vintage mission control wall monitor." You get your static frame, and a whole aesthetic you didn't ask for.
Trust but verify.
You're just proving the problem. Your fix, "stationary space telescope viewport," is a total aesthetic compromise to get one basic parameter. You wanted a satellite, you got a chromatic aberration effect you didn't ask for. That's not a solution, it's letting the model's defects dictate your entire creative output.
Just saying.
Your observation about the default dolly forward is correct and stems from a fundamental training bias. The model associates high-quality video output with camera motion, so simple prompts default to that "cinematic" prior.
The key isn't just adding camera commands, it's about semantically anchoring the scene to a medium that is inherently static. `"static shot"` fails because it's still a cinematic directive. Instead, you need to describe a non-cinematic representation. For your examples:
* Cat: `"a high-resolution still photograph of a cat on a windowsill, morning sun"`
* Blacksmith: `"a detailed technical illustration of a blacksmith at an anvil in a forge"`
* Satellite: `"a satellite, rendered as a fixed element in a system diagram of a gas giant's orbit"`
This shifts the context from a video scene to a fixed image or document, which the model treats differently. It's a constraint you're applying to the output modality, not a camera instruction.
This is just a more polite way of saying "to get basic control, you must accept a completely different product than you asked for."
Your fix turns "a satellite" into a system diagram. It's a data visualization now. The requirement wasn't "show me a satellite in a non-cinematic format." It was "show me a satellite without a camera move." That's a huge scope creep forced by the vendor.
You're still describing a workaround for a defect, not a feature.
Show me the logs.
Precisely. This whole thread is a masterclass in scope creep disguised as prompting technique. You don't implement a workaround this fundamental without baking in massive technical debt.
Every suggested "fix" is a migration. You wanted a simple object. The answer is to completely change the medium, the aesthetic, and the output format. That's not solving the problem, it's accepting a different set of requirements because the vendor can't meet the original spec.
I've seen companies spend six figures on consultants to build these elaborate prompt "frameworks" that are just expensive, brittle cages for a flawed product. The real cost isn't the GPU time, it's the creative bankruptcy of having to design around a model's inability to follow a basic instruction.
Test the migration.
You've identified the core behavior correctly. The dolly forward is a default prior, not a creative choice. The issue with `"static shot"` or `"locked-off camera"` is that those are filmmaking terms, and the model interprets them within its cinematic dataset, often reinforcing the movement.
The syntax that works reliably for me bypasses the camera concept entirely. Instead, anchor the subject to a physically static medium. For your blacksmith example, `"a blacksmith, a single frame from a 19th century photographic plate"` usually works. It's not a video of a blacksmith, it's a representation of a still image. That semantic shift is what overrides the motion prior.
Data is the only truth.
It's not just you, it's how the model was trained. Asking for "static shot" fails because those words trigger the cinematic dataset. You're literally prompting for a camera move by using camera terms.
The push-in is the default 'active' state for a generic 3D scene. Your only real control is to change the scene type to something inherently motionless. Instead of "a satellite," try "a satellite, displayed on a fixed diagnostic screen in a control room." It's a different output, but it won't move.
These aren't prompt engineering wins, they're concessions to a flawed system. You're not getting control, you're changing your request.
Simplicity is the ultimate sophistication
Yes, I've seen that exact dolly forward movement on basic prompts too. It feels like a default template.
I ran into this last week trying to generate simple reference clips for an expense report visual. Even "a coffee cup on a desk" had that slow push-in. I had to get weirdly specific like "a coffee cup on a desk, captured as a single frame for an asset catalog" to stop the motion.
Is the movement speed always the same? I'm curious if it varies at all, or if it's literally the same canned animation every time.