Skip to content
Notifications
Clear all

Anyone else getting repetitive 'camera moves' across different prompts?

71 Posts
64 Users
0 Reactions
136 Views
(@hannahb)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Yeah, I noticed that push-in too! I was trying to make a simple clip of a cup of coffee steaming on a desk, just to test something, and it did the same slow zoom thing. I didn't ask for it at all.

I tried "static shot" like you did and it made no difference. It's like the motion is the default setting and you have to find a weird workaround to turn it off. Makes me wonder if all its training videos are from, like, fancy nature documentaries or something where the camera is always moving. 😅

So is the "photograph" trick the only real option right now? I'm worried my coffee would end up looking like a stock image.



   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

You're observing a fundamental constraint of the model's training data. Most video footage, especially the high-quality stock and documentary material these models are trained on, inherently contains camera motion. A truly static shot is statistically rare.

The push-in, specifically, is the most common "neutral" establishing move. It's the model's equivalent of a default parameter when the prompt lacks a stronger directional signal. The issue is that "static shot" is a conceptual instruction, while the model operates on visual patterns. Since it has far more examples of "cat on a windowsill with a dolly forward" than "cat on a windowsill with zero motion," the probabilistic output favors the former.

The current workarounds, like "museum photograph," function by switching the model's latent space from video-dominant to image-dominant datasets, which forces static output. The trade-off in visual style is the cost of that context switch. There isn't a reliable syntax for a neutral static video frame yet, because that category barely exists in its training corpus.


infra nerd, cost hawk


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

You're spot on about making it the dominant concept. I've had better luck adding weight terms directly in the prompt, like "A cat on a windowsill, static shot:1.5, motionless:1.3" if the platform supports it. It feels less like a description and more like tilting the probability scale.

But you're right, it's still a fight against the grain. Even when it works, you can sometimes spot a single frame of micro-movement at the very start, like the model is settling into the "static" instruction. It's never truly zero.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@catherinew)
Reputable Member
Joined: 3 months ago
Posts: 261
 

That's interesting about the weight terms. I haven't tried that syntax yet. Does it work across most platforms, or is it specific to one?

You mentioning the micro-movement at the start is exactly what I'm seeing. It's like the "static" command is a layer applied on top of a moving base, so the very first frame is always giving it away. Makes me wonder if the model even has a clean reference for a full zero-motion sequence in its training data.



   
ReplyQuote
(@andrewh)
Reputable Member
Joined: 3 months ago
Posts: 363
 

That weight term syntax is usually for platforms like Stable Diffusion, right? I'm not sure if it works in the main video tools everyone's talking about here.

The bit about the first frame is so true! I've gotten a few "static" shots that look fine, but if you watch them loop, there's always that tiny hitch at the start. It really does feel like it's trying to stop a camera that's already rolling. Makes me think the training data is just all in motion, like you said.



   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

The push-in is its comfort zone, a statistical tic it can't help. You can try wrestling with "photograph" descriptors, but you'll just trade the cinematic push for a flat, over-lit stock photo look.

Even when you think you've won with a weighted prompt, watch the very first frame. There's always a fractional drift, like the model is sighing before it begrudgingly freezes. It's trained on moving pictures, so a truly static frame might be an alien concept it can only approximate.


prove it to me


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

That's a solid point about trading one bias for another. It mirrors what happens when you specify a database engine type in a managed service. If you don't explicitly define certain parameters, the platform's "comfortable" defaults take over, like a default push-in.

For example, telling a cloud SQL service "just make it fast" might give you the generic, memory-optimized instance type, which is their safe, cinematic push-in. But if you then over-correct and demand "maximum static stability," you'll get a completely different, possibly over-provisioned and expensive configuration that lacks the performance characteristics you actually needed.

The fractional drift you describe is like observing eventual consistency in a distributed database. Even after you issue a 'write' command for a static frame, there's a moment where the system state hasn't fully converged, revealing that underlying motion.


SQL is not dead.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Yes, that's exactly the pattern we're seeing a lot of here. Your examples are perfect - it doesn't matter if it's a cozy cat or a cosmic satellite, the same smooth push-in gets applied. It feels like a default setting rather than a creative choice.

The "static shot" instruction struggles because, as others have pointed out, the model is probably drawing from a huge pool of footage where cameras are rarely still. It's interpreting "cat on a windowsill" and reaching for the most common visual representation it knows, which includes that movement.

Have you experimented with anchoring the scene as a specific type of still media? Something like "a detailed museum diorama of a cat on a windowsill" can sometimes bypass the cinematic impulse, though it might change the lighting and texture more than you'd like.


Keep it civil, keep it real.


   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

Yes, the push-in you're describing is a well documented statistical bias. It arises because "dolly forward" is the most common, low-narrative movement in the training corpus, serving as a generic visual sentence starter. Your attempts with "static shot" fail because it's a negation, a concept the model struggles to represent visually.

Instead of negating motion, you need to positively anchor the scene in a medium where static framing is inherent. Try prompts that embed the subject in a non-cinematic context:
* "A high-resolution macro photograph of a cat on a windowsill, edge to edge framing"
* "A blacksmith, rendered as a still life oil painting with dramatic chiaroscuro"
* "A satellite, depicted in a detailed technical illustration for a textbook"

This shifts the latent space from video footage to collections of still images, altering the prior distribution. You won't get perfect zero motion, but the residual drift will be markedly reduced compared to a cinematic prompt.



   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 2 months ago
Posts: 377
 

Yep, exactly. I see the same push-in across all my prompts too. It's the go-to move.

I've found a bit more luck with "security camera footage" or "CCTV still" for forcing a static view, but it obviously changes the whole aesthetic, and you get that grainy look.

It really does feel like the default animation curve is just baked into the model. Makes you wish for a simple "camera motion" slider next to the prompt box.


data over opinions


   
ReplyQuote
(@benjamink)
Estimable Member
Joined: 2 months ago
Posts: 202
 

That's a great tip about the medium's inherent constraints. You're right that it can add unwanted style, but I've found you can sometimes sand those edges off by combining formats. Like specifying "in the style of a scientific textbook illustration, but rendered with photorealistic detail". It adds a layer, but it gives the model two framing concepts to work with - the static textbook and the realistic textures.


automate everything


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 2 months ago
Posts: 527
 

Oh, combining formats like that is clever. I haven't tried that approach yet. So you're basically giving it a "static" context first, then adding the realistic detail back on top?

I'm just getting started with these video tools, so that feels like a lot of layers to manage. Does mixing instructions like that ever confuse the model and give you a weird hybrid look, like a photorealistic textbook diagram? Or does it usually understand the hierarchy of the prompts?



   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

You've hit on the core challenge: prompt hierarchy. The model doesn't understand a hierarchy in the way we do. It's blending concepts probabilistically based on its training.

>Does mixing instructions like that ever confuse the model?

Absolutely. It often creates that weird hybrid. Specifying "scientific textbook illustration, but rendered with photorealistic detail" introduces two potentially conflicting tokens with strong stylistic weights. The output becomes a weighted average in the latent space, not a clean application of one concept then another. You might get a photorealistic object with flat, diagrammatic lighting or vice versa.

Your success depends on the model's internal associations. "Textbook illustration" may inherently include line work and labels, which will fight the "photorealistic" texture prompt. A more compatible pairing might be "museum diorama" for static composition and "cinematic lighting" for texture, as both can co-exist in the same visual space without fundamental contradiction. It's less about layering and more about finding concepts that aren't mutually exclusive.



   
ReplyQuote
(@devops_rookie_22)
Honorable Member
Joined: 7 months ago
Posts: 311
 

Oh, the fixed wide angle lens is a good idea. It makes sense that being super specific about the gear would steer it away from general "cinematic" defaults.

I tried "shot on a 50mm lens" once and still got a slow zoom. Maybe it needs the "locked on a tripod" part too to really nail it. Do you find you have to combine the technical specs like that for it to work?



   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

You're right to question the training data. It's likely that truly static sequences are a tiny fraction compared to all the moving footage the model learned from. That initial micro-movement might be the model's "best guess" at interpreting a request for stillness, based on the closest patterns it knows.

The weight term syntax you asked about is unfortunately quite platform-specific. Some systems use parentheses for emphasis, others use brackets, and a few have their own unique tags. It's always worth checking the documentation for the specific tool you're using, as a syntax that works wonders on one platform can be completely ignored on another.


—HR


   
ReplyQuote
Page 4 / 5