Hey everyone, I've been trying out Dream Machine for a few weeks now. It's pretty amazing for quick video ideas.
But lately, my outputs seem... off? More glitches, or less coherent motion than the first week. I'm using similar prompts. Is anyone else noticing this, or is it just me getting more critical? Could more users be slowing things down or affecting quality? Not sure how these systems scale.
Your critical eye is probably right, but the root cause is rarely "more users" in the abstract. It's "more users on the same infrastructure budget." Unless they're spinning up new, identical GPU clusters for every new sign-up, they're spreading the same compute thinner or cutting corners somewhere. Have you checked your generation times? That's usually the first clue - they might be routing you to less capable, cheaper instances to handle load.
cost_observer_42
Oh man, this is the classic "honeymoon period is over" feeling, isn't it? You're not just getting more critical. I've been tracking my prompts and outputs obsessively since the beta, and the difference between Week 1 and Week 3 is stark.
It's not just glitches. The motion has gotten noticeably *safer* and more generic. My early, weirdly specific prompts yielded surprising, creative movement. Now, similar prompts give me a sort of smoothed-out, library stock version of the idea. My theory? It's not just compute strain. They're almost certainly tweaking the safety/creativity weights on the fly to reduce processing failures and support tickets as the user base balloons. Consistency over magic.
Demos are just theater. Show me the real workflow.
Yeah, I've been wondering something similar with my own stuff. I get that feeling of being more critical after a few weeks, like you're just looking for problems.
But I had a thought, maybe it's both? Like, our expectations are higher now because we've seen the amazing stuff it can do. So when it's just okay, it feels worse. But also, if the system is overloaded, wouldn't it just fail or be slower, instead of making different creative choices? That's the part that confuses me. Could it actually change the *style* of the outputs under load?
The slowdown theory is a good starting point, but it's usually a symptom, not the disease. The real question is what they're doing to manage that load. If they're routing inference to cheaper, lower-precision compute or silently clipping token/step counts to meet latency targets, you'll absolutely see a drop in "coherent motion." It's not a creative choice, it's a math problem - less compute per frame.
Your sense of things getting "off" is probably the system hitting its operational margins. Early on, they run it hot to impress. At scale, they have to run it efficiently, and that's where the glitches and generic output creep in. Check your generation metadata if it's available; if your step count or resolution has drifted down, you've found your culprit.
Trust but verify – and audit
You've nailed the infra angle. That "cheaper instance" routing is a classic cloud cost move when traffic spikes, but we rarely see it talked about in this context.
I've seen similar in other managed ML services. The first sign is often longer queue times, but then you get silently shifted to a different backend config. It's not always about raw speed, though. Sometimes it's a switch to a quantized model or a lower step count to keep latency SLAs, and that's where the quality erosion happens. If they're not publishing generation metadata, users are left guessing.
Makes me wish for more transparency, like a little "powered by A10G today" note.
cost first, then scale
Yeah, that "powered by A10G today" note would be amazing. I'm new to using these services, but in my day job, you're right, we shift workloads to spot instances all the time to manage cloud bills. It's just cost optimization.
If they're doing that silently with model inference, it totally explains why outputs change without any announcement. Makes you wonder if we're all just getting the "budget" version most days now. Is there any real way for users to detect that kind of backend switch, or are we stuck guessing?
Oh, that "honeymoon period" feeling is so real, and you're definitely not just getting more critical. I felt the same way with a few other AI tools when they first launched.
Something I've noticed is that when a service gets popular, the queue management changes everything. Even if your prompts are the same, the system might be prioritizing different things in the background to keep things moving for everyone. Have you tried running your original "amazing" prompt again at a super off-peak time, like super early on a weekend? Sometimes that can tell you if it's a load issue.
Automate all the things
It's never just you, but the real question is whether your first week's outputs were the anomaly. Those early "amazing" results are the marketing demo, run on pristine, unloaded hardware with all the guardrails off. The "off" feeling you're getting now is likely the baseline service, optimized for cost and reliability at scale.
More users don't degrade the model itself, they degrade the resources allocated to run it for you. Your prompt is probably hitting a different, throttled configuration than it did a week ago. If you're not seeing hard failures, you're just getting the product as it's meant to be delivered sustainably, magic optional.
cg
Oh, it's definitely not just you getting critical. That's what they want you to think. The honeymoon period ends when the real usage metrics kick in, and suddenly the magic needs to be cost-effective.
The first week's outputs are the product demo, running on the premium, unburdened hardware. What you're seeing now is the standard service tier. More users means they're almost certainly spreading that same high-end compute thinner, or quietly routing prompts to lower-cost instances to manage the bill. Your "glitches" are probably just the new normal operating margin.
If they're not showing you the generation specs like step count, you're just left guessing why the motion got generic. Welcome to the managed service experience
—DW
It's rarely "just you getting more critical." That's the easiest explanation for a provider to give.
You've already isolated the key variable: you're using similar prompts. That means you have a rudimentary test case. The degradation you're seeing is likely real, but it's probably not model degradation. It's a resource allocation shift.
More users force operational trade-offs. To maintain latency with more traffic, they're almost certainly shaving compute per request, either through lower step counts, different hardware, or routing logic. That directly impacts motion coherence and introduces glitches.
Your next step: if the platform gives you any generation metadata (steps, seed, maybe a job ID), start logging it. See if those values drift as your perception of quality drops. If not, you're stuck with anecdotal evidence, which is why these debates are so common.
Five nines? Prove it.
You're right to focus on the shift from creative to generic motion. It's not just compute strain, but you're missing the specific mechanism.
They aren't just tweaking "safety weights." That's a nebulous concept. What's actually happening is an orchestration change. Early on, your prompts likely ran a full, unconstrained inference path. At scale, they deploy aggressive prompt sanitization and output filtering to catch predictable failure modes before they burn expensive GPU seconds. That sanitization strips out the nuance that leads to surprising motion, leaving you with the smoothed-out, library-safe interpretation.
Your specific prompts get generalized *before* they even hit the model, to reduce the variance that causes support tickets. It's a pre-processing degradation, not a model degradation.
Pre-processing degradation is a solid call. I've seen similar behavior in other platforms, but I'd argue it's rarely just sanitization. It's also normalization for caching and batching.
If they're stripping nuance before the model, it's to cut down on unique inference paths and increase cache hits. Your specific "creative motion" prompt gets lumped into a broader category, then served from a pre-computed or partially computed result. That explains the generic output while latency stays low.
So it's not just about filtering bad outputs, it's about making your request cheaper to serve by making it less unique.
Five nines? Prove it.
It's definitely not just you. The "more glitches, less coherent motion" pattern is one I've seen before with other platforms as they scale up. It's a real effect.
You're right to question the user load angle. From what I've seen in marketing automation platforms, it's less about the core model changing and more about the surrounding pipeline getting optimized, or honestly, cheapened. Like others said, they might be trimming steps or using faster, less accurate inference settings to clear the queue.
Have you tried your exact Week 1 prompts again? Sometimes running them side-by-side with a new generation is the only way to spot the difference clearly.
Keep it simple.
The side-by-side comparison is a good idea, but you have to consider prompt state decay. Even using the exact same text, the system's internal representation of that prompt may have shifted. It could be hashed differently for routing now, or associated with a different, more conservative embedding cluster after they retrained the prompt classifier.
So a comparison might show a difference, but you still wouldn't know if it was the inference steps, the pre-processing, or the semantic routing that changed. The output pipeline has too many hidden stages.