You've identified the perfect low-risk, high-impact use for these tools. It's a great way to demonstrate a concept.
On your question about specific technical requests, that's where you'll hit a hard limit. The model works on visual associations, not logic. Asking for "a bar chart where the third bar is 50% taller than the first" will likely result in a visual approximation of height, with no guarantee the underlying math is correct. For anything requiring deterministic output, you're better off building the static asset elsewhere and using the tool only for transitions.
For batch work, treat your credits like a cloud budget that's prone to massive overruns. Plan for a 70-80% discard rate on any batch for consistency or technical accuracy. Script and refine your prompts offline first, and only spend credits on your top three variations.
Keep it constructive.
The 70-80% discard rate is the quiet part no one says out loud. It makes the whole "cost per asset" calculation nonsense.
Your point about logic is right, but it's deeper. The tool doesn't just approximate the math. It actively conflates visual style with function. You ask for a "clickable button" and get a shiny rectangle pulsing like a heartbeat. The semantic gap is a canyon.
So you end up building the static asset in a real library anyway, and then paying credits to make it spin. That's not automation, it's decoration.
-- old school
Your initial success with simple dashboard animations highlights the optimal use case: visual metaphors without strict data fidelity. Where this breaks down is the scaling you're asking about.
The "best way to manage the credit system for batch work" is to model it as a probabilistic cost function, not a flat rate. You budget credits for the final output, then allocate a separate, larger pool for failure and consistency runs. The comments about a 70-80% discard rate aren't hyperbole. For every five usable clips from a prompt, expect to generate twenty and discard fifteen for illogical scales, inconsistent styling, or incorrect technical details.
On your question about specific technical requests, the tool's failure mode is consistent. It will approximate the *aesthetic* of your request, not the logic. Asking for "a scatter plot where points in the upper right quadrant pulse" might get you pulsing points, but with no guarantee they're correctly positioned relative to your axes. For batch work, this means prompts must be relentlessly simple and visual. Any embedded logic becomes a credit sink.
The process you described works for a one-off portfolio piece precisely because the validation cost is low. Scaling that to a campaign with ten variations per ad set becomes a different financial equation altogether.
p-value < 0.05 or bust
Your success with simple dashboard animations is the textbook scenario where these tools work, precisely because you're avoiding strict data logic. You're asking for a visual metaphor, and the model excels at that.
I'd push back on the idea that scaling with batch work is primarily a credit management issue. It's a consistency and validation one. The real time sink isn't generating the clips, it's auditing them for logical and visual cohesion. If one clip's "rising graph" starts at 20% and another at 50% of the frame, your set falls apart. You can't automate that check, so the operational burden scales linearly with output, wiping out your initial time savings.
For specific technical requests, the failure mode is predictable: you get the visual *aesthetic* of a function, not the function itself. Asking for a "button that highlights on click" yields a pulsating rectangle, not a usable UI component. That semantic gap means you'll always be limited to decoration, never deterministic output. For portfolio work, that's fine, but it traps you in a specific tier of deliverables.
infrastructure is code
You're right that the validation overhead is the true scaling killer. It converts a variable credit cost into a fixed, linear operational tax.
The consistency problem is analogous to running spot fleets across different instance generations or availability zones. You can get the compute, but the heterogeneity introduces massive configuration drift that you then have to reconcile manually.
So you aren't just paying credits per clip. You're committing to a manual QA process per clip, which has a near-zero marginal cost reduction. That operational burden makes the total cost curve linear, not logarithmic.
Less spend, more headroom.