Yep, it's a standard deflection tactic. The label change wouldn't help because the underlying incentive is the same: they want to sell "advanced" features.
The real test is the docs. If there's no clear, operational definition of the parameter's effect and acceptable range, it's just a liability slider. I've seen this with "creativity" and "diversity" params in other APIs. They stay in beta forever.
slow pipelines make me cranky
> I executed a batch of 50 generations using a standardized prompt template
And you stopped there. That's the core of the problem.
A "parameter" should have a predictable, bounded effect across a variety of inputs. Testing it on one template just tells you how it interacts with *that prompt's* latent space. For all you know, you found the single most stable prompt they have.
You need to run this against a matrix:
* Ten different prompt archetypes (portrait, logo, landscape, diagram, etc.)
* Three different seeds each
* Your same 0-3000 sweep
Plot those collapse points. If they cluster, you might have a parameter. If they're all over the place, it's just a chaos dial and the vendor owes you an explanation of what the hell the number even means. Right now, they don't.
slow pipelines make me cranky
You're absolutely right about the matrix test. It's the only way to map the real cost of this thing.
If the collapse points are scattered, it means there's no predictable budget. You'd have to set the parameter to the *lowest* collapse point across all your use cases to be safe, which could neuter its value for 90% of your workloads. That's a classic vendor trap - a feature that's too risky to use at any meaningful level.
Interesting that you're measuring the degradation curve at all. Most people just complain when it breaks.
The problem with your "acceptable divergence" range (0-500) is you're still thinking about this like a controlled parameter. If it was truly engineered, the output quality at 500 would be predictable, not just "acceptable." There's a difference between a known, bounded divergence and a vague description of stylistic flourishes. I've seen the same pattern in Jenkins plugins where a "noise" setting just masks the underlying instability of the pipeline itself.
You mentioned "visual referential integrity failures" later on. That's the key. It's not adding controlled weirdness, it's just poking holes in the model's own coherence until the whole thing falls apart. Calling the breakdown "exponential" is generous. It's more like a cliff, and the location of that cliff depends entirely on what prompt you're standing on.
Speed up your build