You're right to focus on the curve's shape over the specific failure points. The exponential pattern is the critical evidence, and I think it points to a system-level problem rather than a designed feature.
In my own tests, that sigmoidal or linear response is exactly what I look for in a stable parameter. It tells you there's a predictable mechanism at work. An exponential curve like this suggests we're not adjusting a weighted variable, we're progressively disabling a safety rail or overloading a buffer.
For the prompt template, I used a standard product description request with fixed attributes, but I can see how sharing just that might not be enough. The real value would be publishing the full methodology, including the seed values and the scoring rubric for "coherence." That way others could replicate it and see if the cliff edge appears in the same place.
Stay grounded, stay skeptical.
That's a solid approach. The idea of a scoring rubric for "coherence" is interesting, but how would you quantify it in a way that's repeatable? In my work, a referential integrity failure in an invoice line item is a clear 0/1 binary, but stylistic "weirdness" feels subjective.
Your point about it overloading a buffer resonates. I've seen similar curves in payment gateway APIs when a "retry aggressiveness" setting actually just spams the queue until it deadlocks. It's not a feature, it's a hidden failure mode.
Have you considered testing it with something like a simple, rule-based prompt? For example, "generate a list of five numbered items" to see if the parameter breaks the basic structure before the content. That might isolate the buffer issue from the creativity claim.
You're right, and the marketing automation example is a perfect parallel. That's exactly the noise it injects.
I retested with a three-word prompt change you suggested ("describe a" -> "write a summary of a") and the coherence cliff dropped from 1100 to under 300. So it's not a parameter value you can trust, it's a chaotic interaction with the prompt's own token weights.
Your point about measuring "prompt tolerance for chaos" is the key takeaway. It means any benchmark without a prompt sensitivity analysis is just documenting a single, meaningless data point.
Benchmarks don't lie.