I've been putting Anyword through its paces for the last three weeks, specifically the Predictive Performance Score feature, as part of a content pipeline load test. The core promise—a consistent, data-backed score to predict engagement—is exactly what my team needs for automated quality gates.
However, I'm seeing a critical failure in consistency. The score is exhibiting extreme variance with minimal, semantically neutral changes to the input. This isn't a slight drift; it's a swing that would arbitrarily pass or fail content in an automated system. My benchmark tests are showing unacceptable standard deviation.
Here's a concrete example from my test log. All prompts were run sequentially within a 5-minute window, using the same "Facebook Ad" template and target audience settings.
**Test Case: "Performance Score for 'Get started today'"**
- Input 1: `Get started today`
- Result: Score **82**
- Input 2: `Get started today.`
- Result: Score **71** (Added a period)
- Input 3: `Get started today!`
- Result: Score **89** (Added an exclamation)
- Input 4: `Get started right now`
- Result: Score **64** (Synonym swap: "right now" for "today")
The delta between the highest (89) and lowest (64) is 25 points. The difference between "Get started today" and "Get started today." is 11 points. This is not noise; this is a fundamental reliability issue.
I've ruled out basic issues:
- No A/B testing features were enabled.
- Target audience parameters were locked.
- Each was a fresh generation, not an edit of the previous.
This behavior makes the score unusable for any automated workflow. If the model is this sensitive to punctuation alone, its internal scoring mechanism is either massively overfitting to trivial patterns or is non-deterministic in a way that isn't documented.
My questions for the community and Anyword team:
1. Has anyone else replicated this level of volatility in the Performance Score?
2. Is there a documented tolerance or confidence interval for the score? (e.g., ±5 points)
3. What are the actual, technical inputs to the score? Is it purely on the generated text string, or does it include hidden metadata like generation timestamps or internal sequence IDs that would cause this drift?
Until this is resolved, I cannot recommend relying on the Performance Score for any mission-critical filtering. The benchmark data simply doesn't support its stability.
—DL
Benchmarks or bust