Yeah, the monitoring system analogy really helped me wrap my head around it, because I was totally lost on that term too. So thanks for asking this.
It sounds like, from what others are saying, you can't really trace it back to a single input because it's a composite score. That's kind of a bummer for someone like me who's used to seeing clear metrics. I guess my follow-up would be, if the baseline is a moving target, how do you even start to trust it for a new project? Do you just ignore the number and only use it to rank your own ideas against each other?
Exactly. You ignore the absolute number.
> how do you even start to trust it for a new project?
You don't. "Trust" is the wrong framework. You treat it as a black-box ranking engine, not a calibrated metric. Your process is:
1. Run all your candidate copy through their API.
2. Take the top-scoring outputs.
3. Test *those* in your own controlled A/B environment against a real baseline you own (like your current best-performing subject line).
The only validation happens in your own systems. Their score just narrows the field for your real test. If the "winning" variant from their ranking consistently loses your actual A/B tests, you stop using the tool.
Least privilege is not a suggestion.
> what are the agency's financial incentives?
Spot on. Their incentive is to keep you renewing the subscription. That means the score must feel novel and "insightful," which often leads to model drift that chases quarterly marketing fads, not sustainable metrics.
So your long-term customer value gets optimized out of the equation. You're aligning with a vendor's retention KPI, not your own business goals.
Keep it simple
I really like your monitoring system analogy, it's a great way to think about it from a technical mindset. You've got the right hunch about historical copy and metrics being the training data.
The frustrating truth is exactly what others have said: you can't see the query logs. Their 'primary label' is a proprietary blend of engagement signals they've cooked up. It's less of a clean regression on CTR and more of an ensemble model trained on whatever composite metric they've decided defines 'performance' in their walled garden.
Where your analogy perfectly extends is in the variance question. You're used to a stable system where a latency spike has a root cause. With these scores, the baseline *itself* is the moving part. A score of 85 might be in the 99th percentile of their corpus this month, but after a retrain on new data that optimizes for a different stakeholder goal (as user184 pointed out), that same text could fall to the 70th percentile. The variance isn't just in your copy, it's in their shifting definition of the target.
don't spam bro
You've nailed the core instability. That shifting baseline is why benchmarking these tools over time is essential.
You'd need to periodically feed the same seed copy into their system and chart the score's movement. If you see significant variance on identical inputs, it confirms the model's target metric has drifted. This turns their black box into a useful signal, not of copy quality, but of the vendor's changing priorities.
BenchMark