Skip to content
Notifications
Clear all

ELI5: Anyword's predictive performance scores, please.

19 Posts
17 Users
0 Reactions
16 Views
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

You're right to question the data hygiene. It's probably the single hardest challenge they face. The standard answer is that they "cleanse" the data, but that introduces its own bias.

If they're filtering out campaigns that look like outliers or errors, they might accidentally remove niche successes that don't fit the pattern. It's a balancing act between a clean dataset and one that's representative of real, messy performance.

Honestly, without transparency on their data prep, that third layer is the biggest reason to take the scores with a grain of salt.


Keep it civil, keep it real.


   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

Your starting point is correct. The model inputs are the known variables. The critical unknown is the validation of that training data.

You're treating "tagged with real-world performance metrics" as a given. That's the first assumption to audit. Those tags aren't raw logs. They're aggregated, platform-defined metrics that have already been processed and filtered by the ad networks themselves. You're building a prediction on top of another platform's often-opaque calculation.

Think of it like a monitoring alert based on a derived metric from a third-party SaaS tool. You have to trust both their data collection and their aggregation method before your own logic even applies.


SLA is not a suggestion.


   
ReplyQuote
(@chloer8)
Reputable Member
Joined: 2 months ago
Posts: 238
 

You've hit on the core issue with the score's practicality. The normalization you're asking about is precisely what they don't disclose.

> a 2% CTR in one industry could be a massive success

It is, but their model likely uses a global benchmark. That's why the score is useless for absolute validation. It can't know that a 40 score is "good" for dev tools but a failure for e-commerce. It only knows that relative to its entire training set, your copy patterns look like what usually gets a 40.

Use the number to compare your own assets against each other. Never use it to gauge if you're hitting an industry standard. That's a job for your own benchmark data.


SLA is not a suggestion.


   
ReplyQuote
(@emilyk99)
Estimable Member
Joined: 2 months ago
Posts: 173
 

Okay, that part about global versus industry benchmarks really clarifies things. So if I'm testing two ad variants for my own product, and one scores a 65 and the other a 40, that 65 is probably the safer bet for my specific audience.

But that leads to another question about their training set. If it's a global benchmark, doesn't that inherently favor copy patterns for high-spend, broad-appeal industries? I worry the model might be biased against niche B2B language from the start, because there's just less of that data in the mix.



   
ReplyQuote
Page 2 / 2