Skip to content
Notifications
Clear all

My results after a month: CTR up 2%, but is it causal?

3 Posts
3 Users
0 Reactions
6 Views
(@martech_trail_blazer)
Trusted Member
Joined: 4 months ago
Posts: 29
Topic starter   [#481]

After a rigorous 30-day trial of Anyword integrated into our HubSpot-powered content and email workflow, I have concrete performance data to share, alongside significant methodological reservations.

Our primary use case was generating and A/B testing subject lines and email body copy for a nurture stream targeting MQLs in the SaaS sector. We ran a controlled experiment: for four consecutive weekly nurture emails, we used our human-crafted copy (Control) against Anyword-generated variants (Variant A and B, using its "Brand Voice" trained on our top-performing historical assets). All other variables—segment, send time, design template—were held constant.

**The observed metrics were as follows:**

* **Overall Email Open Rate:** Increased by 1.1% (not statistically significant given our volume).
* **Overall Click-Through Rate (CTR):** Increased by 2.3% for the Anyword variants aggregate. This was statistically significant (p < 0.05).
* **Performance by Asset Type:** The uplift was not uniform. Subject line variants showed negligible improvement. The CTR gain was almost entirely driven by the call-to-action phrasing and introductory sentence variants within the email body.

**The Critical Question of Causality:**

While the 2.3% CTR lift is a positive output, attributing it directly to Anyword's "AI predictive performance score" requires a substantial leap of faith. My concerns are as follows:

1. **Confounding Variable - Novelty Effect:** The AI-generated copy often had a slightly different syntactic structure than our team's writing. It is plausible that this novelty alone, not inherent "better" quality, caused the incremental engagement. Will this effect decay over time as recipients become accustomed to the style?
2. **Tool vs. Process Improvement:** The act of systematically generating and A/B testing *more* variants itself likely contributed to the win. Would a junior copywriter mandated to produce three variants per email have achieved a similar result? The tool enforces a discipline our process previously lacked.
3. **Metric Selection Bias:** We optimized for CTR, which Anyword predicted well. However, preliminary data shows a 5% lower conversion rate on the landing page from these clicks. This suggests the AI may be optimizing for clicks using slightly more generic or curiosity-driven phrasing, which attracts less-qualified traffic. The business impact is therefore ambiguous.

**Operational and Data Quality Notes:**

* The HubSpot integration via API is stable but requires careful mapping of custom properties for brand voice training.
* The "Brand Voice" training is only as good as the assets you feed it. We had to cleanse our historical data of several outdated campaigns to prevent the AI from learning and replicating obsolete messaging.
* The time-saving claim is partially valid for high-volume, mid-funnel content. It did not replace the need for strategic messaging sessions or deep product-focused copy.

In conclusion, Anyword functioned as a potent ideation and testing accelerator, correlating with a CTR increase. However, I cannot, with confidence, state the lift is *causal*. It may be a combination of enforced testing rigor and temporary novelty. The ROI calculation must factor in the subscription cost against the *net* value of the incremental clicks, considering potential downstream conversion dilution. I am continuing the experiment for another quarter to monitor for novelty decay and to better measure downstream impact.



   
Quote
(@migration_mike_33)
Eminent Member
Joined: 2 months ago
Posts: 23
 

Interesting breakdown. Your point about the CTR lift being driven by CTA and intro phrasing, not subject lines, really jumps out. It suggests the tool's predictive strength might be in micro-copy and value-proposition articulation within the body, rather than the initial hook.

That fits with a pattern I've seen in other platforms. The "brand voice" training often optimizes for proven, middle-of-the-funnel persuasive language it finds in your historical wins. For subject lines, which are more about curiosity and urgency, the dataset might be noisier or the AI's interpretations less effective.

A solid next step, if you continue the trial, would be to isolate and A/B test *just* the CTA block from the AI against your control. That could confirm if the causal mechanism is truly there, or if it was interaction with other elements in the body copy.


test the migration before you migrate


   
ReplyQuote
(@cloud_sec_enthusiast)
Estimable Member
Joined: 2 months ago
Posts: 90
 

Really solid experimental design, isolating those variables is key. That CTR jump is interesting, but I've seen similar patterns where the "significant" result gets misinterpreted because of a lurking variable.

With email campaigns, a change in click-through rate can sometimes be caused by the new copy inadvertently aligning with a recent security or privacy update from a platform like HubSpot. For example, if the AI's phrasing slightly altered the link structure or anchor text, it might bypass a new spam filter rule or look more "authentic" to a client-side security plugin that was recently updated. It's not that the copy was more persuasive, it just tripped a different technical wire.

Might be worth checking if there were any HubSpot email deliverability updates or major browser changes rolled out during your test window. It's a long shot, but in cloud security, we see "causal" links break down on these technical nuances all the time.


security by default


   
ReplyQuote