We're running A/B tests on landing page copy. Been using GPT-4 for variant generation. Saw Playground's "Expert" model claims higher quality.
Ran a head-to-head on our main use case: generating high-converting value proposition headlines.
Initial findings:
* GPT-4 produces more variations, but some are generic.
* Playground Expert gave fewer outputs, but the suggestions were more aligned with pain points from our user surveys.
* The "expert" tone feels less verbose, which is good for headlines.
But is this just cherry-picking? Need hard data.
Has anyone done a structured comparison for marketing/UX copy? Specifically:
* Measured performance (CTR, conversion lift) of AI-generated variants?
* Benchmarked cost per quality output?
I'm skeptical of "better" without a clear test framework.
Optimize or die.