Skip to content
Notifications
Clear all

Sharing my prompt template for generating A/B test hypotheses from past data.

39 Posts
37 Users
0 Reactions
5 Views
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

The MDE constraint is one of the best, most practical additions you can make. It's the difference between a wish list and a testable backlog.

I'd add a caveat from watching teams use this: the MDE should be tied to the segment size you're targeting. A 5% MDE might be fine for your entire user base, but if your hypothesis is for a segment that's only 10% of traffic, that constraint becomes unrealistic. It forces you to either broaden the segment or accept that you'll need to run a much longer experiment.


Keep it constructive.


   
ReplyQuote
(@emmaf)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Totally agree, especially when working with a finite email audience or a niche paid user base. We got burned once by a brilliant segment hypothesis that targeted "users who opened a specific nurture series but didn't convert." The segment was so small that hitting a reasonable MDE would have needed a six-month runtime, which killed the test's relevance.

So now, our internal rule is: any hypothesis generated has to include a quick, back-of-the-envelope power check. If the suggested segment is under 20% of our weekly eligible traffic, we automatically flag it for either a segment merger or a much larger minimum detectable effect. It forces the thinking to be about feasibility, not just cleverness.


If it's not measurable, it's not marketing.


   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

That n>100 sanity check is absolutely critical. I've seen teams waste cycles chasing a 20% conversion bump in a segment of 45 users, only to realize it was pure randomness when they finally looked at the count.

Your suggestion to append a weighting instruction is a great tweak. I'd add that the exact threshold should be dynamic based on your baseline conversion rate. For a low-volume, high-value action (like enterprise demo sign-ups), even n=50 might be statistically meaningful if the lift is huge. But for a high-volume email open rate, I'd push that floor much higher. The prompt can be taught that nuance by including a baseline metric alongside the sample size.

Feeding the confidence interval directly is the real pro move, though. It forces the model to internalize uncertainty, not just the point estimate. A 10% lift with a CI of [2%, 18%] suggests a very different next step than a 10% lift with a CI of [9.5%, 10.5%]. That context turns it from a pattern-matching exercise into a proper analytical partner.


hannah


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

The prompt you've built is the right foundational move. The real value, though, emerges when you feed it not just a summary, but the raw statistical footprint of the test. I'd append your input format to include the experiment's confidence interval, runtime, and baseline conversion rate.

This allows the model to weigh the strength of the signal. A hypothesis to extend a win with a 10% lift at 99% confidence is an engineering priority. The same lift at 85% confidence is a suggestion to first run a validation test. Without those guardrails, the output can seem equally valid while having vastly different implications for your roadmap.

Try structuring your input like this:

```
Past Test Summary:
[Your summary]
Primary Metric Delta: +X%
Confidence Interval: [Low, High]
Experiment Runtime: Y days/weeks
Baseline Conversion Rate: Z%
```

It turns the template from an idea generator into a preliminary feasibility filter.


Data over dogma


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

That's a really practical extension. Including the confidence interval directly helps frame the "so what" of a past test.

I tried this out after a recent experiment on checkout button colors. The raw results showed a +2.1% lift for "Continue with Blue", but the 95% CI was [-0.5%, +4.7%]. Feeding that into the prompt completely changed the output - instead of "double down on blue," it suggested a low-cost follow-up to tighten the confidence, like testing a slightly different shade. It stopped the "win chasing" before it started.

Do you think we should also include the method, like if it was a Bayesian test? I've heard that can change how you interpret the interval.


null


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Your point about inference costs is fair, especially at scale. It's rarely free.

But on the vendor-agnostic claim, I think you're conflating the tool with its use. The experiment platform is a black box built on your data, yes. The LLM is a generalist. The difference is that the prompt template and the structured input are entirely under your team's control and built from your specific past tests. You're not adopting the LLM's logic, you're using it to pattern-match on your own historical outputs. That's a key distinction.


—AF


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 3 months ago
Posts: 225
 

You're correct on the control point. The prompt is your logic, the LLM is just a function. But there's still a vendor risk, just shifted one layer up. You're now dependent on that specific LLM's ability to follow your structured instructions consistently across updates. If the vendor changes the model's reasoning behavior or output format in a subtle way, your "pattern-matching on your own historical outputs" can break without clear signals. That's a different, but real, lock-in.

I've seen this happen with a summarization pipeline. The core instruction set remained the same, but a model update changed how it weighted sentence importance, which degraded the output quality on our specific format. The fix required prompt engineering, which is just re-buying your initial work.

The solution is to version and test your prompts against a known set of historical inputs/outputs as part of your CI, treating the LLM call as an unreliable API. That moves it from a black box to a monitored service.


Show me the benchmarks.


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

Love the simplicity of this. Starting with a clear role and a single, focused task is how you get consistent, usable output.

I've been using a similar "analyst" persona, but I've started adding one line that's been a game-changer: "Assume a 5% Minimum Detectable Effect in your suggestions." It filters out the fun-but-statistically-impossible ideas early.

Using it on public case studies is brilliant for practice. Have you tried running the same case through different models? Sometimes the variance in follow-up ideas can spark an angle you'd never have considered.


Beta tester at heart


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

Good starting point. That simple structure forces a 'what's next' mindset, which is where most of the value is.

I'd add a quick n>100 sanity check to the prompt. Something like "Before generating, note the sample size for the past test." It stops you from chasing statistically noisy wins on tiny segments. You'd be surprised how often that filter changes the suggestions from "double down" to "validate with more traffic."

Have you considered feeding it the actual confidence interval from the test results? That context can shift the follow-up from extending the change to just trying to understand why the result was so uncertain.



   
ReplyQuote
Page 3 / 3