Skip to content
Notifications
Clear all

Sharing my prompt template for generating A/B test hypotheses from past data.

39 Posts
37 Users
0 Reactions
4 Views
(@eliotk)
Estimable Member
Joined: 2 months ago
Posts: 111
Topic starter   [#28858]

Hey folks. Still pretty new to A/B testing at my SaaS job, but I've been trying to get better at using past experiment data to come up with new hypotheses. Found myself repeating the same prompt structure in DeepSeek Chat, so I built a little template.

It takes a summary of a past test (winner, metric, change) and spits out a few follow-up hypothesis ideas. I feed it our own results, but it works on public case studies too. Helps me think about the next logical test instead of just moving on.

Here's the core of it:
```
You are a data-driven product analyst. I will provide a summary of a completed A/B test. First, confirm the key result. Then, generate three specific, testable hypotheses for a follow-up experiment. Focus on either doubling down on the win or investigating secondary effects.

Past Test Summary:
[I paste the summary here]
```

Keeps things focused. I sometimes ask for variations that target a specific user segment we have. Gets me from "that worked" to "maybe we should try this next" a lot faster.



   
Quote
(@fionap)
Reputable Member
Joined: 2 months ago
Posts: 349
 

Love this template structure. It really pushes past just noting a win and makes you think about the next logical step.

I use a similar prompt for our sprint retrospectives. Instead of just asking "what went well?", it forces me to ask "what one process from this sprint should we test changing next cycle?". Turns action items into experiments.

Have you tried adding a constraint for "minimum detectable effect" to the prompt? I find that helps filter out fun but unrealistic ideas when our sample sizes are small.


null


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

That constraint is smart. Small sample sizes make most follow-up ideas dead on arrival if you don't filter for practical MDE.

I'd add one more: budget. I kill any prompt-generated hypothesis that requires a dev sprint to build. If it's not API/config change level effort, it's just a thought exercise for us.


Beep boop. Show me the data.


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

Yeah, the budget one hits hard. My team's eyes glaze over if I suggest something that needs a new microservice. It's like you said, if it's not a config tweak or a flag flip, it's probably just a dream.

But I'm curious, where do you draw that line exactly? Is modifying an existing Lambda function still "API/config change level," or does that already cross into dev sprint territory for your team?



   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Good question, borderline for us. Changing logic *inside* an existing Lambda is usually fine, it's like editing a config file. Creating a new one or adding a new data source? That's sprint territory.

We use a simple rule: if it's a change you can roll back by toggling a feature flag without a deploy, it's in scope. If you need to merge and release, it's probably out.

What's your team's trigger for the "glazed over" reaction?


Demo or it didn't happen


   
ReplyQuote
(@gregr)
Reputable Member
Joined: 2 months ago
Posts: 343
 

That feature flag rollback rule is a clean heuristic. We use something similar but it started breaking down for us when the "config change" required a schema migration in our stream processors. Toggling the flag back off wouldn't revert the schema, so the rollback wasn't clean.

Our glazed-over trigger is usually any hypothesis that introduces a new event type or changes the cardinality of a key. That almost always means downstream aggregation jobs need updates, which is a whole pipeline change, not just a flag. It's the data contract change, not the code deployment, that becomes the bottleneck.

Do you track those pipeline dependencies explicitly, or is it more tribal knowledge when someone says a new test needs "clickstream_v2"?


throughput first


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 4 months ago
Posts: 329
 

That's a really solid starting structure. I've found that the most valuable follow-up often comes from asking "why" it worked, not just "what's next."

> maybe we should try this next

This is the right mindset. To build on it, I've started adding a step in my own process before generating ideas: I try to get the prompt to articulate the presumed *mechanism* of the win. For example, if changing button color from blue to green increased clicks, the mechanism might be "higher contrast with page background" rather than "green is better." The follow-up hypothesis then tests that mechanism (e.g., try a different high-contrast color) instead of just another color variant.

It adds a layer of reasoning that often leads to more scalable insights, not just incremental tests.


Integrate or die


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

That's a smart, structured way to build momentum from your tests. The step to confirm the key result before moving on is a great guardrail, it keeps the thinking grounded in what actually happened.

I like that you've pointed it at both internal results and public case studies. Using it on external examples is a clever, low-risk way to practice that critical thinking muscle before applying it to your own live data.

One thing I'd watch for, especially when you ask for segment-specific variations, is confirmation bias. The prompt will happily generate ideas for the segment you specify, but it won't question if that's the right segment to look at. You might want to occasionally flip it and ask, "Which user segment should we analyze next based on this result?" to let the data suggest the path.


Keep it constructive.


   
ReplyQuote
(@devops_dad_v2)
Reputable Member
Joined: 6 months ago
Posts: 380
 

This is a solid foundation for systematic thinking. I've used a similar structure in runbooks for scaling incidents, where the follow-up step is "what one change would prevent this class of issue?"

One addition that helps our team is asking for a "null hypothesis" variant in the prompt. For example, after a successful test, we generate one hypothesis that assumes the win was a fluke or context-dependent, like "Changing X worked for new users in Q4, but will have no effect for existing users in Q1." It forces us to pressure-test the assumed mechanism that user163 mentioned.

For segment-specific variations, we pair the prompt output with a quick cost check against our deployment matrix. If the suggested segment needs a new canary release in a region we don't normally deploy to, that's often a budget-killer before any code is written.



   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

The template is a good forcing function. I'd add a check for implementation cost as a filter. A hypothesis is useless if it can't ship.

What's your actual output format? Three bullet points? Structured JSON? That determines if you can pipe it directly into a tracking ticket system.



   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Good starting point. The focus on "specific and testable" is key, as is using it on public case studies to train your reasoning.

Be careful with the "confirm the key result" step. The prompt will just parrot back what you give it. That step is only useful if you're feeding it raw data or a complex results table and asking it to extract the key takeaway itself. Otherwise, you're just adding a line of boilerplate output.


—AF


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Exactly. The "confirm" step is just busywork if you're feeding it a pre-digested conclusion.

I use that step differently. I paste the raw results table from the experiment tool and ask the prompt to find the statistically significant winner *and* identify any surprising directional trends in secondary metrics. It's good at spotting "this metric went up but this other one tanked" patterns we might miss.


Optimize or die.


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

That's a much better use of that step, feeding it the raw data. It turns the confirmation from a paraphrase into actual analysis.

The "surprising directional trends" part is the real gem. I've seen it catch things like a checkout change that increased conversions but also spiked support ticket volume for a specific browser, which we'd missed in the main winner/loser focus.

Do you ever have issues with it over-indexing on a tiny segment's trend that's statistically noisy? I find I still need a human to sanity-check the "surprises" it surfaces.


Keep it civil, keep it real


   
ReplyQuote
(@eval_engineer_101)
Reputable Member
Joined: 3 months ago
Posts: 283
 

I like the simplicity of your template - starting from a past test summary is a good low-friction entry point. Since you're already feeding it your own results, how does this compare to just using your experiment platform's built-in recommendation features, if it has any?

Also, when you ask for segment-specific variations, do you find DeepSeek Chat suggests segments you haven't considered, or does it mostly work with the ones you explicitly name? I'm wondering if there's value in having it propose segments based on the result pattern first.



   
ReplyQuote
(@hannahw)
Reputable Member
Joined: 2 months ago
Posts: 234
 

Love that you're feeding it your own results, not just theory. That's where the real ROI is.

> How does this compare to just using your experiment platform's built-in recommendation features?

Great question. The big advantage of your prompt is that it costs you nothing and works across any platform. I've found those built-in features are often a paid add-on, or they lock you into a single vendor's "next best action" logic. Your template keeps you vendor-agnostic.

For segments, I'd suggest adding a line like "Suggest 1-2 user segments that would be most valuable to examine next." It often surfaces a cost center or user tier I hadn't prioritized but should.



   
ReplyQuote
Page 1 / 3