← Back to Prompt Library

📈 Data & Analysis

A/B Test Result Interpreter

Paste your experiment numbers and get a sober read: what the result supports, what could explain it besides your change, and whether to ship.

The Prompt — replace [BRACKETS] with your details

Act as an experimentation analyst who is skeptical by default. Interpret my A/B test.

The experiment:
- What I changed: [DESCRIBE VARIANT B VERSUS CONTROL A]
- Hypothesis stated before the test: [WHAT I EXPECTED AND WHY — say "none written" if you did not write one]
- Primary metric: [METRIC AND HOW IT IS DEFINED]
- Guardrail metrics: [e.g., refunds, support tickets, page load]

The numbers:
- A: [VISITORS] visitors, [CONVERSIONS] conversions
- B: [VISITORS] visitors, [CONVERSIONS] conversions
- Duration: [DAYS], run from [START DATE] to [END DATE]
- Traffic source mix and anything unusual during the test: [PROMOTIONS, OUTAGES, SEASONALITY, PRESS]

Deliver:
1. The observed rates and the absolute and relative difference, computed step by step so I can check the arithmetic
2. Whether this sample size can support a conclusion at all, and roughly what effect size it could detect — show your reasoning
3. Four alternative explanations besides my change: novelty effect, sample ratio mismatch, segment mix shift, and the most likely one specific to my context
4. What I should check in the raw data before believing the result
5. A recommendation: ship, kill, or keep running — with the specific condition that would change your answer
6. If I did not write a hypothesis first, say plainly how that weakens the conclusion

Rules: do not declare a winner from a small difference. If the honest answer is "this test cannot tell you", say that first.

How to use this prompt

  • Include the unusual events during the test window — a promo email is the single most common cause of a fake winner.
  • Check the arithmetic in section 1 yourself; models make careless mistakes with percentages.
  • Treat "keep running" as a real answer rather than a failure to decide.

Why this prompt works

The instinct after a test is to explain why you won. Asking for alternative explanations before the recommendation reverses that order, and requiring the model to show its arithmetic and its power reasoning makes both checkable rather than asserted.

Variations to try

  • Add revenue per visitor and ask it to interpret both metrics together, including the case where they disagree.
  • Ask for a pre-registration template so your next test has a written hypothesis and a stopping rule.
  • Paste three past tests and ask what pattern of mistakes shows up across them.

Common mistakes to avoid

  • Stopping the test the day it looks significant. That single habit invalidates most amateur experiments.
  • Trusting a model's p-value arithmetic without checking it in a calculator or a real stats tool.
  • Ignoring guardrail metrics because the primary metric went up.

Works well with

ChatGPT
Claude
Gemini

Need a custom version of this prompt?

The free prompt generator builds a prompt tailored to your exact goal, framework, and target AI model — or paste this template into the optimizer to refine it.