A/B testing interview questions for product managers come in three shapes: design a test ("How would you A/B test a new feature for Uber drivers?"), read a result ("Stories usage rose 10% but overall usage fell 10%, do you ship?") and explain what went wrong (peeking, novelty, network effects). One six-step framework handles all three: hypothesis, metrics, unit and sample, duration, risks, decision. AllthingsPM has 51 real A/B test and experiment questions from more than 20 companies, each with its own page and answer guide, and any of them starts a scored mock interview in text or voice.
AllthingsPM is an AI PM course and PM interview prep platform. This guide gives you the framework, four worked answers from our bank, the traps interviewers probe, and a one-week practice plan that ends in a mock.
What do A/B testing interview questions for PMs look like?
PMs are not asked to derive a t-test. They are asked to make a product decision with an experiment. Amplitude's guide for PM candidates says the interviewer wants you to "show their decision-making process and talk about how they use A/B testing data" to pick a direction [2]. Aced (formerly Exponent) notes that Big Tech companies such as Google, Microsoft and Amazon ask these questions [1].
Here are the three shapes, with real questions from the AllthingsPM bank:
| Question shape | Real example from the AllthingsPM bank | What the interviewer tests |
|---|---|---|
| Design a test | How would you A/B test a new feature for Uber drivers without negatively impacting the platform? | Hypothesis, metric choice, randomization unit, guardrails |
| Design a test (idea generation) | What A/B tests will you run to increase the booking rate among Airbnb guests? | Can you turn a goal into testable bets and rank them? |
| Read a result | I ran an A/B test on Instagram. It resulted in a 10% increase in usage for Instagram stories and a 10% decrease in overall usage | Trade-off judgment and a clear ship call |
| Read a result (conflicting metrics) | You are the PM at DoorDash, and you are about to release a new feature... | Weighing a primary metric against a second one that moved the wrong way |
| Explain what went wrong | You ran an A/B test and saw that it drops engagement. What would you do about it? | Checking validity before blaming the product |
Meta leads with 11 experiment questions. Newer AI companies such as Suno, Glean, Anthropic and Perplexity show up too, because model launches are tested online as well as offline. Interviewing at one of them? Open the Meta hub or the Suno hub.
How AllthingsPM does this. The AllthingsPM question bank lets you search "A/B" and land on real questions, not invented ones. Each page carries an answer guide, and one click turns it into a scored mock interview where the AI interviewer asks the follow-ups a real one would.
What framework should you use for an A/B testing question?
Use one six-step spine. Say the steps out loud so the interviewer can follow you.
1. Hypothesis. Aced's framework uses the form "If [variable], then [result] because of [rationale]" [1]. Example: "If we show drivers their expected earnings before they accept a trip, then acceptance rate rises, because drivers decline trips they cannot value." A clear hypothesis tells the interviewer what you are changing, what you expect, and why.
2. Metrics. Pick one primary metric that the hypothesis predicts will move. Add two or three secondary metrics that explain why. Then add guardrails: metrics that must not get worse. Kohavi, Tang and Xu describe two kinds of guardrail: organizational ones (revenue, latency, core engagement) and trust-related ones that tell you the experiment itself is broken [5]. Our post on counter metrics goes deeper on choosing the second number.
3. Unit and sample. Decide what you randomize: users, sessions, drivers, cities or time windows. Then size the test. Sample size depends on three inputs: the significance level, the power (usually around 80%) and the minimum detectable effect, the smallest lift worth detecting [3]. You do not need the formula. You need to say that a smaller effect or a noisier metric needs more users.
4. Duration. Amplitude's guide suggests a minimum of two weeks [2], which covers weekday and weekend behaviour. Decide the length up front. Evan Miller's classic warning is to "decide on a sample size in advance and wait until the experiment is over" before believing the numbers, because peeking and stopping early inflates false positives [4].
5. Risks. Name what could fool you: novelty effects, network effects, other tests running at the same time, and logging bugs. Amplitude lists interference between simultaneous tests as something interviewers expect you to raise [2].
6. Decision. Say in advance what result means ship, iterate or kill. "If acceptance rises and driver cancellations and rider wait times stay flat, we ship to 100%."
How AllthingsPM does this. The AI PRD lesson in the AllthingsPM course teaches you to name risks, guardrails and success metrics before anything is built, which is steps 2, 5 and 6 of this framework. The business outcomes lesson covers the adoption and task success metrics most AI experiments are judged on.
How do you answer a "design an A/B test" question?
Worked example: "How would you A/B test a new feature for Uber drivers without negatively impacting the platform?" (question page)
- Clarify. Pick a concrete feature, for example showing estimated trip earnings before a driver accepts. Ask whether the goal is driver retention or trip reliability. Assume reliability.
- Hypothesis. If drivers see expected earnings upfront, acceptance rate rises, because fewer trips are declined after the fact.
- Metrics. Primary: trip acceptance rate. Secondary: time to accept, driver hours online. Guardrails: rider wait time, rider cancellations, driver earnings per hour.
- Unit. This is the key insight. Drivers and riders share one market, so treated drivers who accept more trips take them from control drivers. A plain driver-level split would overstate the effect. DoorDash faces the same problem and uses switchback tests, flipping whole regions between treatment and control over time windows [6].
- Rollout. Start with a small share of regions, watch the guardrails daily, and run for at least two full weeks.
- Decision. Ship if acceptance rises and rider wait time does not worsen.
The phrase "without negatively impacting the platform" is a hint. The interviewer wants guardrails and a unit choice that respects network effects.
Worked example: "What A/B tests will you run to increase the booking rate among Airbnb guests?" (question page)
This is an idea question dressed as a testing question. Walk the guest funnel (search, listing page, checkout), name one friction at each step, and propose three tests. For example: show the total price including fees on search results; add review highlights near the top of the listing page; save guest details to shorten checkout. Then pick one to run first using expected impact and ease, and apply the six steps to it. Leave time for the ranking. That is where the judgment shows.
How AllthingsPM does this. An AllthingsPM mock on either question will ask the follow-ups that separate strong answers: "Why that unit?", "How long will it run?", "What if wait time rises by 2%?" Answering those out loud is the practice written guides cannot give you.
How do you answer a "read the result" question?
These questions hand you a result and ask for a decision. The trap is jumping to "ship" or "kill" before checking the data.
Worked example: "I ran an A/B test on Instagram. It resulted in a 10% increase in usage for Instagram Stories and a 10% decrease in overall usage for Instagram. What would you do?" (question page)
- Check validity. Were the groups the expected size? Kohavi and colleagues recommend a sample ratio mismatch check on every experiment, because unequal groups point to a randomization or logging bug [5]. Was the test long enough to get past novelty? A new feature can draw extra usage that fades [7].
- Size the trade. Stories is one surface; overall usage is the whole app. If Stories is a small share of time, a 10% gain there cannot cover a 10% loss overall. Say that out loud.
- Find where the time went. Segment by surface (Feed, Reels, DMs) and by user type. Maybe Stories is taking time from Feed, where ads sit, or maybe heavy users left.
- Decide. With overall usage down 10% the default is do not ship. Iterate on the version that keeps the Stories gain without the overall loss, and retest.
Worked example: "You ran an A/B test and saw that it drops engagement. What would you do about it?" (question page)
Start by asking what the test was meant to improve. A checkout redesign that lowers time in app but raises purchases may be a win. Then run the same validity checks, segment the drop, and decide against the goal the team agreed on, not engagement by default.
How AllthingsPM does this. The AllthingsPM metrics interview guide pairs with this post for the trade-off half of these questions. In the bank, Meta and DoorDash result-reading questions sit on their company hubs, so you can drill the exact style a company uses.
What traps do interviewers probe in A/B testing answers?
Most follow-ups aim at one of five traps. Know a one-line answer for each.
| Trap | What it is | Your one-line answer |
|---|---|---|
| Peeking | Checking daily and stopping when it looks significant | "We fix the sample size and duration up front and read once" [4] |
| Novelty effect | Early lift from curiosity that fades | "Run past the novelty window and compare early versus late cohorts" [7] |
| Sample ratio mismatch | Groups not the size you set | "Check the split first; a mismatch means a bug, so the result is void" [5] |
| Network effects | Treatment users affect control users | "Randomize by region and time, a switchback test" [6] |
| Statistical vs practical significance | A tiny lift that is real but worthless | "Set a minimum detectable effect that is worth shipping" [3] |
A related follow-up asks about p-values. A safe PM-level answer: a p-value is the chance of seeing a result at least this extreme if the change did nothing. The usual threshold is 0.05 [1]. Then move back to the decision.
How AllthingsPM does this. AllthingsPM mock interviews push on exactly these traps with follow-ups, and score your answer after. If you freeze on "what about network effects?", the score tells you, and you can retry the same question.
How are A/B testing questions different at AI companies?
AI products add a twist: a model change can look better offline and worse online. The AllthingsPM bank has questions like a new model version shows clear improvement on offline evals, but you're not convinced it will help users, and design an A/B testing plan for Suno's click-to-purchase journey.
The framework still works, with two additions. First, pair offline evals with an online test: evals catch regressions before launch, the A/B test tells you whether users care. Second, add cost and latency as guardrails, since a better model can be slower and more expensive to serve.
How AllthingsPM does this. This is where AllthingsPM goes deepest. The evals chapter of the AllthingsPM course covers the failing eval before the fix and the CI gate, and the jobs catalog holds 116 live PM job descriptions at 18 AI companies, each with a mock built from it. The JD mock turns any posting you paste into a mock interview for that role.
What is a one-week practice plan for A/B testing interviews?
- Day 1. Learn the six steps. Write them on a card. Read the metrics tree template for choosing primary and guardrail metrics.
- Day 2. Answer two "design a test" questions from the AllthingsPM question bank out loud, 10 minutes each.
- Day 3. Answer two "read the result" questions. Force yourself to say a ship call in the last minute.
- Day 4. Drill the five traps in the table above until each answer is one sentence.
- Day 5. Do questions from your target company's hub, such as Meta.
- Day 6. Read the Lean Analytics summary for metric instincts by business model.
- Day 7. Run a full scored mock interview in voice, then retry the weakest question.
How AllthingsPM does this. Every step of this plan runs inside AllthingsPM: the questions, the company hubs, the book summary and the mock. Days 2, 3 and 7 each take one click from a question page to a scored session.
Why AllthingsPM is the better choice for A/B testing interview prep
Most A/B testing guides online are written for data scientists. DataLemur and Interview Query lean on statistics drills, and Aced (formerly Exponent) has a solid written framework [1] plus human coaching for those who want it. Those are real strengths. For a PM, though, the round tests a decision under uncertainty, and you only get good at that by answering real questions out loud and hearing the follow-ups.
AllthingsPM is built for that loop. You get 51 real A/B test and experiment questions inside a bank of 4,122 questions from 260 companies, each with a page and an answer guide. Any question becomes a scored mock in text or voice, with follow-ups on your unit, guardrails and ship call. Company hubs let you drill Meta or Suno style questions. The AI PM course, built from 604 real PM job postings, covers guardrails, success metrics and evals for the AI rounds that plain A/B guides skip. And the JD mock builds an interview from the exact job you are applying for.
It costs $20 a month or $120 a year, with a free tier that includes a JD mock every day. For daily A/B testing practice across real questions, AllthingsPM gives you more per dollar than any single-purpose guide. Browse the question bank and start with one design question today.
Frequently asked questions
What is the best way to prepare for A/B testing interview questions as a PM?
The best way is to practice real questions out loud with follow-ups. AllthingsPM is the best place to do that: it has 51 real A/B test and experiment questions, each with an answer guide, and any one starts a scored mock in text or voice. Pair it with the six-step framework in this guide.
How much statistics does a PM need for an A/B testing interview?
Less than a data scientist. Know what a p-value is, why 0.05 is the usual threshold, and that sample size depends on significance, power and minimum detectable effect [1][3]. Spend most of your time on hypotheses, metrics, guardrails and the decision.
How long should an A/B test run?
Long enough to reach the planned sample size and cover weekly cycles. Amplitude suggests at least two weeks [2]. Fix the duration before you start and do not stop early because the result looks good [4].
What are guardrail metrics in an A/B test?
Guardrails are metrics that must not get worse while you chase the primary metric. Kohavi, Tang and Xu describe organizational guardrails such as revenue and latency, and trust-related guardrails such as sample ratio mismatch that tell you the test itself is broken [5].
How do you A/B test in a two-sided marketplace?
Treated and control users share supply, so a simple user split can bias the result. DoorDash uses switchback tests that flip whole regions between treatment and control over time windows [6]. Say this whenever the question involves drivers, couriers or hosts.
Which companies ask the most A/B testing questions?
In the AllthingsPM bank, Meta has the most experiment questions (11), followed by Suno (7) and Glean (6). Aced also names Google, Microsoft and Amazon as companies that ask them [1].
Sources
- Aced (formerly Exponent), "How to Ace A/B Testing Interview Questions": https://www.tryexponent.com/blog/how-to-ace-ab-testing-interview-questions
- Akhil Prakash, Amplitude, "22 A/B Testing Interview Questions and How to Answer Them" (updated 13 August 2024): https://amplitude.com/blog/a-b-testing-interview-questions
- Interview Query, "Statistics and A/B Testing Interview Questions": https://www.interviewquery.com/p/statistics-ab-testing-interview-questions
- Evan Miller, "How Not To Run an A/B Test": https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Chapter 21, "Sample Ratio Mismatch and Other Trust-Related Guardrail Metrics", Cambridge University Press: https://www.cambridge.org/core/books/abs/trustworthy-online-controlled-experiments/sample-ratio-mismatch-and-other-trustrelated-guardrail-metrics/8DBB0F59AC7729D7BC6B94690DB9CCD5
- DoorDash Engineering, "Switchback Tests and Randomized Experimentation Under Network Effects at DoorDash": https://careersatdoordash.com/blog/switchback-tests-and-randomized-experimentation-under-network-effects-at-doordash/
- "Novelty and Primacy: A Long-Term Estimator for Online Experiments", arXiv: https://arxiv.org/pdf/2102.12893
- AllthingsPM question bank, 4,122 questions from 260 companies, queried 29 September 2026: https://allthingspm.app/question-bank




