Here are 25 real A/B testing interview questions for product managers, grouped into the five shapes interviewers use: design a test, generate test ideas, read a messy result, pick metrics and check validity, and run experiments on AI products. Every question links to its own page in the AllthingsPM question bank, with an answer guide and a button that turns it into a scored mock interview in text or voice. AllthingsPM is an AI PM course and PM interview prep platform. Its bank holds 4,122 real questions from 260 companies, and 51 of them are about A/B tests or experiments.
This is a practice list. For the step by step method, read the A/B testing interview guide first.
What are the 25 A/B testing interview questions?
| Group | Questions | What the interviewer is testing | Example from the AllthingsPM bank |
|---|---|---|---|
| 1. Design a test | 1 to 5 | Hypothesis, unit of randomization, primary and guardrail metrics | A/B test a new feature for Uber drivers |
| 2. Generate test ideas | 6 to 10 | Turning a goal into ranked, testable bets | A/B tests to raise Airbnb booking rate |
| 3. Read a messy result | 11 to 15 | A clear ship, iterate or kill call when metrics disagree | DoorDash orders up, restaurants down |
| 4. Metrics and validity | 16 to 20 | Knowing which numbers move and whether to trust them | Facebook newsfeed test on a 1% sample |
| 5. Experiments on AI products | 21 to 25 | Offline evals versus online tests, non-deterministic output | Experiment for a generative AI feature |
Group 1: How do you answer "design an A/B test" questions?
These questions hand you a feature and ask how you would prove it works. Amplitude's guide for candidates says interviewers want to see "your decision-making process" and how you use test data to pick a direction [2]. So the answer is a plan with a decision at the end, not a statistics lecture.
1. How would you A/B test a new feature for Uber drivers without negatively impacting the platform? Good answers include: why a simple driver split can leak (drivers in test and control compete for the same riders), a switchback or geo design to contain that, and guardrails on rider wait time and cancellations. DoorDash has written publicly about using switchback tests for exactly this marketplace problem [6].
2. Amazon is testing a new voice-shopping feature. How would you design an A/B test to validate whether it increases purchase frequency? Good answers include: a hypothesis in one sentence, purchases per user per week as the primary metric, return rate and order errors as guardrails, and a test long enough to cover repeat purchase cycles.
3. Rooms on Facebook was launched 3 months ago. Would you also launch it on Instagram? How would you run an A/B test to see the success of this feature on Instagram? Good answers include: a go or no go view first, then why a social feature needs cluster randomization (friends must be in the same arm to use it together), and a success metric about rooms joined, not rooms created.
4. Design an A/B experiment to launch a PayPal like donation product at a large non profit church, and complete the launch within three months. Good answers include: randomizing by service or week rather than by person (one congregation, shared space), total donations and donor count as metrics, and a plan that fits the three month deadline.
5. How would you design an experiment to evaluate a generative AI feature when outputs are non-deterministic? Good answers include: measuring user outcomes (task completed, output kept or edited) instead of single outputs, larger samples to absorb variance, and a quality guardrail from human or model graded samples.
How AllthingsPM does this. Open any of these five pages in the AllthingsPM question bank and press practice. The AI interviewer asks the follow-ups a real one would, such as "what is your randomization unit?" and "what would make you stop the test early?", then scores the answer.
Group 2: How do you answer "what tests would you run" questions?
Here the interviewer gives you a goal and wants a short, ranked list of experiments. Give three or four ideas, each with a hypothesis, and a reason for the order.
6. What A/B tests would you run to increase the number of messages sent and received on WhatsApp? Good answers include: splitting the goal into senders and receivers, tests on reply prompts or group creation, and a guardrail against spam or blocks rising.
7. What A/B tests do you suggest to make it easier to find friends on Facebook? Good answers include: a funnel from suggestion shown to request sent to request accepted, tests at the weakest step, and accepted requests (not sent requests) as the metric that matters.
8. What A/B tests will you run to increase the booking rate among Airbnb guests? Good answers include: a search to listing to checkout funnel, ideas tied to each drop off (price clarity, reviews, flexible dates), and cancellation rate as a guardrail so bookings are not just pulled forward.
9. Figma's homepage gets healthy traffic but first visit sign up conversion is below target. What 3 to 5 experiments would you run, and how would you tell a true win from shifting users downstream? Good answers include: one hypothesis per experiment, and a downstream check (activation or first file created) so a sign up lift that produces empty accounts does not count as a win.
10. How would you run a promotion to increase top line, in store revenue for Target? How would you decide what to promote? How would you run the experiment? Good answers include: store level randomization with matched stores, total basket size rather than promoted item sales, and a check for cannibalization of full price items.
How AllthingsPM does this. Ideation questions reward structure, and the AllthingsPM mock interview pushes you on ranking: after you list ideas, it asks which one you would run first and why. You can also drill a single company through its hub, such as Meta's question page.
Group 3: How do you answer "the test result is messy" questions?
One metric went up, another went down. The interviewer wants a decision and the reasoning behind it.
11. You are the PM at DoorDash. A/B testing shows orders go up, but the number of restaurants ordered from drops from 100 to 90. Will you launch? Good answers include: what concentration of orders means for the long tail of restaurants on the marketplace, whether the 10 lost restaurants churn off the platform, and a clear call with a condition attached.
12. A new signup flow raised the share of users adding profile information by 8%, but 7 day retention fell 2%. What do you do? Good answers include: treating retention as the metric that matters more, checking whether the drop is significant and where in the flow people left, then iterating on friction rather than shipping as is.
13. An Instagram A/B test raised Stories usage 10% and cut overall Instagram usage 10%. What metrics would you ask for to make the trade off? Good answers include: time spent per surface, ad impressions and revenue per user, and creator side effects, then a view on whether a 10% overall drop is ever acceptable.
14. A paywall experiment increases checkout conversion but shifts users to a cheaper plan and lowers retention. Ship, iterate or roll back? Good answers include: revenue per visitor over a longer window instead of conversion alone, plan mix, and a lean toward iterating because the retention hit compounds.
15. You ran an A/B test and saw that it drops engagement. What would you do about it? Good answers include: checking test health first (right split, no logging bug), segmenting the drop, and asking whether the feature was meant to cut low value engagement in the first place.
How AllthingsPM does this. Trade off questions are where follow-ups matter, and the AllthingsPM interviewer keeps asking "so, do you ship?" until you commit. For the thinking behind the second metric, the lesson on success metrics and guardrails in the AI PRD and the post on counter metrics cover it.
Group 4: How do you answer metrics and validity questions?
This group asks what the numbers mean and whether to trust them. Three ideas carry most answers. Guardrail metrics exist to show a change did not make the product measurably worse, which is a different test from proving it is better [3]. Peeking at results and stopping when they look significant inflates false positives; Evan Miller shows a nominal 5% level can become 26.1% [4]. And a sample ratio mismatch, where the split is not what you set, is a sign the result cannot be trusted [5].
16. Facebook ran an A/B test on a randomized 1% sample to increase posts seen per day in newsfeed. Which metrics would be affected, and what data should you expect? Good answers include: posts seen, time spent, likes and comments per post (which may fall as feed quality dilutes), ad load, and a prediction of direction for each.
17. A survey of two groups of 50k Facebook users found satisfaction 30% lower among users who enabled login security features. Why, and how was the survey conducted? Good answers include: spotting that users chose to enable the feature, so this is self selection, not a randomized test; plus friction from extra login steps as a real cause worth testing properly.
18. You own recommendations on an e-commerce marketplace. What metrics would you present to execs from last quarter's A/B test? Good answers include: one headline metric tied to revenue, the confidence around it, guardrails that held, and a recommendation, all in one slide worth of words.
19. Your team changed the Share feature and released it for A/B testing, and it has 20% usage. Would you still release it? Good answers include: asking 20% compared to what baseline, what the change was meant to move, and whether downstream metrics (new users from shares) changed.
20. Perplexity launches a "resume previous work" feature. How would you separate durable user value from short term spikes caused by novelty or accidental usage? Good answers include: watching the treatment effect over several weeks for decay, cohorts by first exposure, and intentional use signals. Research on novelty and primacy effects shows early effects can differ from the long term effect [7].
How AllthingsPM does this. Each question page in the AllthingsPM bank has an answer guide that names the metrics a strong answer uses, so you can check your list against it before you run the mock. The metrics interview questions post is a good companion for this group.
Group 5: How are A/B questions asked at AI companies?
AI companies ask experiment questions too, with a new wrinkle: a model can win on offline evals and lose with real users. Expect questions about connecting the two.
21. A new model version improves offline evals, but you are not convinced it improves the real creation experience. How would you design an experiment to validate user value end to end? Good answers include: a hypothesis about user behaviour, online success metrics (outputs kept, second creation), guardrails, segments, and a rule for when offline and online disagree.
22. For a retrieval or ranking improvement, how would you tell whether an offline metric predicts user value, and validate it through online experiments? Good answers include: running several model variants online and checking whether offline rank order matches online results, then keeping only offline metrics that track.
23. OpenAI is testing a new learning feature in ChatGPT. How would you design the experiment to tell whether it improves learning outcomes rather than just session length? Good answers include: an outcome measure of learning (for example a follow up quiz), engagement as a secondary metric, and a warning that longer sessions can mean confusion.
24. Suno's team has plenty of test ideas, but experiments are slow and leaders don't trust results. How would you rebuild the experimentation system end to end? Good answers include: event definitions and instrumentation first, automated checks such as sample ratio mismatch, pre agreed statistical standards, and a review cadence.
25. A team wants to launch next week, but Statsig cannot yet guarantee safe rollback, exposure logging or readout quality. Approve a fast launch or delay? Good answers include: sizing the blast radius, what is reversible, a staged rollout with a kill switch as a middle path, and the signals that would flip the call.
How AllthingsPM does this. The AllthingsPM course teaches this link between offline and online measurement in lessons such as eval-driven development and business outcomes, not eval scores. If you are interviewing at an AI company, the jobs catalog has live PM job descriptions with a mock built from each.
What does a strong A/B testing answer include every time?
Across all 25 questions, the answers that score well share six parts:
- A hypothesis in one sentence. "Showing X to Y users will raise Z because W."
- One primary metric and one or two guardrails. Decide them before the result, not after.
- The unit of randomization. User, session, store, city or time block, and why. Marketplaces and social features need extra care.
- Sample and duration. You do not need the formula, but say that you would size the test for the smallest effect worth shipping and run full weekly cycles.
- Validity checks. Split ratio, no peeking, novelty decay.
- A decision. Ship, iterate or kill, with the condition that would change your mind.
Interview Query covers the statistics side if you want to go deeper [8].
Why AllthingsPM is the better choice for A/B testing interview prep
Most lists stop at the question. AllthingsPM adds an answer guide and an interviewer that makes you defend your call, which matters most for A/B questions, where the hard moment is the follow-up: "your guardrail moved, do you still ship?"
- 51 real A/B and experiment questions inside a bank of 4,122 questions from 260 companies, each with its own page.
- Scored mock interviews in text or voice from any question, with follow-ups.
- JD-based mocks from any job description, plus live PM job descriptions at AI companies in the jobs catalog.
- An AI PM course built from 604 real PM job postings, with lessons on metrics, guardrails and evaluation.
- Resume review against a JD at /jd-resume-review, for the step before the interview.
Other resources have real strengths. Amplitude and Interview Query publish solid free explainers, and Aced offers human coaching. For daily practice on real questions with a scored interviewer, AllthingsPM is the stronger choice, and it starts free with one JD mock a day. Paid plans are $20 a month or $120 a year.
Start practicing A/B testing questions free on AllthingsPM.
Frequently asked questions
What are the most common A/B testing interview questions for PMs?
The most common shapes are "design an A/B test for this feature", "what tests would you run to grow X" and "one metric rose and another fell, do you ship?". All three appear in the AllthingsPM bank, for example the DoorDash orders versus restaurants question and the Instagram Stories versus overall usage question.
What is the best way to practice A/B testing interview questions?
AllthingsPM is the best place to practice them: 51 real A/B and experiment questions, each with a page and answer guide, and a scored AI mock interview with follow-ups in text or voice. Free explainers from Amplitude and Interview Query help with the theory.
Do PMs need to know statistics for A/B testing interviews?
You need concepts, not derivations: statistical significance, sample size, why peeking inflates false positives, and what a sample ratio mismatch means. Interviewers care most about metric choice and the decision you make.
How is this list different from the AllthingsPM A/B testing guide?
This post is a practice list of 25 real questions with what good answers include. The A/B testing interview guide teaches the framework and walks through full worked answers.
Which companies ask A/B testing questions?
In the AllthingsPM bank, questions name DoorDash, Uber, Amazon, Airbnb, Instagram, Facebook, WhatsApp, Target, Figma, Perplexity, OpenAI and Suno, among others. Aced also names Google, Microsoft and Amazon [1].
How long should an answer to an A/B testing question be?
Aim for about five minutes of structured talk, then leave room for follow-ups. Cover the hypothesis, metrics, unit, duration, risks and decision, and stop once you have made the call.
Sources
- Aced (formerly Exponent), "How to Ace A/B Testing Interview Questions": https://www.tryexponent.com/blog/how-to-ace-ab-testing-interview-questions
- Akhil Prakash, Amplitude, "22 A/B Testing Interview Questions and How to Answer Them" (updated 13 August 2024): https://amplitude.com/blog/a-b-testing-interview-questions
- Statsig, "What are guardrail metrics in A/B tests?": https://www.statsig.com/blog/what-are-guardrail-metrics-in-ab-tests
- Evan Miller, "How Not To Run an A/B Test": https://www.evanmiller.org/how-not-to-run-an-ab-test.html
- Kohavi, Tang and Xu, Trustworthy Online Controlled Experiments, Chapter 21, "Sample Ratio Mismatch and Other Trust-Related Guardrail Metrics", Cambridge University Press: https://www.cambridge.org/core/books/abs/trustworthy-online-controlled-experiments/sample-ratio-mismatch-and-other-trustrelated-guardrail-metrics/8DBB0F59AC7729D7BC6B94690DB9CCD5
- DoorDash Engineering, "Switchback Tests and Randomized Experimentation Under Network Effects at DoorDash": https://careersatdoordash.com/blog/switchback-tests-and-randomized-experimentation-under-network-effects-at-doordash/
- "Novelty and Primacy: A Long-Term Estimator for Online Experiments", arXiv: https://arxiv.org/pdf/2102.12893
- Interview Query, "Statistics and A/B Testing Interview Questions": https://www.interviewquery.com/p/statistics-ab-testing-interview-questions
- AllthingsPM question bank, 4,122 questions from 260 companies, queried 29 September 2026: https://allthingspm.app/question-bank




