Metrics question
After a pricing and packaging change for a commercial product, checkout conversion drops materially. How would you diagnose whether the issue is caused by demand elasticity, messaging confusion, payment failures, funnel UX regressions, or customer mix changes? What data would you inspect first, and what criteria would you use to decide whether to roll back, iterate, or keep the launch running?
- OpenAI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Tests structured diagnosis of a conversion drop after a pricing change, separating demand, messaging, payment, UX, and mix causes with clear decision criteria.
How to approach it
- Check payment failures first, since a technical regression is the fastest to confirm or rule out and the easiest to fix if found.
- Check funnel UX for regressions introduced alongside the pricing change, such as a broken checkout step or unclear new price display.
- Check messaging clarity, since a pricing and packaging change can confuse returning users about what they are now paying for even if the funnel works correctly.
- Check customer mix, since a pricing change can shift who even reaches checkout, attracting a more price sensitive segment that converts at a naturally lower rate.
- Distinguish demand elasticity, real price sensitivity, from the above operational causes by comparing conversion drop magnitude against what elasticity models would predict.
- Decide based on cause: roll back for a technical regression, iterate messaging or UX for confusion or friction, or keep running if the drop matches expected elasticity and the new pricing still improves overall revenue.
What a strong answer includes
- Checks technical and UX regressions first, since those are fastest to confirm and most likely explanation for a sudden, material drop right after a change.
- Distinguishes real demand elasticity from operational causes by comparing the actual drop against what elasticity would predict, avoiding a premature rollback for expected price sensitivity.
- Ties the final decision explicitly to cause: rollback only for a technical or UX regression, not simply because conversion dropped.
Common mistakes
- Rolling back immediately without first checking for a payment or UX regression that could be fixed without abandoning the new pricing.
- Attributing the drop entirely to demand elasticity without checking for technical or messaging causes that could be fixed quickly.
Likely follow-up questions
- How would you model expected elasticity to know if the drop is larger than predicted?
- What would you do if the drop is concentrated in one customer segment but not others?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop