Metrics question
Suppose OpenAI is testing a new learning feature in ChatGPT, such as guided practice or personalized explanations. How would you design the experiment so you can tell whether it improves learning outcomes rather than just session length or retention, and what success metrics would you use?
- OpenAI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Tests experiment design skill: isolating true learning impact from engagement proxies like session length or return rate.
How to approach it
- Define the actual outcome of interest, measured learning gain on a pre and post assessment, rather than proxy engagement metrics.
- Design a randomized controlled test between the new feature and the current experience, not just a before and after comparison.
- Include a delayed post test, not just immediate assessment, to check whether the learning gain persists rather than just improving short term recall.
- Track session length and retention as secondary metrics, explicitly separate from the primary learning outcome metric, to avoid conflating the two.
- Control for selection effects, since students who opt into a new feature may already be more motivated learners.
- Set a pre registered success threshold for the learning gain metric before running the experiment, to avoid post hoc rationalization of engagement wins.
What a strong answer includes
- Uses a pre and post assessment with a delayed retest, which is what actually separates real learning from short term engagement.
- Explicitly separates primary outcome, learning gain, from secondary engagement metrics instead of letting engagement substitute for the real question.
- Controls for selection bias, since more motivated students opting into a new feature would inflate apparent learning gains.
Common mistakes
- Using session length or return rate as the primary success metric for a learning claim.
- Skipping a delayed retest, which would miss cases where a feature helps short term recall but not durable learning.
Likely follow-up questions
- How would you handle a case where engagement is high but the delayed retest shows no real gain?
- What would you do if a true randomized test is not feasible in a live classroom setting?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop