Metrics question

Suppose OpenAI is testing a new learning feature in ChatGPT, such as guided practice or personalized explanations. How would you design the experiment so you can tell whether it improves learning outcomes rather than just session length or retention, and what success metrics would you use?

Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.

Start a mock interview on this question · Mock interview from a job description

What this question tests

Tests experiment design skill: isolating true learning impact from engagement proxies like session length or return rate.

How to approach it

  1. Define the actual outcome of interest, measured learning gain on a pre and post assessment, rather than proxy engagement metrics.
  2. Design a randomized controlled test between the new feature and the current experience, not just a before and after comparison.
  3. Include a delayed post test, not just immediate assessment, to check whether the learning gain persists rather than just improving short term recall.
  4. Track session length and retention as secondary metrics, explicitly separate from the primary learning outcome metric, to avoid conflating the two.
  5. Control for selection effects, since students who opt into a new feature may already be more motivated learners.
  6. Set a pre registered success threshold for the learning gain metric before running the experiment, to avoid post hoc rationalization of engagement wins.

What a strong answer includes

Common mistakes

Likely follow-up questions

More metrics questions

More questions from OpenAI

Learn the skill behind it

Chapters of the AI PM course that teach what this question tests.

Preparing for a specific role?

Book summaries for this kind of question

Browse all 4,000+ questions in the bank