Metrics question

For a retrieval or ranking improvement, how would you determine whether an offline metric is actually predictive of user value rather than just easy to optimize, and how would you validate that relationship through online experiments?

Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.

Start a mock interview on this question · Mock interview from a job description

What this question tests

Rigor in connecting an offline proxy metric to real user value and designing the online validation to prove or disprove it.

How to approach it

  1. Name the specific offline metric, for example a relevance score on a labeled dataset, and the user value it is meant to proxy, such as task success or satisfaction.
  2. Check the labeling process for bias, since offline labels can reward things easy to game, like keyword overlap, rather than true relevance.
  3. Run a held out online experiment where the ranking change is shipped to a subset of users, measuring downstream behavioral metrics like click through, task completion, or reformulation rate.
  4. Compare the offline metric's movement against the online outcome across multiple experiments over time, not just once, to see if they consistently correlate.
  5. Watch for cases where offline improves but online user value does not, which signals the offline metric is being gamed or missing what users actually want.
  6. Decide whether to keep, recalibrate, or replace the offline metric based on that correlation, and document the finding for future launch decisions.

What a strong answer includes

Common mistakes

Likely follow-up questions

More metrics questions

More questions from OpenAI

Learn the skill behind it

Chapters of the AI PM course that teach what this question tests.

Preparing for a specific role?

Book summaries for this kind of question

Browse all 4,000+ questions in the bank