Metrics question
How would you measure the success of the Netflix recommendation engine?
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Metric design for a machine learning system, balancing engagement outcomes against catalog diversity and discovery health.
How to approach it
- State the goal, help users find content they will enjoy quickly, increasing watch time and reducing decision fatigue.
- Define the primary metric, percent of viewing sessions initiated from a recommended title, combined with completion rate for those titles.
- Define a quality guardrail, recommendation diversity, since an engine that only recommends the same popular titles repeatedly would look good on completion rate but fail at real discovery value.
- Define a satisfaction proxy, rating or thumbs up rate on recommended content, catching cases where users finish content but did not actually enjoy it.
- Consider long term impact, whether users who engage more with recommendations show higher 90 day retention than those who do not.
What a strong answer includes
- Names a specific primary metric, recommended title completion rate, distinct from overall watch time, better isolating the engine's specific contribution.
- Adds a diversity guardrail explicitly, since an engine could maximize short term completion by only ever recommending obvious blockbusters.
- Ties the metric to retention, testing whether good recommendations causally improve 90 day retention, not just immediate viewing behavior.
Common mistakes
- Using only overall watch time as the metric, which does not isolate the recommendation engine's specific effect.
- No diversity or catalog health guardrail, risking an engine that only ever suggests the most popular titles.
- No connection made between recommendation quality and longer term retention.
Likely follow-up questions
- How would you isolate the recommendation engine's effect from a user's own active searching?
- How would you handle the trade off between short term completion rate and long term catalog diversity?
More metrics questions
- A metric for a video streaming service dropped by 80%. What do you do?Google · Metrics · Medium
- How would you measure the success of YouTube?Spotify · Metrics · Medium
- There are about 1 million inactive Netflix users. What would you do about them?Spotify · Metrics · Hard
- Pick a feature from a product of your choice & tell us how will you assess the success of that feature. Assume it is being launched now.Google · Metrics · Medium
- What are the key metrics to be captured for a streaming service product?Spotify · Metrics · Easy
- Netflix's subscription renewal rate is declining year over year. What hypotheses would you propose, and how would you validate them?Netflix · Metrics · Hard
More questions from these companies
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop