Metrics question

What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.

Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.

Start a mock interview on this question · Mock interview from a job description

What this question tests

Metrics design for a dual-objective product where optimizing one goal, helpfulness, could actively work against the other, safety.

How to approach it

  1. Propose a north-star metric that blends both goals conceptually, like the rate of sessions rated by review as both age-appropriate and genuinely helpful to the user's need.
  2. Propose guardrail metrics separately from the north star: rate of confirmed policy-violating content served to under-18 accounts, held to a near-zero threshold regardless of engagement.
  3. Name a metric that could be misleading: overall engagement or session length, since rising engagement could reflect either genuine helpfulness or unsafely permissive behavior, and cannot distinguish the two alone.
  4. Explain how these metrics would change the roadmap: if the guardrail metric worsens even slightly, that should pause any feature expansion, prioritized above growth-oriented improvements to the north star.
  5. Propose combining automated detection with periodic expert human review of a sampled set of teen conversations, since safety violations are often too subtle or context-dependent for automated flags alone.

What a strong answer includes

Common mistakes

Likely follow-up questions

More metrics questions

More questions from OpenAI

Learn the skill behind it

Chapters of the AI PM course that teach what this question tests.

Preparing for a specific role?

Book summaries for this kind of question

Browse all 4,000+ questions in the bank