Metrics question
What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.
- OpenAI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Metrics design for a dual-objective product where optimizing one goal, helpfulness, could actively work against the other, safety.
How to approach it
- Propose a north-star metric that blends both goals conceptually, like the rate of sessions rated by review as both age-appropriate and genuinely helpful to the user's need.
- Propose guardrail metrics separately from the north star: rate of confirmed policy-violating content served to under-18 accounts, held to a near-zero threshold regardless of engagement.
- Name a metric that could be misleading: overall engagement or session length, since rising engagement could reflect either genuine helpfulness or unsafely permissive behavior, and cannot distinguish the two alone.
- Explain how these metrics would change the roadmap: if the guardrail metric worsens even slightly, that should pause any feature expansion, prioritized above growth-oriented improvements to the north star.
- Propose combining automated detection with periodic expert human review of a sampled set of teen conversations, since safety violations are often too subtle or context-dependent for automated flags alone.
What a strong answer includes
- Explicitly names engagement or session length as a metric that could mislead, correctly identifying that it cannot on its own separate genuine helpfulness from unsafe permissiveness.
- Sets a near-zero threshold guardrail for confirmed safety violations, treating it as non-negotiable rather than something traded off against growth metrics.
- States a clear roadmap rule: any guardrail regression pauses feature expansion, giving a concrete answer to how metrics would change priorities.
- Proposes combining automated and human expert review, recognizing that safety in nuanced teen conversations cannot be fully captured by automated signals alone.
Common mistakes
- Proposing engagement or usage growth as the north star without flagging it as a potentially misleading signal for this specific product.
- Failing to state a clear rule for how a guardrail regression should actually change the roadmap, leaving the tradeoff unresolved.
Likely follow-up questions
- How would you set the sample size for human review to catch rare but serious violations?
- What would you do if the guardrail metric and the north star metric moved in opposite directions?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- An early rollout of a multimodal model is driving strong user retention, but harmful and policy-sensitive image and audio outputs are also rising. What safety and user-experience metrics would you define, how would you instrument them, and what thresholds would trigger scaling the rollout, limiting it, or pausing it?OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop