Metrics question
This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?
- OpenAI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Metrics design that connects product usage to real defensive outcomes and forces an explicit tradeoff decision.
How to approach it
- Propose product quality metrics: investigation accuracy, measured against expert-labeled ground truth cases, and false positive or false negative rate on flagged threats.
- Propose operational outcome metrics: mean time to investigate an alert, and defensive outcomes per analyst-hour, the named north star, calculated as investigations completed per hour of analyst time spent.
- Propose user behavior metrics: adoption rate among analysts, and reliance rate, meaning how often analysts accept the AI's investigation findings without significant manual rework.
- Classify leading versus lagging: adoption and reliance rate are leading indicators of habit formation, while mean time to investigate and accuracy are lagging indicators of actual defensive value delivered.
- Address the tradeoff explicitly: if adoption is high but accuracy or safety is weak, pause scaling and prioritize accuracy fixes first, since high adoption of an inaccurate tool actively increases risk rather than reducing it.
What a strong answer includes
- Organizes metrics into the three named categories explicitly, product quality, operational outcomes, user behavior, rather than one undifferentiated list.
- Correctly classifies adoption and reliance as leading indicators and time-to-investigate or accuracy as lagging, showing understanding of how these metrics move at different speeds.
- Makes the tradeoff decision explicit and directionally clear: high adoption with weak accuracy is a reason to slow down, not scale up, since it compounds risk.
- Names a concrete north star aligned to the question's own framing, defensive outcomes per analyst-hour, and defines exactly how it would be calculated.
Common mistakes
- Listing metrics without classifying which are leading versus lagging, missing an explicit part of the question.
- Treating high adoption as an unambiguous success signal without addressing the case where accuracy or safety is weak.
Likely follow-up questions
- How would you build the expert-labeled ground truth needed to measure investigation accuracy?
- What threshold of accuracy would you require before allowing continued scaling despite high adoption?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
- An early rollout of a multimodal model is driving strong user retention, but harmful and policy-sensitive image and audio outputs are also rising. What safety and user-experience metrics would you define, how would you instrument them, and what thresholds would trigger scaling the rollout, limiting it, or pausing it?OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop