Metrics question
What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.
- OpenAI
- Metrics
- Medium
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Whether the candidate can define different success bars for an early pilot versus a scaled rollout, not one static scorecard.
How to approach it
- Propose pilot-stage metrics, deliberately small and qualitative: direct user-value feedback from the five customers, like time saved per task, and trust and quality signals from close-read review of a sample of outputs by legal experts.
- Propose operational pilot metrics: task completion rate and manual correction rate per output, tracked closely given the small, observable sample.
- Propose scaled-rollout metrics: aggregate business metrics like customer retention and expansion, alongside statistically robust quality metrics no longer needing full manual review of every output, just a monitored sample.
- Classify leading versus launch gates: pilot-stage expert quality review is a launch gate for scaling, while aggregate retention and usage growth are leading indicators once scaled.
- Note the transition point explicitly: moving from pilot to scaled rollout requires the pilot's quality bar to hold steady as volume grows, not just that early adopters were satisfied.
What a strong answer includes
- Distinguishes pilot metrics, which can be closely observed and qualitative given small scale, from scaled metrics, which must be statistically robust and operationally lighter-touch.
- Names a concrete pilot-stage metric, expert legal review of outputs, as a hard gate rather than a nice-to-have signal.
- Proposes business metrics, retention and expansion, specifically for the scaled stage, since a five-customer pilot is too small to draw reliable business conclusions from.
- Explains the transition condition clearly, that quality must hold as volume grows, addressing the real risk of a pilot succeeding on hand-holding rather than the product itself.
Common mistakes
- Using the same metrics identically for both stages, ignoring that a five-customer pilot allows close observation that scaled rollout cannot sustain.
- Proposing only business metrics with no mention of quality or trust signals, missing the point of measuring whether the product is actually working.
Likely follow-up questions
- How would you decide the pilot is ready to scale beyond five customers?
- What would you do if quality held in the pilot but degraded once volume increased?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
- An early rollout of a multimodal model is driving strong user retention, but harmful and policy-sensitive image and audio outputs are also rising. What safety and user-experience metrics would you define, how would you instrument them, and what thresholds would trigger scaling the rollout, limiting it, or pausing it?OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop