Metrics question
What metrics would you use to evaluate a conversational shopping experience in ChatGPT? Define a north-star metric and supporting metrics for user value, trust, merchant ecosystem health, and business outcomes, and explain the tradeoffs among them.
- OpenAI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Tests north star and supporting metric design for a three sided shopping marketplace balancing user, merchant, and business interests.
How to approach it
- Choose a north star that reflects genuine user value delivered, such as confident completed purchases or decisions made without regret, rather than raw engagement or GMV.
- Track user value supporting metrics: discovery to decision time, comparison depth used, and post purchase satisfaction or return rate.
- Track trust supporting metrics: disclosure clarity ratings, recommendation accuracy against actual price and availability, and complaint rate.
- Track merchant ecosystem health: merchant participation rate, fair exposure across merchants rather than concentration in a few large ones, and merchant reported conversion quality.
- Track business outcomes: take rate or commission revenue, repeat shopping session rate, and overall GMV as a lagging validation, not the primary target.
- Explain the tradeoff explicitly: optimizing purely for GMV risks pushing users toward more expensive or less suitable options, undermining trust and merchant ecosystem health long term.
What a strong answer includes
- Chooses a north star around decision confidence or completed value rather than raw GMV, avoiding the trap of optimizing for revenue at the cost of trust.
- Includes merchant ecosystem health as its own category, recognizing this is a three sided marketplace, not just a user and business relationship.
- States the GMV versus trust tradeoff explicitly, showing awareness that a purely revenue optimized recommender would undermine long term user and merchant trust.
Common mistakes
- Choosing GMV or conversion rate as the north star, which creates a direct incentive to push users toward suboptimal purchases.
- Ignoring merchant ecosystem health entirely, treating this as a two sided user and business metric problem instead of three sided.
Likely follow-up questions
- How would you detect if the recommendation system is quietly favoring merchants who pay more over what is actually best for the user?
- Which one metric would you protect even if it meant sacrificing short term GMV growth?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop