Metrics question
What metrics would you track to measure the success of ChatGPT Projects?
- OpenAI
- Metrics
- Medium
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Whether the candidate can define metrics for an organizational feature where success means sustained use, not one-time novelty.
How to approach it
- Clarify the feature's job: Projects lets users group conversations and files around an ongoing task, so success means it reduces friction for recurring work.
- Propose adoption metrics: percentage of active users who create at least one Project, and percentage who return to a Project after the first session.
- Propose depth metrics: average number of conversations and files per active Project, showing it is used as intended, not just created once.
- Propose a guardrail: churn rate of Projects, meaning created but abandoned within a set window, say two weeks.
- Tie to business outcome: whether Project users show higher overall retention or upgrade rate than non-Project users.
What a strong answer includes
- Distinguishes creation from sustained use, since a Project created once and never revisited is not a success.
- Proposes a concrete return metric, like percentage of Projects with a second session within 7 days, as the real signal of habit formation.
- Connects Project usage to a business-level outcome, hypothesizing that Project users retain or upgrade at a higher rate, and proposes testing that correlation.
- Flags a guardrail against a misleading vanity metric, like total Projects created, which could rise even if most are abandoned.
Common mistakes
- Measuring only Project creation count without checking whether Projects get revisited.
- Ignoring the connection between this feature and broader retention or monetization outcomes.
Likely follow-up questions
- What would you do if Projects are created often but rarely revisited?
- How would you segment Project usage between individual and team-oriented use cases?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
- An early rollout of a multimodal model is driving strong user retention, but harmful and policy-sensitive image and audio outputs are also rising. What safety and user-experience metrics would you define, how would you instrument them, and what thresholds would trigger scaling the rollout, limiting it, or pausing it?OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop