Metrics question
You launch a new app ecosystem surface in ChatGPT. What north-star, guardrail, and ecosystem-health metrics would you track across users, partners, and the platform? How would your metric set differ between consumer self-serve usage and enterprise deployments with admins and compliance requirements?
- OpenAI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Whether you can define a layered metrics framework for a platform launch and adapt it for very different consumer and enterprise contexts.
How to approach it
- Define north-star metrics: weekly active users engaging with third-party apps, and for partners, active integrations generating repeat usage.
- Define guardrail metrics: safety and policy violation rate per app, user complaint rate, and app approval-to-incident ratio.
- Define ecosystem-health metrics: partner retention, time from submission to approval, and distribution of usage across apps (to catch over-concentration risk).
- For consumer self-serve, weight discovery and activation metrics higher, since the user journey is unassisted and low-friction.
- For enterprise, add admin-specific metrics: percent of admins who configure app allowlists, compliance audit completion rate, and support ticket volume tied to governance.
- Set guardrails tighter for enterprise given compliance requirements, for example a stricter safety-incident threshold before an app can be surfaced to enterprise users.
What a strong answer includes
- Separates north-star, guardrail, and ecosystem-health metrics into three distinct tiers rather than one flat list.
- Adds enterprise-specific governance metrics (allowlist configuration, compliance completion) that a consumer metric set would miss entirely.
- Names a concrete ecosystem-health risk, usage over-concentrated in a few apps, which threatens long-term platform diversity.
- Explains why guardrail thresholds should differ by context, tying it to compliance stakes in enterprise deployments.
Common mistakes
- Using one identical metric set for both consumer and enterprise without adapting for admin and compliance needs.
- No ecosystem-health metric, focusing only on user-side engagement.
- Missing a concrete guardrail for safety or policy violations.
Likely follow-up questions
- How would you detect an app becoming a systemic safety risk before it causes real harm?
- What would make you delist or restrict a popular but risky app?
- How would you balance ecosystem diversity against user preference for a few dominant apps?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop