Metrics question
OpenAI wants to improve order-to-cash for its largest enterprise customers. How would you segment customers, map the highest-friction steps from order creation through invoicing, payment collection, and reconciliation, and prioritize the first three product investments? Include the success metrics you would use and how you would balance requests from Finance, Operations, and Engineering.
- OpenAI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Whether you can map a B2B financial process end to end, prioritize with limited engineering capacity, and manage competing functional stakeholders.
How to approach it
- Segment customers by contract size and payment complexity, for example top 50 enterprise accounts with custom terms versus the broader enterprise base on standard terms.
- Map the order-to-cash chain: order creation, contract terms, invoicing, payment collection, dispute handling, reconciliation, and note where each function (Finance, Ops, Engineering) touches it.
- Interview each function to find the highest-friction step, for example manual reconciliation of usage-based invoices against actual API consumption.
- Prioritize the first three investments using a reach-times-friction score, for example automated invoice-to-usage matching, self-serve dispute resolution, and payment-method flexibility for large accounts.
- Set success metrics: days sales outstanding, percent of invoices requiring manual touch, reconciliation time, and dispute resolution time.
- Resolve competing asks by anchoring every request to its effect on DSO or manual-touch rate, not on who asked loudest.
What a strong answer includes
- Assumes usage-based billing is the main friction source for an AI API business and prioritizes invoice-to-usage matching first.
- Gives a concrete metric target, for example cut manual-touch invoices from an assumed 30% to under 10% in two quarters.
- Explicitly trades off Engineering's platform preference (a general billing rules engine) against Finance's need for a fast point fix, and picks a sequence rather than picking a side.
- Uses DSO as the single cross-functional metric everyone can rally around.
Common mistakes
- Listing generic order-to-cash steps without naming which one is actually highest-friction for this business.
- Prioritizing by stakeholder volume of requests instead of a reach or cost-of-friction score.
- No named success metric, so the investment cannot be judged after the fact.
Likely follow-up questions
- How would you validate that reconciliation, not collections, is really the biggest friction point?
- What would you do if Engineering says the fix requires a full rebuild of the billing system?
- How would you measure whether the fix actually reduced manual work, not just moved it?
More metrics questions
- Weekly active users of Codex dropped 15% after a pricing change. How do you investigate?OpenAI · Metrics · Medium
- What metrics would you track to measure the success of ChatGPT Projects?OpenAI · Metrics · Medium
- OpenAI wants one topline safety metric for frontier model deployments. How would you define it so it is credible for leadership decisions, sensitive enough to detect meaningful changes in harm, and decomposable into drivers that research and engineering teams can act on?OpenAI · Metrics · Hard
- This team cares about measurable improvement in defensive outcomes per analyst-hour. For an AI-assisted threat investigation product, what metrics would you use across product quality, operational outcomes, and user behavior? Which would be leading vs. lagging indicators, and how would you handle tradeoffs if adoption is high but investigation accuracy or safety is weak?OpenAI · Metrics · Hard
- What metrics would you use to judge whether a legal AI product is working in a 5-customer pilot versus a scaled rollout? Be specific about user-value, trust/quality, operational, and business metrics, and explain which ones are leading indicators versus launch gates.OpenAI · Metrics · Medium
- What metrics would you use to determine whether the under-18 ChatGPT experience is both helpful and safe? Define a north-star metric, guardrails, and one metric that could be misleading, then explain how those metrics would change your roadmap.OpenAI · Metrics · Hard
More questions from OpenAI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop