Metrics question
An AI agent for patient access and scheduling has high adoption, but resolution rate is below target because too many conversations escalate to human staff. How would you break down the funnel, determine whether the root cause is workflow design, knowledge gaps, policy boundaries, or user behavior, and prioritize the first three fixes? What leading and lagging metrics would you use to know the agent is becoming core infrastructure?
- Decagon
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Whether you can decompose a resolution-rate shortfall into distinct root causes, prioritize fixes with real reasoning, and define metrics that track the agent becoming trusted infrastructure, not just handling volume.
How to approach it
- Break the funnel by escalation reason: workflow design gaps (the agent cannot complete a valid task type at all), knowledge gaps (the agent lacks information to answer correctly), policy boundaries (the agent correctly defers because a human must handle it), and user behavior (ambiguous or out-of-scope requests).
- Separate true failures from appropriate escalations first, since policy-boundary escalations are working as intended and should not be counted the same as workflow or knowledge failures when prioritizing fixes.
- Prioritize the first three fixes by volume times fixability: a high-volume knowledge gap that is cheap to close (updating a knowledge base) likely outranks a lower-volume but harder workflow redesign.
- Fix knowledge gaps first if they dominate volume, since they are typically the fastest to close, then workflow design gaps, then revisit policy boundaries only if they are miscalibrated, escalating things that could safely be automated.
- Track leading indicators like escalation reason distribution shifting away from knowledge and workflow gaps over time, and lagging indicators like sustained resolution rate improvement plus patient or staff trust signals, such as repeat usage without complaint, to know the agent is becoming core infrastructure rather than a novelty.
What a strong answer includes
- Separates true failures from appropriate policy-driven escalations before prioritizing fixes, a distinction the question implicitly requires to avoid mis-scoping the problem.
- Uses a volume-times-fixability rule to order the first three fixes concretely instead of listing them with no priority.
- Defines core-infrastructure status with both a leading indicator (escalation reason mix shifting) and a lagging one (sustained trust and resolution improvement), not just raw adoption.
Common mistakes
- Treats all escalations as failures, including appropriate policy-driven ones that should not be automated away.
- Lists three fix categories without any prioritization logic for which to tackle first.
Likely follow-up questions
- How would you tell a policy-boundary escalation apart from a knowledge-gap failure in ambiguous cases.
- What would make you conclude the agent has become trusted infrastructure rather than just widely used.
More metrics questions
- A large customer has an AI agent live in production, but adoption is below plan and leadership is hesitating on expansion. What metrics would you review first, how would you isolate whether the issue is workflow selection, agent quality, operational rollout, or stakeholder buy-in, and what actions would you take in the next 30 days to improve adoption and demonstrate business impact?Decagon · Metrics · Hard
- You have inherited a new strategic account and must choose the first customer-support workflows to automate in production. What prioritization framework would you use to decide where the agent goes live first, and which adoption, quality, and business metrics would you require before recommending expansion into additional workflows or channels?Decagon · Metrics · Hard
- A live enterprise agent is generating strong customer demand for expansion, but engineers report unresolved reliability gaps in the current design. How would you decide what to ship next, including what evidence or thresholds you would require to expand safely, what you would defer, and how you would manage the conversation with the customer’s leadership team and internal engineering partners?Decagon · Metrics · Hard
- How would you define a metrics framework for Decagon’s developer experience across APIs, SDKs, and headless deployments? Specify the leading and lagging metrics you’d track from integration start through production launch, and explain how those metrics would change your roadmap priorities.Decagon · Metrics · Hard
- One of Decagon's largest customers has launched an agent, but adoption has plateaued because internal teams will not let it handle higher-value interactions. How would you diagnose whether the bottleneck is model quality, workflow design, integration gaps, or change management, and how would you decide which intervention to make first?Decagon · Metrics · Hard
- You own a customer support agent from first production launch through enterprise-wide expansion. What success metrics would you track in the first 30-60 days versus six months later, and how would you balance business outcomes, customer experience, and operational reliability when those metrics conflict?Decagon · Metrics · Hard
More questions from Decagon
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop