Metrics question
A launched Sierra agent has increased containment significantly, but CSAT is flat and escalations are rising. How would you diagnose the issue? What metrics and instrumentation would you review at the conversation, intent, and handoff levels, how would you segment the problem, and what improvements would you prioritize first?
- Sierra
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Tests structured diagnosis of a specific automation paradox, rising containment but flat CSAT and rising escalations, using layered instrumentation.
How to approach it
- Review conversation level metrics first: are escalations happening after multiple failed automated attempts, suggesting the agent is trying too hard to contain difficult cases.
- Review intent level metrics: segment containment and CSAT by intent to see if certain intents are being over automated relative to their real complexity.
- Review handoff level metrics: when escalation does happen, check if the human agent has full context, since a poor handoff itself can suppress CSAT even after a correct escalation.
- Check if containment gains came from stricter definitions of what counts as contained, which could mask cases that were technically closed but left the customer unsatisfied.
- Segment the problem by customer type or issue complexity to see if the paradox concentrates in a specific segment rather than being uniform.
- Prioritize fixes based on where the data points: narrowing automation scope on over automated intents, or fixing handoff context if that is the bigger driver.
What a strong answer includes
- Checks whether the definition of contained itself changed, a subtle but common cause where a metric improves on paper while the real customer experience does not.
- Reviews handoff context quality specifically, recognizing that CSAT can suffer from a bad escalation experience even when the escalation decision itself was correct.
- Segments by intent and customer type to find where the paradox concentrates rather than treating it as a uniform, system wide problem.
Common mistakes
- Assuming rising containment is unambiguously good without checking whether the definition of contained itself may be too loose.
- Fixing the automation broadly without checking whether poor handoff context, not the escalation decision itself, is driving flat CSAT.
Likely follow-up questions
- How would you tell if the containment definition itself needs tightening?
- What would you do if the paradox is concentrated in one specific intent rather than spread evenly?
More metrics questions
- What metrics prove ROI to a Fortune 500 company deploying Sierra?Sierra · Metrics · Hard
- How would you measure customer satisfaction with an AI support agent?Sierra · Metrics · Medium
- What are the most important metrics for an infrastructure platform powering enterprise AI agents, and how would you organize them into a scorecard? Include how you would measure latency, availability, fault tolerance, and developer productivity, and explain which leading indicators you would monitor to catch problems before they show up in customer impact.Sierra · Metrics · Medium
- After launch, how would you measure whether Sierra’s Agent SDK is succeeding? Define the leading and lagging metrics you would use across developer adoption, implementation quality, and downstream end-user outcomes, and explain how those metrics would influence roadmap decisions.Sierra · Metrics · Medium
- For an agent that collects and routes fraud, waste, and abuse reports, what success metrics would you define for both the institution and the end user? If report volume and completion rates are high but downstream resolution quality is poor, how would you diagnose the problem and prioritize fixes?Sierra · Metrics · Hard
- One live agent has high conversation volume but low containment, with many users escalating to human support. What metrics would you inspect first, how would you segment the problem, and how would you determine whether the main issue is conversation design, model behavior, or the customer’s backend workflow integration?Sierra · Metrics · Hard
More questions from Sierra
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop