Metrics question
If you joined Harvey, what 2 or 3 AI-assisted workflows would you implement first for the product communications team, and why those first? For each workflow, explain the job to be done, the human review points, the main failure modes or brand risks, and the metrics you would use to decide whether it improved speed or quality enough to scale.
- Harvey
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Whether you can pick and justify a small, high-leverage set of AI workflows for a specific function, with real attention to human review points and brand risk given the sensitivity of legal-industry communications.
How to approach it
- Pick draft generation for press releases and briefing documents first, since it has a clear job to be done (turn approved facts into on-brand copy fast) and an obvious human review point before anything goes external.
- Pick message-testing and audience-tailoring second, using AI to draft variants of the same announcement for technical, legal, and business audiences, reviewed by the comms lead before use.
- Pick competitive and media monitoring third, using AI to surface relevant coverage and sentiment for the team to act on, which is lower-risk since it informs rather than publishes.
- For each, name the failure mode: draft generation risks overclaiming AI capability or legal accuracy errors, audience tailoring risks inconsistent core messaging across versions, monitoring risks missing context behind a headline.
- Set the scaling metric per workflow: time saved per press cycle for drafting, message consistency score across audience variants for tailoring, and time to surface relevant coverage for monitoring, each reviewed after a defined pilot period.
What a strong answer includes
- Picks three workflows with increasing risk tolerance (drafting, tailoring, monitoring) and explains why that order makes sense for a legal AI company's own brand risk.
- Names a specific failure mode per workflow instead of a single generic AI risk warning.
- Defines a distinct scaling metric per workflow rather than one blanket productivity number.
Common mistakes
- Picks generic AI workflows that could apply to any comms team, without connecting them to Harvey's specific brand risk as a legal AI company.
- Gives no human review point, which is especially risky for a company whose product claims legal AI reliability.
Likely follow-up questions
- How would you handle a draft that oversells Harvey's AI accuracy claims.
- What would make you pull a workflow back from production use.
More metrics questions
- What metrics prove Harvey's value to a firm like PwC or Paul Weiss?Harvey · Metrics · Hard
- What KPI hierarchy would you use for Vault, from account-level adoption and active matters to search success, document coverage, and workflow outcomes, to measure value for law firms and enterprises? How would those metrics change your roadmap if usage is high but repeat usage in critical workflows is low?Harvey · Metrics · Hard
- What metrics would you define for Command Center across adoption, admin efficiency, and governance coverage, and how would you use those metrics to decide whether the product is actually improving Harvey’s enterprise land-and-expand motion?Harvey · Metrics · Hard
- Harvey is considering expanding its platform to additional regions. How would you decide whether multi-region expansion is the right product investment now versus later? Explain the customer signals, business considerations, and success metrics you would use, and how you would compare this investment against competing platform priorities.Harvey · Metrics · Hard
- For Vault’s search and Q&A experience, what metrics would you define for adoption, engagement, trust, and business value at the workspace and firm level? Which leading indicators would you watch in the first 90 days, and how would changes in those metrics alter your product decisions?Harvey · Metrics · Hard
- A pilot agent deployment has strong qualitative feedback from senior stakeholders, but end-user adoption is inconsistent and the customer’s data environment is messy. How would you structure the deployment plan, define the success metrics, diagnose whether the issue is workflow fit vs data readiness vs change management, and decide if this should graduate into a repeatable product for the broader vertical?Harvey · Metrics · Hard
More questions from Harvey
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop