Metrics question
Scale can invest one quarter either in a demand-side workflow improvement that helps customers create and evaluate tasks faster, or in a supply-side tooling improvement that helps contributors complete more high-quality work. How would you decide between them? Walk through the decision framework, the marketplace and financial data you'd examine, and how you'd compare near-term revenue impact vs. long-term marketplace health.
- Scale AI
- Metrics
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Marketplace strategy judgment: can you weigh a demand-side versus supply-side investment using real marketplace and financial signals, not just intuition about which side feels more urgent.
How to approach it
- Frame it as a marketplace balance question first: is throughput currently constrained by customers not creating enough well-specified tasks, or by contributors not being able to complete enough high-quality work.
- Examine marketplace data: task backlog size and age (a growing backlog signals supply constraint), contributor idle time or task rejection rate (signals demand-side task quality issues), and utilization on both sides.
- Examine financial data: near-term revenue impact of unblocking demand (faster task creation drives immediate billable throughput) versus the compounding value of supply-side tooling (better contributor tools improve quality and retention over many future engagements).
- Decide using a simple rule: invest demand-side if backlog is thin and contributors are underutilized, invest supply-side if backlog is deep and quality or contributor churn is the binding constraint.
- Communicate the choice with the evidence, not intuition, since a wrong call in either direction either starves the marketplace of tasks or burns out the contributor base.
What a strong answer includes
- Uses backlog depth and utilization as the concrete diagnostic for which side is actually constrained, rather than assuming based on which team complains more.
- Explicitly separates near-term revenue impact from long-term marketplace health as two different kinds of value, matching what the question asks to compare.
- Gives a clear decision rule tied to the diagnostic data rather than a generic pros-and-cons list.
Common mistakes
- Picks a side based on which team is louder rather than examining backlog and utilization data.
- Compares near-term revenue and long-term health only qualitatively with no data-backed reasoning.
Likely follow-up questions
- What would you do if the data shows both sides are constrained simultaneously.
- How would you measure whether the supply-side investment actually reduced contributor churn.
More metrics questions
- Scale is considering a new evaluation product for enterprise customers to assess model quality before deployment. How would you choose the first customer use case to support, scope the MVP, and define the launch metrics that would tell you whether to expand the product or shut it down?Scale AI · Metrics · Hard
- For contributor engagement and retention across 500,000+ contributors in 100+ countries, what are the few core metrics you would instrument for activation, repeat participation, and churn? If weekly supply health suddenly dropped, how would you determine whether the root cause was demand mix, onboarding friction, pay, quality gating, or country-specific issues?Scale AI · Metrics · Hard
- The first version of the cybersecurity evaluation suite is in market. What metrics would you track to know whether it is actually helping frontier labs and enterprises measure real security capability rather than benchmark gaming? Separate product adoption metrics from benchmark quality metrics, and explain how each would change your roadmap.Scale AI · Metrics · Hard
- Forward-deployed teams say they are rebuilding too much plumbing on each enterprise deployment. How would you identify the highest-leverage platform blockers, distinguish anecdote from systemic friction, and choose the few metrics you would track to prove the platform is improving time-to-value, reuse, and production reliability?Scale AI · Metrics · Hard
- One enterprise account has launched to production, but adoption and measurable value are uneven across teams. How would you diagnose where the deployment is truly working, decide whether expansion is justified, and avoid confusing executive enthusiasm with real customer value?Scale AI · Metrics · Hard
- You own the multi-turn chat tasking experience used by contributors to generate training and evaluation data. What changes would you make to increase throughput by 20% without degrading quality? Explain which parts of the workflow you would redesign, the key failure modes you would watch for, and how you would validate that faster tasking still produces data customers can trust.Scale AI · Metrics · Hard
More questions from Scale AI
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 14: Get the job: the AI PM interview loop