Estimation question
Estimate the number of actual COVID-19 positive people in India as of today.
- Estimation
- Hard
Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.
Start a mock interview on this question · Mock interview from a job description
What this question tests
Estimation under high uncertainty: can you reason from testing rate and known under-detection factors to build a defensible range rather than a false-precision point estimate.
How to approach it
- State the core challenge: official confirmed case counts vastly undercount actual infections due to limited testing capacity and asymptomatic cases.
- Start from reported confirmed cases (a number that would be publicly known at the time) as the floor of the estimate.
- Apply an under-detection multiplier based on general epidemiological patterns (undercounting factors of several times to an order of magnitude are common for under-tested populations, stated as a pattern not a specific cited statistic).
- Adjust the multiplier based on India-specific factors: large population, limited testing infrastructure per capita, and significant rural/informal population less likely to be tested.
- Present a range rather than a single number, given the wide uncertainty in the multiplier, and state the reported case count and multiplier assumption clearly.
- State what data would tighten the estimate: seroprevalence survey results, which directly measure past infection rates in a sample population.
What a strong answer includes
- Explicitly distinguishes confirmed cases from actual infections and explains why the gap exists (testing capacity, asymptomatic spread).
- Uses an under-detection multiplier as the core estimation mechanism, a standard, defensible approach for this type of problem.
- Adjusts the multiplier for India-specific context (population size, testing infrastructure, rural access) instead of applying a generic global assumption.
- Presents a range and names seroprevalence survey data as the best real-world data source to validate or replace the estimate, showing methodological awareness.
Common mistakes
- Treating official confirmed case numbers as the actual answer without adjusting for known under-detection.
- Presenting a single falsely-precise number instead of a defensible range given the genuinely high uncertainty.
Likely follow-up questions
- What real-world data source would most improve the accuracy of this estimate?
- How would your multiplier assumption differ between an urban and a rural state?
More estimation questions
- YouTube Red is a premium service without ads. Assume that 1.5% of initial YouTube user base signs up for the service. What is the lifetime revenue Google generates from those users?Google · Estimation · Hard
- How would you design a service like Instagram? Estimate server and storage requirement for peak traffic.Google · Estimation · Hard
- How much storage is required to store all of the images on Google Maps?Google · Estimation · Hard
- A company has launched a new drug that eliminates the need for sleep. How would you price this drug?Google · Estimation · Hard
- How much storage space is required to host all the images of Google Street View?Google · Estimation · Hard
- Imagine you created a new type of product that would replace the mobile phone. How would you determine how many units to manufacture?Google · Estimation · Hard
Learn the skill behind it
Chapters of the AI PM course that teach what this question tests.
- Chapter 2: Data fluency: SQL, logs, and reading the truth yourself
- Chapter 9: Prove it paid off: outcomes, economics, and pricing
- Chapter 14: Get the job: the AI PM interview loop