Estimation question

Estimate the cost advantage of Groq's LPU vs. NVIDIA H100 for a high-volume workload.

Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.

Start a mock interview on this question · Mock interview from a job description

What this question tests

Structured cost-comparison estimation across two different hardware architectures for a defined workload.

How to approach it

  1. Define the workload: assume a high-volume inference task, say serving 1 billion tokens per day for a mid-size language model.
  2. Estimate H100 throughput and cost: assume an H100 processes roughly 2,000 to 4,000 tokens per second for this model size, at a cloud rental cost of roughly 2 to 4 dollars per GPU-hour.
  3. Estimate LPU throughput and cost: assume Groq's LPU delivers a meaningfully higher tokens-per-second rate for the same model, say 3 to 5 times higher, due to its architecture, at a comparable or somewhat higher per-chip hourly cost.
  4. Compute cost per million tokens for each: divide hourly cost by tokens processed per hour, showing the LPU's higher throughput can offset its potentially higher hourly rate, yielding a lower cost per token overall.
  5. State the conclusion as a range: assume the LPU comes out roughly 30 to 50 percent cheaper per token at high volume, clearly marked as an illustrative estimate, not a verified figure.

What a strong answer includes

Common mistakes

Likely follow-up questions

More estimation questions

More questions from Groq

Learn the skill behind it

Chapters of the AI PM course that teach what this question tests.

Preparing for a specific role?

Book summaries for this kind of question

Browse all 4,000+ questions in the bank