Estimation question

Estimate the gross margin on serverless inference vs. renting GPU clusters.

Practice this question out loud. An AI interviewer asks it, follows up like a real interviewer would, and scores your answer. Type or speak.

Start a mock interview on this question · Mock interview from a job description

What this question tests

Structured cost and margin estimation comparing two different infrastructure business models.

How to approach it

  1. Define both models: serverless inference bills per token with shared, multi-tenant GPU utilization, while dedicated clusters bill a flat reserved rate regardless of utilization.
  2. Estimate serverless margin: assume GPU cost of roughly 2 dollars per hour, serving many customers' requests on shared capacity at high utilization, say 70 percent, yielding a gross margin in the range of 60 to 70 percent.
  3. Estimate dedicated cluster margin: assume the same GPU cost but a customer often runs it at lower average utilization, say 40 percent, since they pay for guaranteed capacity whether or not they use it fully, giving Together a similar or better margin on the same hardware.
  4. Compare: dedicated clusters can have higher or comparable margin per GPU-hour sold because the customer, not Together, absorbs the underutilization risk, while serverless margin depends heavily on achieving high shared utilization.
  5. Conclude with the key driver: utilization rate is the single biggest swing factor for serverless margin, while committed contract length is the biggest factor for dedicated margin.

What a strong answer includes

Common mistakes

Likely follow-up questions

More estimation questions

More questions from Together AI

Learn the skill behind it

Chapters of the AI PM course that teach what this question tests.

Preparing for a specific role?

Book summaries for this kind of question

Browse all 4,000+ questions in the bank