Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

Unit economics of AI features: cost per task

An AI feature's real cost is its cost per successful task: tokens per attempt, times attempts, plus the cost of every failure, divided by the success rate. This guide shows the math with September 2026 model prices, and AllthingsPM teaches it in a graded course chapter.

AllthingsPM·September 29, 2026·17 min read
A product manager at a workbench laying out a small set of tools in a row: a laptop, a plug adapter, a notebook of tables and a coffee mug
Every AI answer comes with a receipt. The PM decides whether the receipt is worth it.

Short answer: the cost of an AI feature is not the price per million tokens. It is the cost per successful task: what one attempt costs in tokens and tools, times the attempts it takes, plus what every failure costs you (a retry, a human fix, a refund), divided by how often the feature actually succeeds. A cheaper model that fails more often can cost more per task than a pricier one. AllthingsPM is an AI PM course and PM interview prep platform, and its Prove it paid off chapter teaches exactly this math, with graded case studies, for $120 a year with a free tier.

Below: the formula, a worked example with real September 2026 API prices, the levers that move the number, and how to defend it in an interview.

What does an AI feature actually cost per task?

Start with one attempt. A model call is billed on input tokens (everything you send: system prompt, instructions, retrieved documents, conversation history, tool definitions) and output tokens (everything the model writes, including reasoning on reasoning models). Output tokens cost more. On Anthropic's price list, Claude Sonnet 5.5 is $2 per million input tokens and $10 per million output tokens, a 5x gap, and OpenAI's gpt-6-sol is priced the same way at $2 and $10 [1][2].

Then add everything around the call:

  • Tool calls. Tool definitions are input tokens on every request, and Anthropic adds a tool use system prompt of 286 tokens on Sonnet 5.5 whenever tools are present. Server tools can bill separately: web search is $10 per 1,000 searches [1].
  • Retries and multi-step runs. An agent that loops five times pays for five calls, and each call re-reads the growing history.
  • Guardrails and graders. A safety classifier or an LLM judge in the request path is a second model bill on every task.
  • Failures. A wrong answer is not free. It costs a retry, a support ticket, a human who rewrites the draft, or a churned customer.

So the formula a PM should own is:

The last line matters. A blended "our OpenAI bill was $40,000" tells you nothing about which feature is losing money.

How AllthingsPM does this: the lesson Cost per successful task, the latency SLO, and the gross margin you defend walks through this exact build, and the free Foundations lesson on the usage block in an API response shows where the token counts come from. You do the math on your own carried product in each chapter's graded case study.

A worked example: pricing a support reply drafter

Here is one feature, costed three ways. The feature drafts a reply to a customer support ticket. The token counts are assumptions for illustration; the prices are Anthropic's published list prices, checked September 29, 2026 [1].

Assumed tokens per attempt: a 4,000 token system prompt with policies and tone rules (identical on every call, so cacheable), 1,500 tokens of ticket and history, 2,500 tokens of retrieved help articles, and a 500 token reply. That is 8,000 input tokens and 500 output tokens.

Model (price per million tokens, input / output)Cost per attempt, no cachingCost per attempt, 4,000 token prefix cached
Claude Haiku 4.5 ($1 / $5)$0.0105$0.0069
Claude Sonnet 5.5 ($2 / $10)$0.0210$0.0138
Claude Opus 5.5 ($4 / $20)$0.0420$0.0268

Prices from Anthropic's pricing page, checked September 29, 2026. Cache reads cost 0.1x the base input price on Haiku and Sonnet and 0.05x on Opus 5.5 [1]. Cache write costs are ignored here because one write serves many reads.

Two cents a reply looks trivial. It stops looking trivial once you divide by success.

Now add assumed quality. Say, again for illustration, that support agents send the Sonnet draft with light edits 80% of the time and the Haiku draft 55% of the time, and that a rejected draft costs about $2 of agent time to rewrite from scratch. Then:

  • Sonnet, cached: $0.0138 + 0.20 x $2 = about $0.41 per ticket
  • Haiku, cached: $0.0069 + 0.45 x $2 = about $0.91 per ticket

The model that is half the price per token is more than twice the price per ticket. That is the whole point of unit economics for AI features: the token bill is usually the small number, and the failure cost is the large one. Your job is to measure the success rate with an eval, not guess it. Our guide to AI evals for product managers covers how.

How AllthingsPM does this: the Evals chapter teaches you to get the success rate from a golden dataset, and its lesson on paying for the cheapest evaluator that can see the failure applies the same cost logic to the grader itself. The two chapters are built to be used together.

Which levers cut AI feature cost the most?

Once you know cost per successful task, you can move it. In rough order of effort:

1. Cache the stable part of the prompt

If the start of every request is identical (system prompt, policies, product docs), prompt caching bills it at a fraction of the input price after the first write. On Anthropic, a cache hit is 0.1x the base input price, a 5 minute cache write is 1.25x, and caching pays off after a single read [1]. OpenAI lists cached input for gpt-6-sol at $0.20 against $2 standard [2]. In the example above, caching cut Sonnet's cost per attempt by about a third. The product decision is prompt order: put stable content first and user content last.

2. Route the easy majority to a cheaper model

Not every request needs the frontier model. A router or cascade sends simple cases to a small model and escalates hard ones. The spread is large: OpenAI's gpt-6-luna lists at $0.10 input and $0.50 output per million tokens against $10 and $50 for gpt-6-astra, a 100x difference [2]. The catch is the example above: route only where your evals show the cheap model still succeeds, and own the escalation rate as a metric.

3. Batch anything that can wait

Nightly summaries, backfills, tagging and eval runs do not need an instant answer. Both Anthropic and OpenAI discount batch processing by 50% on input and output [1][2]. Anthropic notes that batch and caching discounts stack [1].

4. Shrink what you send and what you ask for

Trim retrieved context to what the answer needs, cap output length, and use structured outputs so the model does not pad. Every retrieved chunk you send is paid for on every call.

5. Watch the hidden multipliers

Three are easy to miss. First, tokenizers change: Anthropic says Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text, so a model switch can raise cost at the same list price [1]. Second, data residency: US-only inference on Claude 4.6 and later models carries a 1.1x multiplier [1]. Third, tool schemas: a large browser toolset adds about 6,600 input tokens to each request on Anthropic [1].

How AllthingsPM does this: the course lesson on routing, batching and cost per successful task turns these levers into a decision you defend, and Context as a budget covers what to cut from the prompt. You can see how caching, routing and evals connect in the AI PM knowledge graph.

Why do cheaper tokens not mean cheaper AI features?

Model prices keep falling. Andreessen Horowitz measured that, for a model of equivalent capability, inference cost has been dropping about 10x per year, and called it "LLMflation": a GPT-3 level result that cost $60 per million tokens in 2021 cost $0.06 when they wrote in November 2024 [3].

But usage per task is rising faster in many products. Ethan Ding argues that as models handle longer tasks, "what used to return 1,000 tokens is now returning 100,000," so flat subscriptions get squeezed even as unit prices fall [4]. Agents are the clearest case: one user request can trigger dozens of calls, each re-reading a growing context. Anthropic's own worked example for its managed agents prices a one hour coding session at 50,000 input and 15,000 output tokens on Opus 5 at about $0.71, or about $0.53 with caching [1].

This shows up in company margins. Bessemer's State of AI 2025 report put the fastest growing AI "Supernovas" at an average gross margin of about 25%, "often negative," and the "Shooting Stars" cohort at about 60% [5]. Classic software businesses are usually judged against far higher margins, which is why investors and finance teams now ask PMs for per feature cost.

How AllthingsPM does this: the agents chapter has a lesson on when multi-agent is worth its cost, and our explainer on agents vs workflows shows why a workflow is often the cheaper design for the same outcome.

How do you turn cost per task into price and margin?

Cost per successful task is the input to three decisions.

Gross margin per user. Multiply cost per successful task by tasks per user per month and compare it with what that user pays. Using the illustrative Sonnet figure of $0.0138 per attempt, a user who triggers 600 drafts a month costs about $8.28 in model tokens before any failure costs. Heavy users are where margin breaks, so look at the distribution, not the average.

Pricing model. Seat pricing is simple but exposes you to heavy users. Usage or credit pricing tracks cost but can cause bill shock. Outcome pricing (charge per resolved ticket) matches your cost per successful task directly, but you carry the failure risk. Real interview questions test exactly this, such as how you would redesign Cursor's credit based pricing or defend Sierra's outcome based pricing.

The ship call. If the feature cannot reach a target cost per successful task at an acceptable success rate, you change the scope, the model mix or the price, or you do not ship. Writing that down before launch is what makes an AI PM credible with finance.

How AllthingsPM does this: the lessons Price it: seat, usage, and outcome and The business case, and the ship or no ship call finish the chapter by turning your cost model into a pricing recommendation and a launch decision you would defend to leadership.

How do interviewers test AI feature cost?

AI PM interviews now include cost estimation and pricing questions, especially at AI companies. The AllthingsPM question bank has real examples, each with its own page and answer guide:

A strong answer follows the same structure as this guide. State the task. Estimate tokens in and out per attempt, and say where they come from. Apply a named price. Multiply by attempts and volume. Then say what failure costs and how you would measure the success rate. Finish with the lever you would pull first. Interviewers are grading whether you separate cost per call from cost per successful task.

How AllthingsPM does this: you can practise any of these questions in an AI mock interview that follows up on your assumptions and scores the answer, or paste an AI company job description into the JD mock to get the cost and pricing questions that role is likely to ask. Company pages such as Anthropic interview questions group them by employer.

What should a PM track on a cost dashboard?

Keep it to a handful of numbers per AI feature:

MetricWhy it matters
Cost per successful taskThe number you defend; combines price, volume and quality
Success rate (from evals and user accepts)The denominator; small changes swing cost a lot
Average attempts and tool calls per taskShows agent loops and retries creeping up
Cache hit rateTells you whether the stable prefix is really stable
Share of traffic per model tierTells you whether routing is working
p95 cost per task and top 5% of usersWhere margin actually breaks

Review it with every model change, prompt change or tokenizer change, because each one can move the number with no code change on your side.

How AllthingsPM does this: the course treats cost as a product metric next to quality and latency, which is why the same chapter covers the latency SLO. If you are building proof for a job search, a cost dashboard for a feature you built is a strong portfolio piece.

Bar chart of list prices in US dollars with AllthingsPM (us) first and highlighted: one month $20, one year $120; then the AI PM Bootcamp on Maven $2,500 and the Product Faculty AI PM Certification $5,000
AllthingsPM teaches AI feature cost inside a $120 a year course. Sources: allthingspm.app/pricing and Maven course pages, checked September 28, 2026

Why AllthingsPM is the better choice for learning AI feature cost

Most material on AI costs is written for engineers: provider pricing pages, caching docs, infrastructure blogs. They are accurate and worth reading, but they stop at the token. A PM needs the next step, from the token bill to the success rate, the price and the ship decision, and then needs to say it clearly in an interview.

AllthingsPM puts all of that in one place. The AI PM course has a full chapter on outcomes, economics and pricing, next to chapters on evals and agents that feed its numbers, and every chapter ends in a graded case study on a product you carry through the course. The curriculum is built from 604 real PM job postings and updated weekly, so it follows what AI companies ask for. You then practise with real cost and pricing questions from the question bank, rehearse in AI mock interviews built from any job description, and browse live AI PM roles that each come with their own mock.

Live cohorts such as the AI PM Bootcamp on Maven offer instructors and a peer group, which some people value. For learning this skill and practising it every day, AllthingsPM costs $120 a year against $2,500 for one bootcamp cohort, with a free tier to start. Open the Prove it paid off chapter and work through the formula on your own feature.

Frequently asked questions

How much does an AI feature cost to run?

It depends on tokens per task, the model, attempts per task and the success rate. In our illustrative support drafter, one attempt cost between about $0.007 and $0.042 depending on model and caching, using Anthropic's September 2026 prices. Once you add the cost of failed drafts, the cost per successful ticket was far higher, which is why cost per successful task is the number to track.

What is cost per successful task?

It is the total cost of getting one task done correctly: tokens and tools for every attempt, plus the cost of failures such as retries or human fixes, spread over the tasks that succeed. It lets you compare models and designs fairly, because a cheap model that fails often can cost more per task.

Are output tokens more expensive than input tokens?

Yes, on current price lists. Claude Sonnet 5.5 and OpenAI's gpt-6-sol both list output at $10 per million tokens against $2 for input, a 5x gap [1][2]. Long answers, verbose reasoning and chatty agents raise cost quickly.

How do I reduce the cost of an AI feature?

Cache the stable prompt prefix, route easy requests to a cheaper model, batch work that can wait (50% off at Anthropic and OpenAI), trim the context you send, and cap output length. Check every change against your evals so cost savings do not lower the success rate.

What is the best way to learn AI feature unit economics as a PM?

AllthingsPM is the best fit for most PMs: its AI PM course has a full chapter on cost per successful task, routing, pricing and the business case, with graded case studies, plus real cost estimation interview questions and AI mock interviews, for $120 a year with a free tier. Provider pricing docs are a useful reference alongside it.

Do AI interviews ask about inference cost?

Yes, especially at AI companies. The AllthingsPM question bank includes real questions such as estimating ChatGPT's free tier inference cost or Cursor's LLM cost per Pro user, each with an answer guide you can practise in a mock interview.

Ready to put numbers on your next AI feature? Start the AllthingsPM AI PM course free and work through the Prove it paid off chapter.

Sources

  1. Anthropic, Pricing (Claude API docs) (model prices, prompt caching multipliers, batch discount, tokenizer note, data residency multiplier, tool token overhead, web search price, managed agents worked example; checked September 29, 2026)
  2. OpenAI, API pricing (gpt-6 model prices, cached input, 50% batch discount; checked September 29, 2026)
  3. Andreessen Horowitz, "Welcome to LLMflation: LLM inference cost is going down fast" (10x per year decline, $60 to $0.06 per million tokens; November 2024)
  4. Ethan Ding, "tokens are getting more expensive" (token use per task rising, flat subscription squeeze)
  5. Bessemer Venture Partners, The State of AI 2025 (Supernova gross margin about 25%, often negative; Shooting Stars about 60%)
  6. Maven, AI Product Management Bootcamp by Marily Nika ($2,500 cohort price; checked September 28, 2026)
  7. Maven, Product Faculty AI Product Management Certification ($5,000 price; checked September 28, 2026)
  8. AllthingsPM pricing ($20 a month, $120 a year, free tier)
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free