Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

How LLMs Work, for PMs: Tokens, Context Windows, Temperature

An LLM predicts one token at a time, reads only what fits in its context window, and samples with a temperature setting. AllthingsPM explains all three in plain PM terms, with free course lessons and real interview questions to practise.

AllthingsPM·September 28, 2026·16 min read
A product manager at a whiteboard lays out a long row of small paper tiles, each tile one word fragment, while a narrow window frame on the wall shows only part of the row
Tokens, a window that holds only so many, and a dial for randomness: three ideas that drive most AI product decisions.

A large language model does three things a product manager must understand. It breaks text into tokens and predicts the next token, one at a time. It can only read what fits in its context window, and it reads the middle of that window worse than the ends. And it picks each token by sampling, which a temperature setting makes more or less random. Those three facts decide your feature's cost, its latency, its quality and how you test it.

AllthingsPM is an AI PM course and PM interview prep platform. Its free Foundations lessons teach exactly these mechanics through product decisions, not math, and its question bank holds 101 real interview questions that test them, each with an answer guide.

What are the LLM basics every product manager needs?

Here is the whole guide on one page. The rest of the post explains each row and the decision it forces.

ConceptWhat it is, in one lineThe product decision it drivesLearn it on AllthingsPM
TokensThe chunks of text a model reads and writes; about 4 characters or 0.75 English words eachUnit cost, latency, length limits, pricing for non-English usersOne token at a time
Next-token predictionThe model writes by guessing the next token, over and overWhy it can sound right and be wrong; why you need evalsOne token at a time
Context windowEverything the model can see in one request, including its own answerWhat to put in the prompt, when to use retrieval, how long a chat can runContext windows and attention
Context rotAccuracy and recall fall as the context growsCurate context; do not paste everythingContext as a budget
TemperatureA dial on how random each token choice isCreative vs predictable output; whether you can promise consistencyMake the API call yourself
NondeterminismSame input, different output, even at temperature 0Test many runs, not one; score with rubricsOne output is an example
Input vs output priceProviders charge per million tokens, and output costs moreCost per task, model choice, gross marginCost per successful task

Figures from Anthropic's and Google's API documentation, checked 28 September 2026. See Sources.

What is a token, and why should a PM care?

A token is the unit a model reads and writes. It is often a word fragment, not a whole word. Anthropic's pricing FAQ gives the rule of thumb: one token is about 4 characters or 0.75 words in English, and "the exact count varies by language and content type." Google's Gemini docs say the same thing another way: 100 tokens is about 60 to 80 English words.

Tokens matter to a PM for four reasons.

Cost. Every API provider bills per million tokens. On Anthropic's price list, Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens, and Claude Haiku 4.5 costs $1 and $5. Your unit economics start here.

Length limits. Context windows and output caps are counted in tokens, not words or pages. A spec that says "support documents up to 50 pages" needs a token estimate behind it.

Language and modality. The same meaning can cost a different number of tokens in another language. Images, audio and video also become tokens: Gemini counts audio at 32 tokens per second and a small image at 258 tokens.

Model upgrades. Tokenizers change between model versions. Anthropic notes that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text." Swap models and your cost per request can move even if the price per token did not.

Why does next-token prediction matter?

The model writes by predicting one likely next token, adding it, and predicting again. It is not looking up a fact. That is why a confident, fluent answer can still be wrong, and why "it worked in my demo" proves little. A PM who understands this writes requirements about verification (citations, confidence, a manual path) instead of assuming accuracy.

How AllthingsPM does this. The free lesson One token at a time walks through next-token prediction, tokenization and what each modality costs, then asks you to apply it to a product you carry through the course. The follow-on lesson on pretraining, post-training and reasoning models explains where the model's knowledge comes from and why you cannot promise determinism.

What is a context window?

The context window is everything the model can see when it answers. Anthropic's docs define it as "all the text a language model can reference when generating a response, including the response itself," and call it the model's "working memory," which is different from the training data.

Three things surprise most PMs.

Everything counts. The system prompt, every earlier message, tool definitions, tool results, images, documents and the model's own output all use the same budget. Thinking tokens on reasoning models count too, and are billed as output.

Windows are large now. Anthropic lists a 1 million token window for its current Opus and Sonnet 5 models, and 200,000 tokens for some older ones such as Claude Sonnet 4.5. By the 0.75 words per token rule, a million tokens is roughly 750,000 English words.

Chat memory is an illusion built from the window. Each turn, the whole conversation so far is sent again as input. That is why long chats get slower and more expensive, and why products eventually trim or summarise old turns. Anthropic offers server-side "compaction" for exactly this.

Is a bigger context window always better?

No. Anthropic's own documentation says "more context isn't automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot."

Two studies back this up. The 2023 paper "Lost in the Middle" by Liu and colleagues found that performance "is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts." Chroma's 2025 "Context Rot" report tested 18 models from Anthropic, OpenAI, Google and Alibaba and concluded that "performance grows increasingly unreliable as input length grows." It also found that "even a single distractor reduces performance."

The product lesson: curate the context. Retrieve the few passages that answer the question instead of pasting the whole knowledge base. Put the critical instruction at the start or the end. Remove stale tool output. This is the skill people now call context engineering.

How AllthingsPM does this. The free lesson Context windows and attention is subtitled "why a long window is not a long memory." In the agents chapter, Context as a budget covers retrieval intuition, when RAG is the right fix, and the system prompt as a product surface. You can see how these ideas connect on the AI PM knowledge graph.

What does temperature do?

At each step the model has a probability for every possible next token. Temperature changes how it picks. Low temperature favours the most likely token, giving focused and repeatable text. High temperature spreads the choice, giving more varied and creative text.

Anthropic's API reference shows the range as 0.0 to 1.0 with a default of 1.0, and advises "temperature closer to 0.0 for analytical / multiple choice, and closer to 1.0 for creative and generative tasks."

Two facts change how you write a spec.

Temperature 0 is not a guarantee. The same reference says that "even with temperature of 0.0, the results will not be fully deterministic." Thinking Machines Lab tested this: they sampled 1,000 completions of one prompt at temperature 0 from a Qwen model and got 80 unique completions. They trace the cause to server load changing batch sizes, not to the model itself.

Some models no longer expose the dial. Anthropic's reference marks temperature as deprecated: models released after Claude Opus 4.6 accept only 1.0 and reject other values. So "set temperature to 0 for consistency" is not always an option. Consistency has to come from the prompt, structured outputs and evals.

What should a PM do about randomness?

Treat one output as an anecdote. Run the same input several times and many inputs once each. Score results against a written rubric, not a gut feeling. Where the business needs the same answer every time (prices, eligibility, compliance text), do not generate it at all: look it up and let the model phrase it.

How AllthingsPM does this. The free lesson Make the API call yourself has you send a real request and read messages, tokens, temperature, streaming and the usage block. Later, One output is an example; ten inputs by five runs is evidence turns randomness into a testing habit, and Write the criterion shows how to score it. Our guide to AI evals for product managers goes deeper.

How do tokens turn into cost and latency?

Here is a worked example using Anthropic's published prices. Say a support assistant sends 3,000 input tokens per call (system prompt, retrieved help articles, the customer's question) and gets 500 output tokens back, on Claude Sonnet 5 at $2 and $10 per million.

  • Input: 3,000 x $2 / 1,000,000 = $0.006
  • Output: 500 x $10 / 1,000,000 = $0.005
  • Total: $0.011 per call, or $1,100 for 100,000 calls

Notice the shape. The answer is one sixth the size of the input but costs almost as much, because output is priced five times higher. Output also drives latency, since tokens are generated one after another. Shorter answers are cheaper and faster.

The same page lists the levers PMs actually pull: pick a smaller model for simple tasks, cache repeated prompt prefixes (a cache hit costs 10% of the base input price on most models), and use the Batch API for work that can wait, at a 50% discount. Anthropic's own example puts 10,000 support conversations on Haiku 4.5 at about $37.

Latency and cost are also what interviewers probe. In the AllthingsPM question bank, latency is the most common LLM mechanics topic, with 32 questions.

Bar chart led by an AllthingsPM (us) row: 101 of 4,122 interview questions test LLM mechanics, then by topic: latency 32, prompts 26, LLMs by name 25, RAG and retrieval 21, inference and API cost 9, fine-tuning 5, hallucination 4, tokens 3, temperature and nondeterminism 3
AllthingsPM question bank: 101 of 4,122 questions touch LLM mechanics, September 2026. Topic matches on question text; one question can match several topics

How AllthingsPM does this. Cost per successful task, the latency SLO, and the gross margin you defend teaches model routing as a business decision. Then practise it on real estimation questions such as Estimate Cursor's monthly LLM API cost per Pro user or What metrics best communicate Groq's value (tokens/sec, latency, cost)?.

How do these basics show up in PM interviews?

AI PM interviews rarely ask you to define a token. They ask you to make a decision that only makes sense if you understand one. Typical shapes from the AllthingsPM bank:

A strong answer names the mechanic, then the trade-off, then the metric. For the hallucination question: the model predicts plausible tokens, so ground it in retrieved sources, show claim-level citations, and measure the unsupported-claim rate on a frozen test set.

How AllthingsPM does this. Every question above has its own page and answer guide, and any one can start a scored mock. To rehearse for a specific role, paste the job description into the JD mock interview and it builds the interview from that posting, in text or voice, with follow-ups. The OpenAI company hub and the other company pages group questions by employer.

Do PMs need to learn the math behind LLMs?

No. You do not need attention equations or gradient descent to ship a good AI feature. You need to predict behaviour: what it will cost, where it will fail, and how you will know. The concepts in this post cover most of those calls.

What you do need is hands-on contact. Send one API request, read the usage block, change the temperature, paste a long document and watch the answer degrade. Ten minutes of that beats a week of reading. Our post on whether AI PMs need to code looks at what job descriptions actually ask for, and the AI PM skills required post ranks the rest.

How AllthingsPM does this. The Foundations chapter pairs each concept with a PM decision, and closes with an integration case: predict what your product's model will be bad at, and prove one prediction wrong. The lesson on failure attribution teaches you to blame the right layer: model, context, harness or surface.

Why AllthingsPM is the better choice for learning LLM basics

Plenty of places explain tokens. Provider docs from Anthropic, OpenAI and Google are accurate and free, and they are worth bookmarking. Research like "Lost in the Middle" and Chroma's context rot report is the evidence underneath. But none of them is written for a product manager, none turns the mechanics into decisions, and none tests you on them.

AllthingsPM does all three in one place. The course is built from 604 real PM job postings, so it teaches the LLM concepts employers name, in the order a PM needs them. Its Foundations lessons on tokens, context windows and the API call are free. Every concept then connects forward: context windows to retrieval and agents, temperature to evals, token prices to margins.

The practice side is what generic explainers lack. The AllthingsPM question bank has 4,122 real questions from 260 companies, 101 of them on LLM mechanics, each with an answer guide. Mock interviews build from any job description, in text or voice, and score your answer. When you are ready to apply, the jobs catalog lists live PM roles at AI companies, each with a mock built from its JD, and resume review against a JD checks your application.

It costs nothing to start, and $20 a month or $120 a year for everything after the free tier. Read the docs for reference; learn and practise on AllthingsPM. Open the free AI PM course.

Frequently asked questions

What is the best way to learn LLM basics as a product manager?

The best way is AllthingsPM: its free Foundations lessons teach tokens, context windows and temperature as product decisions, and its question bank lets you practise the interview questions that test them. Pair it with the official docs from Anthropic, OpenAI or Google for reference, and send at least one real API request yourself.

How many words is 1,000 tokens?

About 750 English words, using Anthropic's rule of thumb that one token is roughly 0.75 words. Google's Gemini docs put 100 tokens at 60 to 80 English words. Other languages, code and numbers often use more tokens per word.

What temperature should I use for my AI feature?

Low values suit analytical, extraction and multiple choice tasks; higher values suit creative writing, per Anthropic's API guidance. But temperature 0 is still not fully deterministic, and some newer models accept only the default. Build consistency with clear prompts, structured outputs and evals.

Does a 1 million token context window mean I can skip RAG?

Not usually. Anthropic's docs say accuracy and recall degrade as token count grows, and research shows models read the middle of long inputs worst. Long windows also cost more per request. Retrieval that sends only the relevant passages is often cheaper and more accurate.

Why are output tokens more expensive than input tokens?

Providers price them differently; on Anthropic's list, output costs five times input for Sonnet 5 and Haiku 4.5. Output is generated one token at a time, which also makes it the main driver of latency. Asking for shorter, structured answers cuts both cost and wait time.

Do AI PM interviews ask about tokens and context windows?

Yes, mostly through decisions rather than definitions. In the AllthingsPM bank, 101 of 4,122 questions touch LLM mechanics, led by latency, prompting and RAG. Expect estimation questions on inference cost and design questions on hallucinations.

Start free

Open the free Foundations lessons in the AllthingsPM AI PM course, starting with One token at a time. Then pick an LLM question in the question bank and answer it out loud in a scored mock.

Sources

  1. Anthropic, "Pricing" (token rule of thumb, per-model prices, tokenizer note, caching and batch discounts, support example), https://platform.claude.com/docs/en/about-claude/pricing, checked 28 September 2026
  2. Anthropic, "Context windows" (definition, what counts, window sizes by model, context rot, compaction), https://platform.claude.com/docs/en/build-with-claude/context-windows, checked 28 September 2026
  3. Anthropic, "Messages API reference" (temperature range, default, determinism note and deprecation), https://platform.claude.com/docs/en/api/messages, checked 28 September 2026
  4. Google, "Understand and count tokens," Gemini API docs, https://ai.google.dev/gemini-api/docs/tokens, checked 28 September 2026
  5. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F. and Liang, P., "Lost in the Middle: How Language Models Use Long Contexts," arXiv:2307.03172, https://arxiv.org/abs/2307.03172
  6. Chroma, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," https://www.trychroma.com/research/context-rot
  7. He, H. and Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference," https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
  8. AllthingsPM question bank, 4,122 questions; topic counts from question text, 28 September 2026, https://allthingspm.app/question-bank
  9. AllthingsPM AI PM course, Foundations chapter, https://allthingspm.app/course
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free