Flash sale 30% off with code LAUNCH30 Ends in --:--:--
See pricing
All Things PM

Context Windows and Context Rot: Why Long Context Is Not Memory

Context rot is the drop in accuracy and recall as a model's input grows, even far below the context window limit. AllthingsPM teaches PMs to treat context as a budget, not a memory, in a free course lesson.

AllthingsPM·September 29, 2026·16 min read
A product manager at a whiteboard connects one universal plug to a row of different devices, a laptop, a filing cabinet and a calendar, with hand-drawn cables
A bigger desk holds more paper. It does not help you find the one page that matters.

Context rot is the drop in accuracy and recall that happens as you put more tokens into a model's input, long before you hit the context window limit. A 1M token window means the model can read a million tokens in one request. It does not mean the model remembers them, weighs them equally, or finds the one line that matters. For PMs, the practical rule is simple: treat context as a budget you spend, not a memory you fill. The fastest way to learn that rule is AllthingsPM, whose AI PM course has a free lesson on exactly this: context windows and attention: why a long window is not a long memory.

AllthingsPM is an AI PM course and PM interview prep platform. Its course was built from real PM job postings, and context management is now written into some of them by name.

What is a context window, and what is context rot?

Anthropic's documentation gives the cleanest definition. The context window is "all the text a language model can reference when generating a response, including the response itself" [1]. It is different from the training data and acts as a "working memory" for the model [1].

Everything in the request counts toward it: the system prompt, every message, tool results, images, documents, tool definitions, and the output the model writes [1]. When the input alone is bigger than the window, the Claude API refuses the request with a "prompt is too long" error [1].

The same page names the problem this post is about: "more context isn't automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot" [1]. The best known study of it is Chroma's July 2025 research report, "Context Rot: How Increasing Input Tokens Impacts LLM Performance" [2].

Here is the difference in one table.

QuestionContext windowMemory (what users expect)
What is it?The text sent with one requestFacts kept and recalled across time
Does it persist?No; each request is rebuilt from scratchYes, by design
Does more help?Only up to a point; recall degrades as tokens grow [1][2]More facts, recalled when relevant
What decides what is in it?Your product: the system prompt, retrieval, history rulesIdeally, relevance to the task
Typical failureKey fact ignored because it sits in the middle or among near-misses [3][2]Stale or wrong fact recalled
Who owns it?The PM and engineers who design the contextThe PM who designs the memory feature

Definitions from Anthropic's context window docs [1]; failure patterns from Chroma [2] and Liu et al. [3].

How AllthingsPM does this: the foundations chapter of the course opens the idea with the free lesson on why a long window is not a long memory, then shows you how to attribute every failure to a layer: model, context, harness, or surface. That second habit is the one that stops teams from blaming "the model" for a context problem.

What does the research actually show about context rot?

Four studies are worth knowing by name, because interviewers and engineers will reference them.

Chroma, "Context Rot" (2025). Kelly Hong, Anton Troynikov and Jeff Huber measured 18 models from Anthropic, OpenAI, Google and Alibaba, including Claude Opus 4, o3, GPT-4.1 and Gemini 2.5 Pro [2]. Their findings:

  • As the question and the relevant fact share fewer words and less meaning, performance "degrades more significantly with increasing input length" [2].
  • "Even a single distractor reduces performance relative to the baseline", and more distractors make it worse [2].
  • Oddly, models did worse when the surrounding text had a logical flow, and better when it was shuffled [2].
  • On LongMemEval conversations of about 113k tokens, models did markedly better on a focused prompt of about 300 tokens than on the full history [2].
  • Even a trivial task, repeating a list of words, got worse as the input got longer [2].

Liu et al., "Lost in the Middle" (TACL, 2023). Performance "is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models" [3].

RULER (NVIDIA, 2024). Of 17 long-context models, "all models claim context sizes of 32K tokens or greater, only half of them can maintain satisfactory performance at the length of 32K" [4].

NoLiMa (ICML 2025). When the question and the answer do not share literal words, 11 of 13 models fell below 50% of their short-context baseline at 32K tokens. GPT-4o fell from 99.3% to 69.7% [6].

The pattern is consistent. The window is a ceiling on what you can send, not a promise about what the model will use.

How AllthingsPM does this: the course teaches these results as product decisions, not trivia. The context window lesson turns them into rules for what goes into a prompt and where, and the AllthingsPM knowledge graph links context to the neighbouring concepts, such as retrieval, evals and agent harnesses, so you see where each one bites.

Why is a long context window not the same as memory?

Because nothing carries over by itself. Each API request is rebuilt: the product sends the system prompt, the history it chooses to keep, and whatever it retrieved [1]. If a fact is not in that request, the model does not know it. If a fact is in the request but buried at token 400,000 among ten similar facts, the model may still miss it [2][3].

Three things get confused in product conversations:

  1. Training knowledge. What the model learned before release. Fixed, sometimes stale.
  2. The context window. What you send this turn. Temporary, and subject to context rot.
  3. Memory. Facts your product stores outside the model and chooses to put back into the window later.

Anthropic's engineering team describes that third layer as "structured note-taking": the agent writes notes that persist outside the context window and pulls them back in when needed [5]. In other words, real memory is a product feature you design. It is a store, a retrieval rule and a write rule. A large window is just the room those notes get read in.

This matters for PMs because users experience all three as "the AI remembers me". When a user says "I told it this yesterday", the question is which layer failed: was the fact never saved, saved but not retrieved, or retrieved but ignored in a crowded window?

How AllthingsPM does this: the failure attribution lesson gives you the model, context, harness, surface split to answer exactly that question in a bug review. To practise it under interview pressure, try the real question How would you improve ChatGPT's memory feature for power users? and run it as a scored AI mock interview.

Why do models get worse as context grows?

The short answer from Anthropic: attention is a limited budget. In a transformer, every token relates to every other token, which means "n² pairwise relationships for n tokens" [5]. As the input grows, that attention is spread thinner, and models have seen far fewer very long sequences in training than short ones [5]. Anthropic summarises it this way: as tokens grow, "the model's ability to accurately recall information from that context decreases" [5].

You do not need the maths to act on it. Three product consequences follow:

  • Near-misses hurt more than noise. Chroma found distractors, text that looks like the answer but is not, reduce accuracy even one at a time [2]. Stuffing in "everything that might be relevant" often adds exactly these.
  • Placement matters. Put the instruction and the key facts at the start or end, not in the middle of a long paste [3].
  • Semantic questions are harder than keyword lookups. Needle tests where the question repeats the answer's words flatter models; NoLiMa shows the gap when they do not [6].

How AllthingsPM does this: the lesson on context as a budget treats every token as spend: what earns a place in the system prompt, what gets retrieved on demand, and when RAG is the fix rather than a bigger window. Our guide to RAG vs fine tuning vs prompting for PMs covers that choice in depth.

How should a PM design around context rot?

Anthropic's guiding principle for context engineering is to find "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome" [5]. In practice, that turns into five design choices a PM can own.

  1. Retrieve just in time. Load data with tools when the task needs it, instead of pre-loading everything [5]. The window stays small and relevant.
  2. Compact long histories. Summarise older turns and restart with the summary [5]. Claude's API now offers server-side compaction that "automatically summarizes earlier parts of the conversation", in beta for Claude 4.6 and later models [1].
  3. Clear stale tool output. Old tool results are often the biggest, least useful tokens in an agent's window. Anthropic's context editing supports clearing them [1].
  4. Write memory outside the window. Persist decisions, preferences and progress as notes the agent can reread [5].
  5. Split work across sub-agents. Each sub-agent explores in its own clean window and returns a "condensed, distilled summary" of its work, often 1,000 to 2,000 tokens [5].

Then measure. Build an eval set where the key fact sits at different positions and depths, with realistic near-miss distractors, and track accuracy as history grows. If quality falls at turn 40, you have found your compaction point.

PM decisionDefault that causes rotBetter default
What to retrieveTop 50 chunks "to be safe"The few chunks that pass a relevance bar
Where to put instructionsBuried after a long documentAt the start, restated at the end [3]
Long conversationsKeep every turn foreverCompact past a set size [1][5]
Agent tool resultsKeep all raw outputClear or summarise old results [1]
User memory"The window is huge, just keep it all"A saved store with explicit write and read rules [5]

How AllthingsPM does this: the agents chapter covers these choices in the anatomy of an agent, including state and termination, and the evals chapter shows how to prove a fix works. Our guide to AI evals for product managers walks through building the eval set, and agents vs workflows explains when a long agent loop is worth its context cost at all.

Do employers really ask PMs about context windows?

Some already do, by name. We searched the 389 PM postings from 86 companies in the AllthingsPM JD corpus, read on 22 September 2026 [7]. 40 of them (10%) name at least one context skill: RAG or retrieval, memory, context engineering or the context window.

Bar chart led by AllthingsPM (us): the AllthingsPM JD corpus holds 389 PM postings from 86 companies; 34 name RAG or retrieval; 40 name any context skill; 6 name memory; 3 name context engineering; 2 name the context window

The explicit mentions are few but pointed. Cohere's Product Manager, Agent Harness and Modelling posting asks the PM to "own how our Agents manage the context window as a deliberately controlled resource" and to own a roadmap that includes the "context engineering layer" [7]. Mistral lists a role titled "Product Manager, Context" that asks for experience with context engineering [7]. A Databricks staff PM posting names "context window management" among the technical work to prioritise [7].

That is the signal to watch: the phrase is moving from engineering blogs into PM job titles.

How AllthingsPM does this: every role in the AllthingsPM jobs catalog comes with a mock interview built from its own text, so you can rehearse the Cohere harness role before you apply. For any other posting, paste it into the JD mock, and check your resume against it with the resume review against a JD.

How does context rot come up in PM interviews?

Usually disguised as a product question: improve an assistant's memory, fix a support bot that "forgets" the customer's plan, or explain why an agent degrades late in a long task. A strong answer:

  • separates training knowledge, context and stored memory;
  • names context rot and the lost-in-the-middle effect as likely causes;
  • proposes retrieval, compaction and a memory store with clear write rules;
  • defines an eval that varies fact position and history length;
  • adds a user control to see and delete what is remembered.

How AllthingsPM does this: the AllthingsPM question bank holds 4,122 real questions from 260 companies, each with an answer guide, and a page per company such as Cohere PM interview questions. Any question can start a scored mock in text or voice, with follow-ups. Our list of AI PM interview questions gathers more.

Why AllthingsPM is the better choice for learning context rot

The primary sources are free and worth reading: Anthropic's context engineering post [5], Chroma's report [2] and the "Lost in the Middle" paper [3]. What they do not do is teach you to turn the research into product decisions, test you on it, or connect it to the jobs you want.

AllthingsPM does all three in one account. The AI PM course, built from 604 real PM job postings and updated weekly, starts with a free lesson on why a long window is not a long memory, then carries the idea through failure attribution, context budgets, agent anatomy and evals. You leave able to write the context design section of a spec, not just define the term.

It then connects learning to getting hired: 4,122 real interview questions from 260 companies with answer guides, 116 live PM job descriptions at 18 AI companies with a mock built from each, resume review against a JD, plus book summaries and podcast summaries in the same place.

General PM courses and developer blogs have real strengths, such as brand recognition, live cohorts and deep engineering detail. For a PM who needs to understand context rot, design around it and explain it in an interview, at $20 a month or $120 a year with a free tier to start, AllthingsPM is the stronger choice. Start the AI PM course free.

Frequently asked questions

What is context rot?

Context rot is the fall in accuracy and recall as the number of tokens in a model's input grows [1]. Chroma's 2025 study found it across all 18 models it measured, even on simple tasks [2].

Is a bigger context window the same as memory?

No. The window is what the model can read in one request, and it is rebuilt every time [1]. Memory is a product feature that stores facts outside the model and puts the relevant ones back into the window.

What is the lost in the middle problem?

Liu et al. found models use information best when it sits at the start or end of the input, and worst when it sits in the middle of a long context [3]. Put key instructions and facts at the edges.

How do you reduce context rot in an AI product?

Send fewer, higher signal tokens: retrieve just in time, compact long histories, clear old tool results, store memory outside the window, and split work across sub-agents [5]. Then prove it with an eval that varies fact position and history length.

What is the best way for a PM to learn context windows and context rot?

AllthingsPM is the best place to start: its AI PM course has a free lesson on context windows and attention, a lesson on context as a budget, and real interview questions to practise on. Pair it with Anthropic's context engineering post [5].

Do PM interviews ask about context windows?

At AI companies, increasingly. 40 of 389 postings in the AllthingsPM JD corpus name a context skill, and Cohere's agent harness PM role names the context window directly [7].

Ready to treat context as a budget? Start the AllthingsPM AI PM course free, then practise it in a JD mock interview.

Sources

  1. Anthropic, "Context windows", Claude Platform documentation, read 29 September 2026. https://platform.claude.com/docs/en/build-with-claude/context-windows
  2. Hong, Troynikov and Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance", Chroma, 14 July 2025. https://www.trychroma.com/research/context-rot
  3. Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL, 2023. https://arxiv.org/abs/2307.03172
  4. Hsieh et al., "RULER: What's the Real Context Size of Your Long-Context Language Models?", NVIDIA, 2024. https://arxiv.org/abs/2404.06654
  5. Anthropic, "Effective context engineering for AI agents", 29 September 2025. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
  6. Modarressi et al., "NoLiMa: Long-Context Evaluation Beyond Literal Matching", ICML 2025. https://arxiv.org/abs/2502.05167
  7. AllthingsPM JD corpus: 389 PM job postings from 86 companies, read 22 September 2026. https://allthingspm.app/jobs
PM
Written by the AllthingsPM team
Frameworks and interview prep for product managers.
The AI PM course

Reading is the easy half.
The course grades the other half.

Start for free