The context window and its practical limits
A shared token budget, not durable memory: what the window is and where it runs out.

Every call to a language model carries a fixed budget of text, and the model can use only what fits inside it. Anthropic’s documentation defines the “context window” as “all the text a language model can reference when generating a response, including the response itself”, and separates it from “the large corpus of data the language model was trained on”. The window is therefore not what the model knows, but what it can read in this one call. Google’s Gemini documentation states the same fact as a budget: “Each Gemini model has a maximum number of tokens it can handle. The context window defines the combined limit of input and output tokens.” The number is a capacity published per model, not a promise about quality or durable memory.
The window as a shared budget
One call, one budget: the text you send and the text the model writes back are charged to the same limit. Drawing it as a container with named occupants makes the sharing visible:
one call = one token budget (its size is fixed per model)
+--------------------------------------------------------------+
| context window |
| |
| system instruction ...................... application-set |
| conversation history .................... grows each turn |
| retrieved material ...................... retrieval is L3 |
| tool definitions ........................ |
| the current user message ................ |
| |
| +----------------------------------------------------------+|
| | output reservation: the response this call will write ||
| +----------------------------------------------------------+|
+--------------------------------------------------------------+
^ every item above is counted before the model answers
v the response it writes becomes input to the next call
The output is not extra space appended at the end; it is part of the same account. Before answering, the model must fit the answer it is about to produce into the space it is reading from — which is why input and output are one combined limit, not two.
The request also carries its own upper bound on the tokens generated for the reply — set by the application, not by the window — and what that bound really limits is the subject of the sibling sheet on output limits.
Terminology
- Context window — the whole text one call can reference, the response included (Anthropic).
- Token — the unit the budget is counted in; what a token is belongs to the sibling sheet on tokens.
- Input and output phase — the input carries the accumulated history plus the new message; the output is the generated response (Anthropic).
What counts, and why input and output share it
Anthropic lists the occupants: “Everything in the request counts toward the context window: the system
prompt, every message in messages (including tool results, images, and documents), and your tool
definitions.” The output the model generates for the turn, thinking included, counts too.
Google says the same of the system instruction — “System instructions are counted as part of the input
tokens” — and calls the window the “combined limit of input and output tokens”. The two are one account
because the output does not vanish at the end of the turn: “each user message and assistant response
accumulates within the context window, and previous turns are preserved completely.” What the model writes
now is part of what it reads next. That is the default inside one accumulating conversation, not the only
documented mechanism: the same page describes server-side compaction — “summarizes earlier parts of the
conversation on the server, so the conversation can continue past the context window limit” — and some
interfaces evict instead, on a rolling “first in, first out” basis.
When the window overflows
Sooner or later the request does not fit; the OpenAI API specification documents two opposite behaviours. With truncation set to auto, “the model will truncate the response to fit the context window
by dropping items from the beginning of the conversation”. With the default, disabled, “the request will
fail with a 400 error”. (The specification marks that field deprecated as read here, so the semantics may
move; the two behaviours are what it documents today.) One behaviour silently discards the oldest material; the other refuses. When the
budget is exhausted something is lost either way — the beginning of the conversation, or the request.
Anthropic notes that some chat interfaces manage the window on a rolling “first in, first out” basis
as well.
What a larger window changes, and what it does not
| Question | What the sources say |
|---|---|
| Can the model read more at once? | Yes — “A larger context window allows the model to handle more complex and lengthy prompts” (Anthropic). |
| Does more room mean the material is used better? | No — “more context isn’t automatically better” (Anthropic). |
| Is every position in the window used equally? | No — performance is highest with the relevant text at the beginning or the end and degrades in the middle (Liu et al.). |
| Does a larger window cost the same? | No — the cost of a call “is determined in part by the number of input and output tokens” (Google). |
| Does a larger window change what the model knows? | No — the window is not the corpus the model was trained on (Anthropic). |
The first row is real capacity; the next two are why it is not a quality guarantee. Anthropic is explicit: “more context isn’t automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot. This makes curating what’s in context just as important as how much space is available.” The empirical picture comes from Lost in the Middle, which reports that performance is often highest when the relevant information sits at the beginning or the end of the input and degrades when it must be found in the middle, even for explicitly long-context models — a pattern its authors draw as a “U-shaped performance curve”. A larger window buys room to hold more, not the ability to use the same item equally well wherever it sits.
Limits, and the common conceptual error
The common error is to read the window size as a statement about the model. The window is not the training corpus, and it is not durable memory: it is the text carried in this one call, gone when the call ends unless something outside the model keeps it. The size is a property of a specific model, not of language models in general.
Cost follows from the same accounting: the budget is not free, because “the cost of a call to the Gemini API is determined in part by the number of input and output tokens”. Time is my own framing, not a sourced measurement: processing more tokens takes longer; no latency figure is claimed. What is not framing is the curation decision: because the window is finite, and more content does not mean better use of it, choosing what goes in context is part of designing the system.
What to remember: the window is a shared, per-model token budget for one call; the response is spent from it and re-enters the next call as input; overflow loses material in one of two documented ways, and some products summarise or evict instead; and a larger window is more room, not better use of what is in it.
Level and prerequisites
L1 — concepts, terminology and limits only: no chunking, no retrieval, no parameter settings, no cost model. Prerequisites: none; the sibling sheet on tokens explains the unit named here.
Where to go next
- Automation & AI — the area this sheet belongs to.
- Retrieval — selecting external material for the window — is L3 material, named here but not developed.
References
- Anthropic, Context windows (engineering documentation) — the definition of the context window and its separation from the training corpus, the list of what counts toward it, progressive accumulation across turns, the rolling FIFO behaviour of chat interfaces, and context rot.
- Google, Gemini API documentation — Tokens and the context window — the combined input-and-output limit, system instructions counted as input tokens, token counting, and the billing note that cost depends on input and output tokens.
- OpenAI, OpenAI API specification (
openapi.yaml) — the two documented overflow behaviours (truncate by dropping items from the beginning of the conversation; fail with a 400 error) and the upper bound on generated tokens. - Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang, Lost in the Middle: How Language Models Use Long Contexts (arXiv:2307.03172v3, 20 November 2023) — the position effect and the “U-shaped performance curve”.