Definition
A context window is the maximum amount of text a large language model can take in and work with at one time. It’s measured in tokens, the small units of text a model reads, and it covers everything involved in a single exchange: the instructions, the conversation so far, any documents supplied, and the response the model writes.
It’s often described as the model’s working memory. Whatever is inside the window, the model can use. Whatever falls outside it doesn’t exist for the model unless it’s supplied again.
A context window isn’t long-term memory. A model doesn’t remember last week’s conversation because of its context window. Products that appear to remember across sessions do so by storing information separately and placing it back into the window when needed.
See also: LLM Tokens, Large Language Models (LLM), Retrieval-Augmented Generation, Prompt Engineering
Why it matters for marketing
The context window sets a practical ceiling on what a marketer can hand to an AI tool in one go. A team that wants an AI assistant to write in the brand’s voice has to fit the brand guidelines, the brief, the product details and a few examples into the window, with room left for the draft. If the material doesn’t fit, something gets left out or cut off.
It also explains a behavior many users notice without knowing the cause. In a long chat session, the assistant starts ignoring instructions given at the beginning, or contradicts something it said earlier. Usually the conversation has grown past what the window holds, and the tool has dropped or compressed the oldest parts.
Cost is tied to it as well. AI services generally charge by the token, so sending a 200-page research report with every question costs more than sending the three relevant pages.
For teams deploying chatbots or AI agents, the window determines how much product catalog, policy text and conversation history the system can consult when answering a customer.
How it works
Everything a model processes in one request draws on the same token budget.
| What uses the window | Example |
|---|---|
| System instructions | The hidden prompt that tells an assistant how to behave |
| Conversation history | Every earlier message and reply in the session |
| Supplied documents | A pasted brief, an uploaded PDF, a spreadsheet |
| Retrieved content | Passages pulled in from a knowledge base or web search |
| Tool definitions and results | Descriptions of connected tools and the data they return |
| The response | The output the model generates |
When the total would exceed the limit, something has to give. Depending on the product, the oldest messages are dropped, earlier parts of the conversation are replaced by a summary, or the request is rejected with an error.
How large are context windows?
They’ve grown quickly. GPT-3, released in 2020, supported 4,096 tokens, which is roughly 3,000 words. GPT-4 Turbo reached 128,000 tokens in late 2023. Google’s Gemini 1.5 Pro made a 1 million token window broadly available, and several current models from Anthropic, Google and OpenAI now offer windows of 1 million tokens or more.
Sizes change with nearly every model release. Check the provider’s documentation for current figures.
Bigger isn’t automatically better
A model can accept a very long input and still use it unevenly. Two research findings are widely cited.
Lost in the middle. A 2023 study by Nelson Liu and colleagues at Stanford found that models were best at using information placed at the beginning or end of a long input. Performance dropped noticeably when the relevant information sat in the middle. Later testing suggests newer models have improved on this. A 2025 paper by a Google researcher found that Gemini 2.5 Flash retrieved simple facts accurately from any position, even near its input limit.
Context rot. A July 2025 report from Chroma tested 18 models and found that performance declined as input length grew, particularly when the input included distracting or loosely related material. In one test, models did better with a focused history of about 300 tokens than with a full chat log of about 113,000.
The practical conclusion from both is the same. What goes into the window matters more than how much the window can hold.
How to calculate token needs
A rough estimate helps decide whether material will fit and what it will cost.
Estimated tokens ≈ word count ÷ 0.75
For English text, one token is about three-quarters of a word, so 1,000 words is roughly 1,300 tokens. The ratio varies by model and by language.
Then add up everything going into the request.
Total tokens = instructions + conversation history + documents + expected response
An illustrative example for a content task:
| Component | Words | Estimated tokens |
|---|---|---|
| Instructions and brief | 600 | 800 |
| Brand guidelines | 6,000 | 8,000 |
| Research report | 30,000 | 40,000 |
| Three example articles | 4,500 | 6,000 |
| Expected response | 1,500 | 2,000 |
| Total | 42,600 | 56,800 |
A request of about 57,000 tokens fits comfortably in a 128,000-token window and wouldn’t fit in a 32,000-token one. If it didn’t fit, the research report is the obvious thing to trim, since it uses about 70% of the budget.
How to utilize the context window
Supply reference material. Put brand guidelines, product sheets and style examples into the window so the model works from them instead of guessing.
Analyze long documents. Summarize or query reports, transcripts and contracts in a single pass when they fit.
Keep sessions focused. Start a new conversation for a new task so old, irrelevant history doesn’t crowd the window.
Restate what matters. In a long session, repeat key instructions, since early ones may have been dropped.
Use retrieval for large knowledge bases. When the material is far larger than the window, retrieval-augmented generation pulls in only the relevant passages for each question.
Place important content deliberately. Put the most important instructions and facts at the start or end of a long prompt.
Budget for agents. An AI agent running many steps accumulates context with each one. Plan for summarization points in long workflows.
Comparison with related concepts
| Term | What it is | Relationship to the context window |
|---|---|---|
| Context window | The total text a model can process in one request | The subject of this entry |
| LLM Tokens | The units of text a model reads and writes | The unit the window is measured in |
| Training data | The text a model learned from before release | Fixed knowledge. The window holds what’s supplied at the time of use |
| Memory (product feature) | Stored information a product re-supplies across sessions | Sits outside the window and is loaded into it when needed |
| Retrieval-Augmented Generation | Fetching relevant content and adding it to the prompt | A way to work with more material than the window holds |
| Maximum output tokens | The cap on the length of a single response | A separate, smaller limit inside the window |
| Knowledge base | A store of documents an AI system can draw on | The full library. The window holds the subset in use |
| Context engineering | Deciding what information goes into the window | The practice of managing it |
| Model Context Protocol (MCP) | A standard for connecting AI systems to tools and data | A way of bringing external data into the window |
Context window and training data are the two most often confused. Training data is what the model already knows. The context window is what it’s looking at right now.
Best practices
Include what’s relevant and leave out the rest. Extra material costs more and can lower the quality of the answer.
Trim documents before uploading. Send the relevant section of a report, not the whole file, when only part of it bears on the question.
Watch for drift in long sessions. If an assistant starts forgetting instructions or contradicting itself, start a new session with a short summary of what’s been decided.
Keep standing instructions in a reusable place. Many tools offer project instructions or custom settings that are loaded into every session.
Leave room for the response. A window filled entirely with input has no space left for output.
Check the limits of the specific tool. A consumer chat product may apply a smaller limit than the underlying model supports.
Don’t treat the window as storage. Anything that needs to persist should be saved somewhere outside the conversation.
Future trends
Larger windows. Windows of 1 million tokens are now offered by several providers, and larger ones have been announced.
Attention to effective context. Research such as Chroma’s has shifted focus from how much a model can accept to how well it uses what it’s given.
Context engineering as a discipline. The practice of selecting and arranging what goes into the window is becoming a recognized skill, especially for building AI agents.
Automatic context management. AI products increasingly summarize or prune long conversations on their own, so the window is managed without the user’s involvement.
Persistent memory features. More assistants now store information between sessions and reload it, which reduces the need to restate background each time.
Pricing changes. Some providers have charged higher rates for very long inputs, and techniques such as caching repeated content are being used to bring costs down.
Frequently asked questions
What is a context window? It’s the maximum amount of text, measured in tokens, that an AI language model can process in one request, including both the input and its response.
How big is a context window? It varies by model. Early models handled a few thousand tokens. Several current models handle 1 million or more.
What happens when the context window is full? The tool drops or summarizes older content, or returns an error. In a chat, that’s when an assistant appears to forget earlier parts of the conversation.
Is a context window the same as memory? No. It’s short-term working memory for a single session. Memory across sessions is a separate product feature that stores information and reloads it.
How many words fit in a context window? For English, roughly three-quarters of a word per token. A 128,000-token window holds about 96,000 words.
Does a bigger context window give better answers? Not necessarily. Research shows performance can decline as inputs get longer, especially when they include irrelevant material.
What is “lost in the middle”? A 2023 research finding that models used information at the start and end of a long input better than information in the middle. Newer models appear to have improved on this.
What is context rot? A term for the decline in model performance as the amount of input grows. It was described in a 2025 report by Chroma.
How does the context window affect cost? Most AI services charge per token, so larger inputs cost more for each request.
How is the context window different from training data? Training data is what a model learned before it was released. The context window holds the information supplied during a specific conversation.
Related terms
- LLM Tokens
- Large Language Models (LLM)
- Retrieval-Augmented Generation
- Prompt Engineering
- Model Context Protocol (MCP)
- Generative Pre-trained Transformer (GPT)
- AI Agent
- Generative AI
- Multimodal AI
- Conversational AI
Sources
- Liu, Nelson F., et al. “Lost in the Middle: How Language Models Use Long Contexts.” Transactions of the Association for Computational Linguistics, vol. 12 (2024). Preprint: https://arxiv.org/abs/2307.03172
- McKinnon, Max. “Retrieval Quality at Context Limit.” arXiv, 2025. https://arxiv.org/pdf/2511.05850
- PromptHub. “Context Rot: Why LLMs Fail as Context Windows Grow.” https://prompthub.substack.com/p/context-rot-why-llms-fail-as-context
- 13 Labs. “Context Windows and Token Limits: Why a Bigger Window Does Not Fix It.” https://www.13labs.au/guides/context-window-token-limits
- memX. “Context Window” (glossary). https://memx.app/glossary/context-window/
- Metavert. “Context Windows.” https://metavert.io/context-windows
- Maven AGI. “Context Window” (glossary). https://www.mavenagi.com/glossary/context-window
