Context Window
The maximum amount of text an LLM can process in a single interaction — everything you send it plus everything it has generated so far must fit within this limit.
What it means
The context window is the maximum amount of text — input plus output combined — that a large language model can process at once. Text is measured in tokens (roughly 0.75 words per token on average). Everything the model can “see” during a conversation — your prompt, any documents you uploaded, and the full conversation history — must fit within the context window.
Current context window sizes for common models (September 2026):
- Claude (Anthropic): ~200K tokens (~150K words)
- GPT-4o (OpenAI): ~128K tokens (~96K words)
- Gemini 1.5 Pro (Google): up to 1M tokens (long-context version)
Why it matters for researchers
Practical effect: documents get cut off. If you upload a 300-page PDF and ask a model with a 128K context window to analyze it, the model either truncates the document or samples from it. It cannot read the whole thing at once. This is why splitting long documents or using multiple sessions is sometimes necessary.
Common research scenarios where context window matters:
- Full manuscript review: A 15,000-word draft with reviewer comments may be 25–30K tokens — fits in any current major model. A whole book does not.
- Multi-document synthesis: Loading five 40-page papers simultaneously requires ~100–150K tokens — needs a long-context model
- Long conversation history: Extended back-and-forth conversations eventually push earlier context out of the window; the model “forgets” what was said at the start
“Lost in the middle” effect: Research has shown that LLMs tend to recall information from the beginning and end of a context window better than information in the middle. For critical content in a long document, placing it near the start or explicitly asking the model to summarize it improves recall.
A longer context window is not always better. Models with very long contexts are sometimes less attentive to details buried in the middle than shorter-context models are over shorter inputs. Verify retrieval of key claims in long-context tasks rather than assuming the model read everything.