Attention Mechanism
The component in transformer models that lets the model weigh which parts of the input are most relevant to each word it generates — the core innovation behind modern LLMs.
What it means
The attention mechanism is the mathematical operation inside transformer-based AI models (including GPT, Claude, and Gemini) that determines which parts of the input text are most relevant when generating each word of the output. Instead of reading text sequentially like earlier models, attention allows the model to look across the entire input at once and assign relevance weights to every other word.
When you ask an LLM a question, the model uses attention to decide which words in your prompt matter most for each word it generates in response — effectively “paying attention” to different parts of the input simultaneously.
Self-attention (or “multi-head attention”) is the specific variant used in transformers: the model attends to its own input, learning which parts of a sentence relate to which other parts. This is how the model understands that “it” in “the protein folded because it was unstable” refers to “the protein” rather than “folded.”
Why it matters for researchers
- Context window size is a direct consequence of how attention is computed. Longer contexts require more computation, which is why LLMs have context limits (typically 128K–200K tokens).
- Attention maps can be inspected to understand what a model “focused on” when making a prediction — this is used in interpretability research and is sometimes reported in scientific papers using transformer-based models.
- “Flash Attention” — a term you may see in model release notes — refers to an efficient implementation of the attention computation, not a different kind of model; it just makes the same operation faster and more memory-efficient.