Attention Mechanism
Attention Mechanism
Definition: A technique that lets a model weigh the relevance of different parts of the input when producing each part of the output, instead of treating all input equally. Every output position computes a distribution of “attention weights” over the input positions, then builds its representation as a weighted sum guided by those weights. It replaced the fixed-length bottleneck of earlier sequence models with a mechanism that can look at, and selectively pull from, the entire input at every step. Attention is the computational core of the Transformer Architecture, and by extension of nearly every modern Large Language Model (LLM).
How It Works
The Bottleneck Problem It Solves
Before attention, sequence-to-sequence models (encoder-decoder RNNs) compressed an entire input sentence into a single fixed-length vector — the final hidden state of the encoder — and forced the decoder to reconstruct the whole output from that one vector. This worked for short sentences and degraded badly on long ones: information from early tokens got diluted or overwritten by the time the encoder reached the end of the sequence, a symptom of the same vanishing-gradient dynamics that make plain Recurrent Neural Network (RNN) training hard over long spans. Bahdanau et al. (2014) fixed this by letting the decoder, at each output step, look back at every encoder hidden state and compute a weighted combination of them, with weights learned to reflect relevance to the current decoding step. That was the first attention mechanism. Vaswani et al. (2017) then made the radical move of removing recurrence entirely and building a model — the Transformer — out of nothing but attention and feed-forward layers, which is the architecture underlying essentially every large-scale language model today.
Queries, Keys, and Values
Attention is formalized as an information-retrieval operation using three learned projections of each input vector:
- Query (Q) — represents “what this position is looking for.”
- Key (K) — represents “what this position offers,” a label that queries get matched against.
- Value (V) — represents “what this position actually contributes” once it’s selected.
Given an input embedding , three learned weight matrices , , project it into query, key, and value vectors:
The retrieval analogy is exact: think of a lookup table where every entry has a key and a value. A query doesn’t have to match a key exactly — instead of a hard lookup, attention computes a soft match between the query and every key in the sequence, then blends the corresponding values in proportion to how well each key matched. Nothing here is hand-designed; , , and are ordinary learned parameters trained end-to-end via Backpropagation, so the model itself decides what “relevant” means for the task.
Scaled Dot-Product Attention
The full operation, as defined in the original Transformer paper, is:
Reading it left to right:
- — every query vector is dot-producted against every key vector, producing an matrix of raw similarity scores for a sequence of length . A high dot product means the query and key point in a similar direction in vector space — this is the same intuition behind Cosine Similarity, just without the normalization step.
- — the scores are divided by the square root of the key dimension. As grows, the variance of the raw dot products grows with it, pushing values into the extreme tails of the softmax where gradients are nearly flat. Dividing by keeps the variance roughly constant regardless of dimension, which keeps training stable.
- softmax — applied row-wise, converting each query’s row of scores into a probability distribution: non-negative weights that sum to 1 across all keys.
- — each query’s output is the weighted sum of every value vector, weighted by that query’s attention distribution. Relevant tokens contribute more; irrelevant tokens contribute close to nothing.
The overall data flow looks like this:
Worked Example: Computing Attention Weights by Hand
Take the 3-token sequence “The cat sat” and, for simplicity, use toy 2-dimensional query/key/value vectors (real models use hundreds or thousands of dimensions, but the arithmetic is identical). Suppose the learned projections have already produced:
| Token | Query (Q) | Key (K) | Value (V) |
|---|---|---|---|
| “The” | [1, 0] | [1, 0] | [1, 2] |
| “cat” | [0, 1] | [0, 1] | [3, 1] |
| “sat” | [1, 1] | [1, 1] | [0, 1] |
For the query “The” (), the raw dot products against all three keys are , , and . Scaling by gives . Exponentiating and normalizing (softmax) yields weights of roughly — “The” attends almost equally to itself and to “sat,” and less to “cat.” Repeating the same steps for every query row produces the full attention weight matrix:
Key:"The" Key:"cat" Key:"sat"
Query:"The" | 0.40 | 0.20 | 0.40 |
Query:"cat" | 0.20 | 0.40 | 0.40 |
Query:"sat" | 0.25 | 0.25 | 0.50 |
Each row sums to 1.0 — it’s a probability distribution over “where to look.” Multiplying each row by the value matrix produces the final output for every token in the sequence:
| Query token | Weights (The, cat, sat) | Output vector |
|---|---|---|
| “The” | 0.40, 0.20, 0.40 | |
| “cat” | 0.20, 0.40, 0.40 | |
| “sat” | 0.25, 0.25, 0.50 |
None of these output vectors are copies of the original value vectors — each is a fresh blend, unique to how that token’s query matched every key in the sequence. Stack this operation across every position and every layer, and it’s the mechanism by which a transformer builds context-aware representations out of context-free token embeddings.
Attention Inside a Transformer Block
A single attention computation is never used in isolation in a real model — it’s one sub-layer inside a repeated block that also includes a position-wise feed-forward network, residual (“skip”) connections, and normalization. One encoder block looks like this:
The residual connection — adding the block’s input back to its output before normalizing — is not decorative. Stacking dozens of attention layers (modern LLMs commonly have 32 to 100+) would otherwise face the same vanishing-gradient problem attention was invented to escape in RNNs; the residual path gives gradients a direct route back to earlier layers regardless of how deep the stack gets. The feed-forward network that follows attention in every block applies the same two-layer nonlinear transformation independently to each position — attention is the only sub-layer that lets positions exchange information with each other; the feed-forward network processes what attention already gathered.
Why It Matters
- Solved the long-range dependency problem that made RNN-based sequence models unreliable past a few dozen tokens — a query can attend directly to a token 10,000 positions away with no signal degradation, unlike recurrence, which has to propagate information step by step.
- Enabled full parallelization during training. Because every position’s attention computation is independent of the others (no sequential hidden-state dependency), the entire sequence can be processed in one matrix multiplication on a GPU or TPU, which is what made training models with billions of parameters computationally feasible.
- Became the foundational primitive of modern AI. GPT, Claude, Gemini, Llama, and effectively every frontier language model is a stack of attention layers — understanding attention is a prerequisite for understanding how any of them work.
- Generalized far beyond text. Vision Transformers (ViT) apply the same mechanism to image patches for Computer Vision tasks; AlphaFold uses attention over amino acid residue pairs for protein structure prediction; Whisper uses it over audio spectrogram frames for speech recognition.
- Opened a (partial) window into model internals. Attention weights can be visualized as heatmaps over the input, giving researchers one of the few directly inspectable signals inside an otherwise opaque network — a recurring tool in Explainable AI (XAI) research, even though the weights alone don’t fully explain model behavior (see Pitfalls).
- Drove a wave of systems and hardware research. Because naive attention costs in sequence length, scaling context windows required new algorithms — FlashAttention, sparse and windowed attention variants, and KV-caching schemes that now dominate inference-serving engineering.
- Made transfer learning the default paradigm. Pretraining a single large attention-based model and then cheaply adapting it — via Fine-Tuning or plain prompting — replaced training bespoke architectures per task, reshaping how the entire industry builds NLP products.
- Underlies retrieval and search infrastructure. Vector databases, semantic search, and Retrieval-Augmented Generation (RAG) pipelines all lean on the same query-key similarity-and-weighting idea attention formalized, just applied at the level of documents instead of tokens.
- Enabled in-context learning. Because attention lets a model condition its output on arbitrary spans of provided text, a pretrained model can perform new tasks from a handful of examples placed in the prompt, without any weight updates — the basis of modern Prompt Engineering.
- Surfaced unexpected internal phenomena. Researchers found that early tokens — often just the very first token — consistently absorb a disproportionate share of attention weight across heads and layers, an “attention sink” effect now deliberately exploited by streaming-inference systems to keep long-running conversations stable once older tokens get evicted from a fixed-size cache.
Multi-Head Attention: Why One Attention Pattern Isn’t Enough
A single attention operation produces exactly one weighted average per query — one notion of “relevance” per token, per layer. But language has many simultaneous relationships worth tracking at once: subject-verb agreement, coreference (“it” resolving to “animal”), adjacency, topical similarity, syntactic dependency structure. Cramming all of that into one softmax distribution forces destructive compromises.
Multi-head attention runs several attention operations in parallel, each with its own learned , , projections into a lower-dimensional subspace ( for heads), then concatenates the results and projects back to the model dimension with a final learned matrix :
Each head is free to specialize during training. Interpretability research on trained transformers has repeatedly found heads with clean, interpretable roles. No one designs these roles by hand — they emerge from gradient descent because splitting the representational work across independent subspaces gives the optimizer room to specialize:
- Positional heads — attend almost entirely to the immediately preceding or following token, effectively encoding local adjacency.
- Syntactic heads — track grammatical relationships such as subject-verb agreement or determiner-noun binding.
- Coreference heads — link pronouns back to the entities they refer to, sometimes across long spans of text.
- Induction heads — copy a previously seen pattern of the form
[A][B] ... [A] -> [B], a mechanism directly tied to in-context learning in autoregressive models. - Rare / near-null heads — contribute little to the final output and are frequently the ones removed by pruning with negligible quality loss.
Multi-head attention doesn’t change the asymptotic cost of attention — splitting across heads and running smaller attention operations costs about the same as one large one — but it does add real overhead in practice (more matrix multiplications, more memory traffic for intermediate tensors). The number of heads is a hyperparameter with diminishing returns: too few heads and the model can’t represent multiple relationship types simultaneously; too many heads and each one gets too little capacity ( shrinks) to represent anything useful. Head-pruning studies have found that many trained heads can be removed post-hoc with minimal quality loss, suggesting real transformer models are often over-parameterized in head count relative to what any single task needs.
Positional Encoding: Giving Attention a Sense of Order
Attention as defined above is permutation-invariant — nothing in the -softmax-weighted-sum computation depends on token order. Shuffle the input tokens and, absent extra information, the set of attention outputs is identical, just reassigned to different positions. Transformers restore order by injecting a position-dependent signal into each token’s embedding before the first attention layer.
- Sinusoidal encoding (original Transformer) — a fixed, non-learned pattern where each dimension of the positional vector is a sine or cosine wave of a different frequency. Because it’s deterministic rather than learned, it can in principle be evaluated at positions never seen during training.
- Learned absolute positional embeddings (BERT, GPT-2) — a trainable embedding vector per position index, looked up the same way as a token embedding. Simpler to implement, but the model has no way to represent a position past whatever maximum length it was trained with.
- Relative position encodings — encode “how far apart are this query and this key” rather than “where am I in the sequence,” injected directly into the attention score computation. This generalizes better to unseen lengths, because a 50-token offset looks the same near the start or the end of a much longer document.
- RoPE (Rotary Position Embedding) — rotates each query and key vector by an angle proportional to its position before the dot product, so the dot product between two rotated vectors naturally encodes their relative offset. RoPE is the positional scheme behind most current open-weight LLMs (Llama, Mistral, Qwen).
- ALiBi (Attention with Linear Biases) — leaves the vectors untouched and instead subtracts a penalty from the raw attention score that grows linearly with query-key distance, applied directly inside the softmax input. Models trained with ALiBi have shown strong measured extrapolation to sequence lengths well beyond their training length.
The original sinusoidal formulation defines even and odd dimensions of the positional vector separately:
| Method | Encodes | Learned? | Extrapolates beyond training length? |
|---|---|---|---|
| Sinusoidal | Absolute position | No | Moderate |
| Learned absolute embedding | Absolute position | Yes | Poor |
| RoPE | Relative position, via vector rotation | No | Good |
| ALiBi | Relative position, via score bias | No | Best measured |
| Relative bias table (T5-style) | Relative position, via learned bucketed bias | Yes | Moderate |
Without one of these schemes, attention alone would treat “the dog bit the man” and “the man bit the dog” as sets of the same three attended-to concepts, with no representation of which noun preceded the verb and which followed — the exact detail that determines the sentence’s meaning.
Causal Masking and Encoder vs. Decoder Attention
Not all attention layers see the same slice of the sequence, and the distinction matters for both correctness and how models are trained.
Bidirectional (encoder) self-attention lets every position attend freely to every other position, past and future — this is how BERT-style encoder models build representations, and it’s appropriate when the whole input is available up front (e.g., classifying a complete sentence, or encoding a source sentence for translation).
Causal (masked) self-attention restricts position to attending only to positions . This is enforced mechanically by setting every score above the diagonal of the matrix to before the softmax, so those positions get an attention weight of exactly zero. This is mandatory for autoregressive generation: during training, a decoder predicts token from tokens , and if it could peek at token through attention, it would trivially learn to copy the answer instead of learning to predict it — a train-time shortcut that collapses the moment the model has to generate token-by-token at inference with no future tokens to peek at. GPT-style decoder-only LLMs use causal self-attention exclusively.
Cross-attention is the third configuration: queries come from one sequence (typically a decoder) while keys and values come from a different sequence (typically an encoder’s output). This is how the original encoder-decoder Transformer connects the two halves for translation — each decoder position queries over the entire source sentence’s encoded representations — and it’s the same mechanism multimodal models use to let a text decoder attend over image or audio embeddings.
A practical consequence of causal masking: because position ‘s output never depends on future tokens, its key and value vectors never change once computed. Inference engines exploit this with a KV-cache — computed keys and values are stored and reused for every subsequent generation step rather than recomputed from scratch, which is the single biggest lever for making autoregressive text generation fast in production.
Efficient Attention at Scale: Taming the O(n²) Cost
Standard scaled dot-product attention materializes the full score matrix, so both its compute and its memory footprint grow quadratically with sequence length. At a few hundred tokens this is irrelevant; at the context windows modern LLMs now advertise — well into the hundreds of thousands of tokens — naive attention becomes the dominant cost of both training and serving. The growth is easy to underestimate until it’s written out:
| Sequence length () | Score matrix entries () |
|---|---|
| 100 | 10,000 |
| 1,000 | 1,000,000 |
| 10,000 | 100,000,000 |
| 100,000 | 10,000,000,000 |
A 100x increase in sequence length is a 10,000x increase in score-matrix size — which is why context-window growth in production LLMs has tracked almost exactly with progress in the techniques below. Several complementary techniques attack different parts of this cost.
- FlashAttention — an exact, non-approximate reimplementation that never writes the full score matrix to slow GPU memory. It processes attention in small tiles, keeping intermediate results in fast on-chip memory, and recomputes what it needs during the backward pass instead of storing it. The output is mathematically identical to standard attention — the entire gain comes from being memory-bandwidth-aware, not from approximating anything.
- Sparse / windowed attention (Longformer, BigBird) — restricts each query to attend only to a local window of nearby tokens plus a handful of designated global tokens, cutting cost from to roughly . Any relationship not covered by the local window or a global token is simply invisible to that layer.
- Sliding window attention (Mistral) — a simpler variant: every token attends only to the previous tokens. Because this repeats across many stacked layers, the effective receptive field still grows with depth (layer ‘s window reaches roughly tokens back), even though no single layer sees the full sequence.
- Linear attention approximations — reformulate the softmax similarity as a kernel feature map so the output can be computed without ever forming explicitly, reducing both compute and memory to , usually at some cost to model quality relative to full softmax attention.
- Multi-Query Attention (MQA) — keeps multiple query heads but collapses all of them down to a single shared key and value projection. This mainly shrinks the KV-cache — the stored keys and values from every previous token that autoregressive generation must keep in memory — which is usually the real memory bottleneck at inference time.
- Grouped-Query Attention (GQA) — a middle ground between full multi-head attention (one KV pair per query head) and MQA (one KV pair total): query heads are split into groups, and each group shares one key/value projection. Llama 2/3 and Mistral use GQA because it recovers most of the quality lost to MQA while keeping most of the KV-cache savings.
None of these techniques change what attention is conceptually — they change how, and how cheaply, the same query/key/value computation gets executed. A model built with GQA, RoPE, and FlashAttention is still, mechanically, doing scaled dot-product attention under the hood; the engineering just makes it affordable at production scale.
| Technique | Complexity | Exact or approximate? | Primary benefit |
|---|---|---|---|
| FlashAttention | compute, low memory | Exact | Memory-bandwidth speedup with no quality change |
| Sparse / windowed | Approximate — restricted connectivity | Sub-quadratic cost for long sequences | |
| Linear attention | Approximate — kernel feature map | Constant memory per generated step | |
| MQA / GQA | compute, unchanged | Exact computation, reduced KV capacity | Smaller KV-cache, faster inference |
Reading an Attention Map in Practice
Visualizing a layer’s attention weights as a heatmap — rows and columns both indexed by token, cell intensity equal to weight — is one of the most common diagnostic tools used when inspecting a trained transformer. A few recurring patterns show up often enough to be worth recognizing on sight:
- A strong diagonal — each token attends mostly to itself and its immediate neighbors, typical of layers doing local, syntax-level processing.
- Vertical stripes — one or a few key positions (frequently the first token) receive high attention from almost every query regardless of content; this is the attention-sink effect described above, not a sign those tokens are semantically important.
- Block patterns — in retrieval or multi-document contexts, attention clusters within a passage boundary, indicating the model treats each retrieved chunk as a largely self-contained unit rather than freely mixing across them.
- Diffuse, near-uniform rows — a query with no strong preference among keys, often a sign the layer isn’t doing much discriminative work for that token, or that the relevant signal lives in a different head.
A stylized comparison — a diagonal-heavy pattern typical of a lower, local-syntax layer, against a broader pattern typical of a higher, more semantic layer:
Lower layer (local) Higher layer (semantic)
A B C D E A B C D E
A [# . . . .] A [# . # . .]
B [. # . . .] B [. # . . #]
C [. . # . .] C [# . # . .]
D [. . . # .] D [. . . # .]
E [. . . . #] E [. # . . #]
None of this replaces rigorous interpretability methods, but it’s often the first, fastest signal an engineer checks when a model’s output looks wrong and the question is whether attention is even looking at the right part of the input.
Comparison
| Type | Query source | Key/Value source | Sees future tokens? | Typical use |
|---|---|---|---|---|
| Additive (Bahdanau) attention | Decoder RNN state | Encoder RNN states | Yes | Historical predecessor to dot-product attention in RNN encoder-decoder models |
| Self-attention (encoder) | Same sequence | Same sequence | Yes | Encoding a full input (BERT, encoder side of translation) |
| Causal self-attention (decoder) | Same sequence | Same sequence | No (masked) | Autoregressive generation (GPT-style LLMs) |
| Cross-attention | Decoder sequence | Different (encoder) sequence | N/A — separate sequence | Translation, image captioning, any encoder-decoder or multimodal fusion |
| Multi-head attention | Any of the above | Any of the above | Depends on masking used | Running several attention subspaces in parallel within any of the above types |
| Sparse / linear efficient attention | Same or cross, per configuration | Same or cross, per configuration | Depends on masking used | Long-context and resource-constrained serving, trading connectivity or exactness for sub-quadratic cost |
| Grouped-Query Attention (GQA) | Same sequence, split into groups | Shared per group | Depends on masking used | Inference-optimized decoder-only LLMs (Llama, Mistral) balancing quality against KV-cache size |
Real-World Use Cases
- Machine translation — encoder-decoder Transformers use cross-attention so each output word can draw directly from the most relevant source words, regardless of word-order differences between languages.
- Chat and coding assistants — GPT-, Claude-, and Llama-style models use stacked causal self-attention to generate each token conditioned on everything said so far in the conversation or codebase context.
- Retrieval-augmented generation — after a Retrieval-Augmented Generation (RAG) system fetches candidate passages, the generator’s attention layers decide, token by token, which retrieved sentences actually matter for the current answer.
- Vision Transformers (ViT) — images are split into fixed-size patches treated as a token sequence, and self-attention lets any patch relate to any other patch directly, without the local-neighborhood limitation of convolutions.
- Visual question answering and image captioning — cross-attention lets a text decoder query over image patch embeddings to ground generated words in specific image regions.
- Speech recognition — Whisper-style models apply encoder self-attention over audio spectrogram frames and cross-attention from the text decoder into those audio representations.
- Protein structure prediction — AlphaFold’s Evoformer uses attention over pairs of amino acid residues to model which residues are likely to be spatially close in the folded structure.
- Document summarization — long-document summarizers use attention (often in sparse or windowed form) to identify which sentences across a long input carry the information worth compressing into a summary.
- Recommendation systems — sequence-aware recommenders apply self-attention over a user’s interaction history to weigh which past actions are most predictive of the next one.
- Time-series forecasting — attention-based forecasting models weigh historical time steps unequally instead of assuming a fixed lookback window matters uniformly.
- Long-document legal and contract review — tools that process filings well beyond a few thousand tokens lean on sparse or windowed attention variants to stay within a usable compute and memory budget.
- Code repository assistants — attending across many files at once lets a coding assistant resolve a function call to its definition several files away from the current cursor position.
- Anomaly detection in system logs — attention over sequences of log events highlights which prior events a model considers relevant to flagging the current one as anomalous.
- Drug discovery and molecule generation — graph and sequence transformers apply attention over atoms or molecular fragments to predict properties or propose novel candidate structures.
- Fraud and transaction monitoring — attention over a sequence of a user’s past transactions helps a model weigh which prior behavior is most relevant to scoring the current one as suspicious.
Common Pitfalls
- Mistaking attention weights for explanations. High attention from output to input correlates with influence but is not a causal or complete account of why a model produced its output — gradient-based and ablation-based interpretability methods routinely find important paths through the network that attention weights alone don’t reveal.
- Ignoring the cost. Attention’s compute and memory scale quadratically with sequence length — doubling context length roughly quadruples the cost of a naive implementation, which surprises teams that scale up context windows without budgeting for it.
- Skipping the scaling factor. Implementing attention from scratch and forgetting the scale divisor lets dot-product magnitudes grow with dimensionality, saturating the softmax and stalling gradient flow — a subtle bug that shows up as poor convergence rather than a crash.
- Confusing self-attention with cross-attention. Wiring keys and values to come from the wrong sequence (e.g., feeding decoder states into what should be encoder cross-attention) silently produces a model that ignores its actual conditioning input.
- Forgetting the causal mask during autoregressive training. Leaving a decoder’s self-attention unmasked lets it “see” future tokens during training, producing artificially low training loss that collapses the moment the model has to generate sequentially at inference with no future to peek at.
- Assuming more heads always helps. Beyond a task-dependent point, additional heads mostly get redundant or get pruned away with no quality loss — head count is a real hyperparameter to tune, not a dial to maximize.
- Assuming attention alone captures word order. Attention itself is permutation-invariant — shuffle the input tokens and, absent positional information, the attention computation doesn’t know the difference. Transformers need explicit positional encodings or embeddings layered on top to represent order at all.
- Numerical instability in a naive softmax. Exponentiating large raw scores without first subtracting the row-wise maximum causes floating-point overflow; production implementations always use the numerically stable “subtract-max” softmax variant.
- Assuming a bigger context window guarantees better use of it. Empirical work on long-context models has repeatedly found a “lost in the middle” effect — information placed in the middle of a long context gets attended to less reliably than information at the very start or very end, regardless of the window’s nominal size.
- Treating MQA/GQA as free efficiency. Sharing key/value projections across query heads measurably reduces representational capacity in exchange for a smaller KV-cache; treating the choice as costless rather than a tuned trade-off leads to quality regressions that get blamed on the wrong part of the system.
- Conflating attention with memory. Attention only operates over whatever tokens are in the current context window — it is not a persistent store of facts. A model attending well within a long prompt says nothing about what it can recall once that content falls outside the context window.
- Reading a single head’s map as the model’s overall behavior. With dozens of heads per layer and dozens of layers, cherry-picking one head that shows a clean, interpretable pattern and presenting it as “how the model works” overstates what any one head’s weights actually establish about the full computation.
- Benchmarking attention variants on short sequences only. Efficient-attention techniques (sparse, linear, windowed) are built to pay off at long sequence lengths; testing them only on short inputs where quadratic cost was never a problem hides both their overhead and their real trade-offs.
- Forgetting that positional encoding choice constrains deployment. A model trained with learned absolute positional embeddings capped at, say, 4,096 tokens cannot simply be run on a 16,000-token input — there’s no positional embedding defined past its trained maximum, unlike models built with RoPE or ALiBi.
Related Terms
- Transformer Architecture
- Large Language Model (LLM)
- Tokenization
- Embeddings
- Natural Language Processing (NLP)
- Retrieval-Augmented Generation (RAG)
- Recurrent Neural Network (RNN)
- Backpropagation
Example
Consider translating the English sentence “The trophy didn’t fit in the suitcase because it was too big” into French. The word “it” is genuinely ambiguous in isolation — it could refer to the trophy or the suitcase — and a model without context-aware attention would have no principled way to resolve it. Inside the encoder, self-attention lets the representation being built for “it” attend heavily back to “trophy” and “big,” because those tokens’ key vectors produce high-scoring dot products against “it“‘s query once the model has learned, from vast training data, that size adjectives near “didn’t fit” tend to modify the object that’s too big to fit, not the container. That resolved, context-infused representation of “it” then flows into the decoder.
During generation, the decoder produces the French output one token at a time using two attention operations at each step: causal self-attention over the French tokens generated so far (so “il” or “elle” agrees in gender with whatever noun the model has already decided “it” maps to), and cross-attention over the entire English source sentence’s encoder output, letting the decoder re-check “trophy” versus “suitcase” at the exact moment it needs to choose the correctly gendered pronoun. Neither attention operation exists in isolation — the causal mask keeps generation autoregressive and consistent with what’s already been produced, while cross-attention keeps every generated word grounded in the actual source meaning rather than in a single compressed summary vector.
Scale the same scenario up to a production system translating a 40-page contract instead of one sentence, and nothing about the underlying computation changes — only the engineering around it does. The model still forms queries, keys, and values and still runs the same softmax-weighted lookup; it just does so under RoPE positional encodings so the offsets between clauses pages apart stay meaningful, under grouped-query attention so the KV-cache for such a long document fits in memory, and inside a FlashAttention kernel so the score matrix for tens of thousands of tokens never has to be fully materialized at once. This is precisely the failure mode attention was invented to eliminate: with a fixed-length bottleneck architecture, “it” would have had to be resolved, or guessed, once, early, and baked into a single vector long before the model ever got to choose a French pronoun — with attention, that resolution happens fresh, with full source access, at the moment it’s actually needed, whether the source is one sentence or four hundred pages.
Referenced by