Attention and Language Models
This article starts from the problem a language model is trying to solve, then derives the attention mechanisms used to solve it. We will connect next-token prediction to logits, probabilities, cross-entropy, causal masking, and the components of a modern GPT-style language model. The goal is not to memorize an architecture diagram, but to understand why each computation is present.
1. Language Modeling
1. 语言模型在学习什么?
1.1 Next-Token Prediction
1.1 下一个 Token 预测
A language model assigns probabilities to token sequences. After a tokenizer maps text to $x_1,x_2,\ldots,x_T$, the chain rule writes the probability of the complete sequence as
\[q_\theta(x_1,\ldots,x_T)=\prod_{t=1}^{T}q_\theta(x_t\mid x_{<t}),\]where $x_{<t}=(x_1,\ldots,x_{t-1})$ is the context available before token $t$, and $\theta$ denotes all trainable parameters. The large sequence problem has therefore become one repeated prediction problem: given a prefix, produce a probability distribution over the next token.
At position $t$, the network produces a vector $z_t\in\mathbb R^{\lvert\mathcal V\rvert}$ with one number for each token in vocabulary $\mathcal V$. These raw scores are logits. They are not probabilities and need not sum to one. Softmax converts them into the next-token distribution:
\[q_\theta(v\mid x_{<t})=\frac{e^{z_{t,v}}}{\sum_{u\in\mathcal V}e^{z_{t,u}}}.\]During training, the correct next token is known at every position. The model is penalized by $-\log q_\theta(x_t\mid x_{<t})$ and these terms are averaged across tokens. This is next-token cross-entropy; Use of Information Theory in Learning Theory explains why it is both a prediction loss and an expected coding length.
What must the network do before it can produce useful logits? It must turn each token into a representation that contains the relevant information from its prefix. A token embedding by itself only identifies that token; the embedding for “bank” does not say whether the surrounding sentence concerns money or a river. Attention is the mechanism that makes a token representation depend on its context.
1.2 Autoregressive Training and Generation
1.2 自回归训练与生成
At inference time, GPT cannot produce an unknown continuation all at once. Starting from a prompt $x_1,\ldots,x_m$, it computes $q_\theta(x_{m+1}\mid x_{\le m})$, selects or samples one token, appends that token to the context, and repeats. The distribution at the next step depends on the token just chosen, so that step cannot be computed beforehand. Generation is therefore sequential, and it stops when an end token is selected or another stopping rule is reached.
During training, however, the complete text $[x_1,\ldots,x_T]$ is already available. The model receives $[x_1,\ldots,x_{T-1}]$ and is supervised against the one-position-shifted targets $[x_2,\ldots,x_T]$. This is often described as teacher forcing: every training position receives the actual preceding tokens from the data rather than tokens sampled from the model. An early prediction error therefore cannot corrupt the context used to train later positions.
Because every training token is known before the forward pass, all positions can enter causal self-attention in one matrix. The mask in Section 2.4 preserves the rule that position $t$ may only use positions up to $t$. The model assigns a probability to the correct next token at every position, and these probabilities are trained with next-token cross-entropy; Use of Information Theory in Learning Theory derives exactly what that loss measures.
The distinction is worth making explicit. During generation, if the model has produced “Welcome to,” the next prediction must condition on the entire available prefix, not only on “to.” During training, the shifted sequence and causal mask expose exactly the corresponding prefix at every position. Training is parallel because those prefixes come from known data; generation is sequential because each new prefix contains a token that has not yet been chosen.
2. Self-Attention
2. 自注意力机制
The language-modeling problem above tells us what context is available, but not how to use it. Self-attention gives each token a query, compares that query with the keys of visible tokens, and uses the resulting scores to mix their values. We will begin with one attention head, write it in matrix form, and understand what its weights mean. Sections 2.4 and 2.5 then add causal masking and multiple heads for an autoregressive language model.
2.1 Single-Layer Attention
2.1 单层注意力
The diagram below illustrates single-layer self-attention. Assuming an input sequence $x_{1,2,3}$, an embedding layer generates corresponding embeddings $a_{1,2,3}$ for each token. We then define three matrices, $Q, K, V$, as model parameters. For token embedding $a_1$, we multiply it with matrices $Q$ and $K$ to obtain vectors $q_1$ and $k_1$, respectively. Multiplying these two vectors results in an initial attention score, $at_{11}$ (often denoted as $\alpha$). It’s crucial to understand that $at_{11}$ is a scalar value, not a vector. Applying softmax to all attention scores produces normalized values, denoted as $st_{11}$. Simultaneously, $a_1$ is multiplied with matrix $V$ to yield a value vector $v_1$. Multiplying the normalized qk token with the v token gives us a qkv token, $wt_{11}$. By multiplying $q_1$ with the k vectors derived from the second and third tokens in the input sequence, we obtain $wt_{12}$ and $wt_{13}$. Summing these three tokens yields $b_1$. Repeating this process for $q_2$ and $q_3$ yields $b_2$ and $b_3$, respectively. This constitutes the core algorithm behind self-attention.
Here are some noteworthy points about the diagram:
- The only parameters in a self-attention layer are the three matrices: ${W}^Q,{W}^K,{W}^V$.
- The output token corresponding to each input token is essentially a weighted sum of key-value pairs from all tokens (including itself) and its own query.
- $q, k, v$ are essentially semantic representations of the corresponding token $x$ in the latent space. Token $x$ enters the latent space as $q, k, v$ through ${W}^Q$, $W^K, W^V$.
- The output token does not depend on the hidden state of any previous time step.
The fourth point may raise a question: if every position is computed from the same input matrix, how does the model know token order? Order is supplied explicitly by a position representation, which is developed in Section 3.1.
2.2 Matrix Form
2.2 矩阵形式
Based on the fourth point, this algorithm can be parallelized using matrices. By concatenating the three vectors $e_{1,2,3}$ into a matrix, we get the following diagram. Since our input sequence has 3 tokens, the Inputs matrix on the left has $n = 3$. This input matrix is multiplied by three parameter matrices to obtain $Q, K, V$. Multiplying $Q$ by $K^T$ produces the attention score matrix. In the diagram, self-attention seems to be applied to each of the three tokens ($x_1, x_2, x_3$) separately. However, their embedding vectors can actually be concatenated for parallel processing. Similarly, concatenating $at_{11, 12, 13}$ forms a row of the attention score matrix (think about why it’s a row, not a column?). Applying row-wise softmax to this attention score matrix yields the normalized attention score matrix $A$, where each row sums to 1. The shape of matrix $A$ is $n\times n$. We will delve deeper into the mathematical significance of matrix $A$ later. Finally, multiplying matrix $A$ with $V$ produces the final matrix $Z$, with a shape of $n\times d_v$. This signifies that we have $n$ tokens, with each token now possessing a value of length $d_v$.
Now, let’s examine the official formula for self-attention:
\[AT(Q,K,V)=softmax(\frac{QK^T}{\sqrt{d_k}})V\]Here, $Q, K, V$ represent the hidden state matrices resulting from multiplying the input matrix $X$ with three parameter matrices, respectively. $d_k$ represents the number of columns in the matrix $W_k$.
Why divide by $\sqrt{d_k}$? One unscaled score is a dot product,
\[s_{ij}=q_i^\top k_j=\sum_{m=1}^{d_k}q_{im}k_{jm}.\]For intuition, suppose the components are independent, have mean zero and variance one. Each product $q_{im}k_{jm}$ then has mean zero and variance one, so the sum has
\[\mathbb E[s_{ij}]=0,\qquad \operatorname{Var}(s_{ij})=d_k.\]Its typical magnitude therefore grows like $\sqrt{d_k}$. Large score differences push softmax toward nearly one-hot outputs, where most probabilities—and their gradients—are tiny. Scaling gives
\[\operatorname{Var}\!\left(\frac{s_{ij}}{\sqrt{d_k}}\right)=\frac{1}{d_k}\operatorname{Var}(s_{ij})=1,\]keeping the scale entering softmax roughly stable as the key dimension changes. The independence assumption is only a motivating approximation, but it explains the normalization.
To isolate the scaling effect, the small figure below uses a deliberately simple toy setting. One query compares exactly two independent candidate keys. The query and key coordinates are assumed independent, with zero mean and unit variance, just as in the variance argument above; neither key has a systematic score advantage. The head width is fixed at 64. Under these assumptions, a representative one-standard-deviation positive score gap is about 11.3 before scaling and about 1.41 after scaling; these are the red and blue marker positions.
With two candidates, softmax depends only on the signed score gap—the first key’s score minus the second key’s score. As that gap ranges from negative to positive, the first key’s weight follows exactly a sigmoid curve, not an arbitrary fitted curve. A negative gap gives the first key little weight, zero gap gives the keys equal weight, and a positive gap gives the first key most of the weight. The left panel of the figure draws this complete sigmoid. The right panel draws the local slope of that same sigmoid: how much the attention weight responds to a tiny change in the gap. It is one softmax sensitivity factor through which gradients pass, not the complete loss gradient. The unscaled red marker lies in a flat tail; the scaled blue marker remains in a region with visible slope. Only the gap matters—a large common offset changes neither marker.
2.3 Essence of Self-Attention
2.3 自注意力的本质
Let’s delve into the essence of the official self-attention formula. Firstly, it’s essential to acknowledge that the three matrices $Q,K,V$ are essentially linear transformations of the input matrix $X$, representing $X$ semantically in the latent space. In other words, it’s possible to train the model without the matrices ${W}^Q,{W}^K,{W}^V$, but the complexity would be insufficient, impacting the model’s performance. For clarity, we’ll use $X$ as a toy substitute for all three matrices, while retaining the scale factor. Each entry of $XX^{\mathsf T}$ sums $d_k$ coordinate products. If those products are roughly independent with unit variance, the dot product has variance $d_k$ and a typical magnitude proportional to $\sqrt{d_k}$. Dividing by $\sqrt{d_k}$ keeps the score scale approximately constant as the vector width grows, preventing softmax from becoming artificially sharp merely because more coordinates were added. The toy formula is therefore:
\[AT(X)=softmax\!\left(\frac{XX^{\mathsf T}}{\sqrt{d_k}}\right)X.\]Consider the sentence “Welcome to Starbucks.” If the embedding layer employs simple 2-hot encoding (e.g., “Welcome” is encoded as 1010), we can represent the input matrix $X$ as shown on the left side of the diagram below. Multiplying this matrix with its transpose yields a matrix that’s essentially an attention matrix, as depicted on the right side of the diagram.
What does this attention matrix represent? Examining the first row, we see that this row essentially calculates the similarity between the token “Welcome” and all other tokens in the sentence. The essence of similarity between word vectors is attention. If token A and token B frequently co-occur, their similarity tends to be high. For instance, in the diagram, “Welcome” exhibits high similarity with itself and “Starbucks,” indicating that these two tokens should receive higher attention when inferring the token “Welcome.”
Normalizing this result using softmax gives us the normalized attention matrix shown on the right side of the diagram below. After normalization, this attention matrix becomes a coefficient matrix, ready to be multiplied with the original matrix.
The final step involves right-multiplying the normalized attention matrix $\alpha$ with the input matrix $X$, resulting in the matrix $\hat X$, as illustrated below. What does this step essentially achieve? The highlighted first row of the left matrix will be multiplied and summed with each column of the input matrix $X$ to compute each value in the first row of the output matrix. Since the first row of the $\alpha$ matrix represents the attention values of the token “Welcome” towards all tokens, the first row of the output matrix $\hat X$ becomes the attention-weighted embedding of the token “Welcome.”
In summary, given an input matrix $\mathsf{X}$, self-attention outputs a matrix $\hat X$, which is the attention-weighted semantic representation matrix of the input matrix.
2.4 Query, Key, Value, and Causal Masking
2.4 Query、Key、Value 与 Causal Mask
While the matrices $W^Q, W^K, W^V$ aren’t strictly necessary, and we know that attention weighting can be achieved using $X$ alone, the performance would be suboptimal. It’s natural to wonder how the QKV concept came about and its underlying rationale.
Why choose the names QKV? Q stands for Query, K for Key, and V for Value. The database analogy is useful if applied carefully: a query is compared with searchable keys, while each key is paired with the value that will be returned if that key receives weight. There is one key and one value per token position, so K and V contain the same number of rows. Unlike an ordinary database lookup, the result is usually not one value; it is a weighted mixture of several values.
The following interactive snippet makes the tensor contractions in a minimal implementation explicit. Select either einsum expression to inspect which axes are retained and which axis is summed out.
Shifting the targets is not sufficient by itself. Without a mask, the representation at an early position could attend to later tokens in the same training sequence. It would then use the answer it is supposed to predict—a direct form of data leakage. Causal self-attention enforces the language-modeling factorization by allowing row $i$ to use only columns $j\le i$.
For query $q_t$, this means that only $k_1,\ldots,k_t$ may receive weight and only their paired values $v_1,\ldots,v_t$ may enter the output:
\[o_t=\sum_{s\le t}\alpha_{t,s}v_s, \qquad \alpha_{t,s}=0\ \text{for }s>t.\]In Figure 8, select $q_1$, $q_2$, or $q_3$. Query $q_1$ can use only $(k_1,v_1)$; $q_2$ can use the first two pairs; $q_3$ can use all three. The dashed future paths contribute nothing. Training can compute all three query rows in parallel, but parallel computation does not imply future access.
Switch the figure to Matrix view to follow the complete computation and the shape of every matrix:
\[\boxed{ S=\frac{QK^{\mathsf T}}{\sqrt{d_k}} \quad\longrightarrow\quad \widetilde S=S+M \quad\longrightarrow\quad A=\operatorname{softmax}_{\mathrm{row}}(\widetilde S) \quad\longrightarrow\quad O=AV.}\]The additive mask makes the causal rule concrete:
\[M_{t,s}=\begin{cases} 0,&s\le t,\\ -\infty,&s>t. \end{cases}\]Thus the allowed region is lower triangular, while the upper-triangular future scores are replaced by $-\infty$ before softmax. Figure 16 isolates this pattern.
The dense multiplication may temporarily compute $S_{1,3}=q_1^{\mathsf T}k_3/\sqrt{d_k}$. Computing that number is not yet the same as transmitting information from $v_3$. The mask changes it to $-\infty$, row-wise softmax produces $A_{1,3}=0$, and the multiplication $AV$ therefore assigns $v_3$ coefficient zero in $o_1$. The order matters: if ordinary softmax ran before masking, the future score would enter the denominator, and simply zeroing its weight afterward would leave the allowed weights incorrectly normalized.
The database analogy now receives its final qualification: each output row represents a token after integrating information only from its available prefix, not from the complete sequence. Later rows have a larger searchable prefix than earlier rows. Backpropagation learns Q and K projections that produce useful routing weights and a V projection that produces useful content.
2.5 Multi-Head Attention
2.5 多头注意力
A single head produces one similarity matrix and therefore one way of routing context at each layer. GPT-style language models use multiple heads so the same token can form several attention patterns in parallel—for example, one pattern may retrieve a nearby syntactic cue while another retrieves a distant name. Multi-head attention is the same operation derived above, applied to several learned subspaces.
The diagram below illustrates the structure of multi-head attention. Formally, multi-head attention is defined as:
\[\text{MultiHead}(Q, K, V) = \text{Concat}(\text{head}_1, \ldots, \text{head}_h) W^O, \quad \text{head}_i = \text{Attention}(Q W_i^Q, K W_i^K, V W_i^V)\]where \(W_i^Q \in \mathbb{R}^{d_{\text{model}} \times d_k}\), \(W_i^K \in \mathbb{R}^{d_{\text{model}} \times d_k}\), \(W_i^V \in \mathbb{R}^{d_{\text{model}} \times d_v}\), and \(W^O \in \mathbb{R}^{hd_v \times d_{\text{model}}}\). This mechanism divides the three matrices $W^Q,W^K,W^V $ into multiple smaller matrices. For instance, in a two-head attention setup, $W_Q$ is split into two smaller matrices, $W^{q_1}$ and $W^{q_2}$. Consequently, the $q$ matrix generated from $a_1$ can also be divided into two smaller matrices, $q_{11}$ and $q_{12}$, which we call attention heads. After obtaining multiple heads, the corresponding qkv heads perform single-layer attention separately, resulting in multiple outputs. For example, two heads would yield $b_{11}$ and $b_{12}$ as outputs. These outputs from different heads are then aggregated into a single output vector, $b_1$. Notice a detail in the diagram: $q_{21}$ is not used when calculating $b_1$. Think about why. Since $b_1$ represents the latent space representation of the query $a_1$, it cannot involve the query $a_2$. Refer back to the Query, Key, Value section if this isn’t clear. Once all heads have completed their calculations, an affine matrix $W_o$ is applied to aggregate information from all heads. The shapes of the matrices are indicated in the diagram. Assuming the maximum sentence length is 256 tokens and the embedding size is 1024, the shape of $a_1$ would be (256, 1024). The shapes of other matrices are also shown in the diagram.
Is the sole purpose of adding heads merely to increase the number of parameters? If so, we could simply enlarge the hidden size of $W^Q, W^K, W^V$. Why achieve this by adding heads?
Usually it does not increase the QKV parameter count. If the model width is $d$ and each of $h$ heads has width $d_h=d/h$, then all query projections together still map $d$ dimensions to $h d_h=d$ dimensions. The point is structural: every head has its own softmax distribution. Enlarging one head gives a wider value vector but still only one attention pattern; using $h$ heads produces $h$ independently normalized patterns and lets $W^O$ combine their retrieved information.
During the training of the multi-head attention mechanism, due to differences in parameter initialization, we have $q_{11} \neq q_{12}$. Similarly, we have $st_{111} \neq st_{121}$ and $b_{11} \neq b_{12}$. However, since $b_{11}$ and $b_{12}$ are concatenated, the gradient flow during backpropagation is symmetrical for these two paths. Different initialization methods lead to heads learning different feature selection capabilities.
Analyses of trained language models often find heads that emphasize syntax, local context, delimiters, or rare identifying tokens. This specialization is not hard-coded and is not guaranteed: some heads become redundant. It is an optimization outcome made possible by separate projection matrices and separate softmax maps. The gradients that create these roles are computed by backpropagation; Backpropagation develops that mechanism from the chain rule.
Let’s ponder another question: what are the significant drawbacks of using multi-layer attention (stacking multiple single-layer attention layers) instead of multi-head attention? There are substantial parallelization limitations. Multi-head attention can be easily parallelized because different heads receive the same input and perform the same computations. In contrast, due to the stacked structure of multi-layer attention, upper layers must wait for computations in lower layers to complete before proceeding, hindering parallelization. The time complexity increases linearly with the number of layers. Therefore, from a parallelization standpoint, multi-head attention is often preferred.
More heads are not automatically better. At fixed model width, increasing $h$ makes every head narrower, so head count trades per-head expressivity against the number of distinct attention maps. What makes the computation efficient is that all heads act on the same input and can be evaluated in parallel before concatenation.
3. Embeddings
3. Embeddings
Attention operates on vectors, but a tokenizer initially gives the model only discrete token IDs. The embedding stage must therefore answer two separate questions: what token is this? and where does it occur? Token embeddings represent identity and learned lexical features; position embeddings represent order. Their sum forms the initial residual-stream vector consumed by the first block.
3.1 Position Embeddings
3.1 语言模型如何表示位置?
Self-attention by itself is permutation equivariant: if we reorder the rows of $X$, the output rows are reordered in exactly the same way. The operation can compare token contents, but nothing in $QK^T$ says that one row came before another. A language model, however, must distinguish “dog bites man” from “man bites dog” and must know which tokens belong to the prefix of the current prediction.
We therefore give the model a vector $p_t$ that identifies or relates to position $t$. The simplest integration is additive: if token embedding $e(x_t)$ and position representation $p_t$ both have width $d$, the layer receives
\[h_t^{(0)}=e(x_t)+p_t.\]The vector $p_t$ can be learned directly as a parameter for every allowed position, generated deterministically, or applied to queries and keys as a relative-position operation. The following deterministic construction is useful because its behavior can be inspected exactly.
A classic construction is sinusoidal positional encoding. For a token at position $t$, embedding dimension $d$, and dimension-pair index $i$, even and odd coordinates use sine and cosine at the same frequency:
Equivalently, $\omega_i=10000^{-2i/d}$, so the pair is $[\sin(t\omega_i),\cos(t\omega_i)]$. Here $t$ identifies the token position, while $i$ identifies a pair of embedding coordinates; they are different indices. Notice that the resulting $p_t$ has the same length $d$ as the token embedding. It can therefore be added directly to that semantic embedding. Extending $t$ across the sequence produces a matrix. For example, given “I am a Robot” with 4 tokens and $d=4$, we obtain the following $4\times4$ matrix. Each row is one $p_t$.
The first row corresponds to $t=0$. Every sine coordinate is then $\sin(0)=0$ and every cosine coordinate is $\cos(0)=1$. For $d=4$, the two frequencies are $\omega_0=1$ and $\omega_1=0.01$, so the next row is $[\sin(1),\cos(1),\sin(0.01),\cos(0.01)]$. Repeating this calculation produces the matrix.
How do we interpret this method? It leverages the periodicity of the cosine function. In fact, as long as we can find multiple functions with different periods but similar characteristics (e.g., the waveform of cosine functions is the same), they can theoretically be used for positional encoding. For instance, we can employ binary positional encoding instead of sinusoidal encoding. As shown in the diagram below, assuming a sequence with 16 tokens and an embedding size of 4, we can obtain the positional embedding for each token using binary encoding:
Clearly, the least significant bit changes very rapidly (period of 2), while the most significant bit changes the slowest (period of 16). Each coordinate thus operates at a different scale, and together the bits uniquely identify a position. Binary encoding illustrates the multi-scale idea; sinusoidal encoding gives a smooth version. GPT-style language models may instead learn an absolute position embedding or encode relative displacement by rotating queries and keys. The latter construction is developed in RoPE and M-RoPE.
3.2 Token Embeddings
3.2 Token Embedding 如何表示词元?
Let the vocabulary contain $\lvert\mathcal V\rvert$ tokens and let the model width be $d$. A learned embedding table
\[E\in\mathbb R^{\lvert\mathcal V\rvert\times d}\]stores one $d$-dimensional row for every token ID. If the tokenizer emits ID $x_t$, the lookup $e(x_t)=E[x_t]$ selects that row. This is equivalent to multiplying a one-hot vector by $E$, but a lookup avoids constructing a mostly zero vector.
The table is learned with the rest of the model. Tokens that play similar roles may acquire useful geometric relationships, but the lookup alone is context-free: the same token ID always starts from the same row. Attention later turns this base vector into a context-dependent representation. This distinction is why the word “bank” can begin with one token embedding yet end with a different hidden representation in a financial sentence than in a sentence about a river.
Many language models also reuse the same table at the output. With weight tying, a final hidden state $h_t\in\mathbb R^d$ produces vocabulary logits such as
\[z_t=Eh_t+b.\]The input side selects a row from $E$; the output side compares $h_t$ with every row at once. Sharing these weights reduces parameters and connects the space used to read tokens with the space used to predict them.
4. Feed-Forward Network
4. 前馈网络
Attention moves information between token positions by mixing value vectors. The feed-forward network performs a different job: it transforms the features within each position. The same two-layer MLP is applied independently to every row $x_t$ of the sequence matrix:
\[\operatorname{FFN}(x_t)=W_2\,\phi(W_1x_t+b_1)+b_2.\]The first projection expands the model width $d$ to a larger hidden width $d_{\mathrm{ff}}$; the activation $\phi$ adds nonlinearity; the second projection returns to width $d$ so the result can re-enter the residual stream. Without this nonlinear feature transformation, attention would mainly route and average existing features rather than construct richer ones. The choice of $\phi$ is itself an architecture decision. A common choice is GELU, explained with the corresponding implementation in Section 5.3; many modern language models instead use a gated variant.
The parameter count explains why this visually simple component matters. Ignoring biases, the attention projections $W^Q,W^K,W^V,W^O$ contain about $4d^2$ parameters. A plain FFN contains $2d\,d_{\mathrm{ff}}$. With the common choice $d_{\mathrm{ff}}=4d$, that is $8d^2$: roughly twice the parameters of attention in the same block. These weights also have to be stored and trained; LLM Optimization Basics: Memory follows their contribution to training memory.
5. Architecture Designs
5. Architecture Designs
Attention and the FFN specify the two main transformations, but they do not by themselves make a deep language model trainable or turn its final vector into a probability distribution. Residual connections preserve an information path, LayerNorm controls feature scale, GELU supplies a nonlinear gate inside the FFN, and softmax normalizes scores where the model must choose among alternatives.
The following minimal model places every component on one depth-ordered path. Inside the expanded block, gray arrows are learned transformations and orange arrows are the two residual identity paths. Click any node to inspect what enters and leaves it.
5.1 Residual Connections
5.1 残差连接
Residual connections follow the same one-per-sublayer pattern as the normalization modules, but they are not normalization. The attention sublayer has one skip path from $x^{(\ell)}$ to its addition; the FFN sublayer has a second skip path from $a^{(\ell)}$ to its addition. The two patterns are structurally identical, but they carry different states and surround different learned functions:
\[a^{(\ell)}=x^{(\ell)}+\Delta_{\mathrm{attn}}^{(\ell)}, \qquad x^{(\ell+1)}=a^{(\ell)}+\Delta_{\mathrm{ffn}}^{(\ell)}.\]GPT-style language models stack many attention and FFN sublayers. If every sublayer replaced its input entirely, information and gradients would have to survive a long chain of transformations. A residual connection instead asks each sublayer to learn a change to a persistent residual stream. In the common pre-norm arrangement,
\[y=x+f(\operatorname{LayerNorm}(x)),\]where $f$ is either causal self-attention or an FFN.
The addition creates a direct identity path. Locally, its Jacobian contains an identity term,
\[\frac{\partial y}{\partial x}=I+\frac{\partial f(\operatorname{LayerNorm}(x))}{\partial x}.\]Consequently, both activations and gradients have a route through the block that does not depend entirely on the learned transformation $f$. This does not guarantee perfect optimization, but it makes deep stacks far easier to train. Backpropagation explains how the Jacobians compose; Basics of Optimizers explains how the resulting gradients update parameters.
5.2 LayerNorm
5.2 LayerNorm
Where does LayerNorm appear? A language-model block usually contains two learned sublayers—attention and an FFN—and each has its own residual addition. In the common pre-norm arrangement, normalization appears immediately before each learned sublayer:
\[\begin{aligned} a^{(\ell)} &=x^{(\ell)} +\operatorname{Attention}_{\ell} \!\left(\operatorname{LN}^{\mathrm{attn}}_{\ell}(x^{(\ell)})\right),\\ x^{(\ell+1)} &=a^{(\ell)} +\operatorname{FFN}_{\ell} \!\left(\operatorname{LN}^{\mathrm{ffn}}_{\ell}(a^{(\ell)})\right). \end{aligned}\]Read the first line from the inside outward: normalize $x^{(\ell)}$, pass the normalized vector through attention, then add the result back to the original, unnormalized $x^{(\ell)}$. The sum is $a^{(\ell)}$. The second line repeats the same pattern for the FFN. Thus there are normally two separate normalization modules per block, with distinct learned $\gamma$ parameters (and $\beta$ when the chosen norm includes a bias). There are also two residual additions—not one residual connection wrapped around the entire attention-plus-FFN block.
The identity branch never passes through LayerNorm:
\[\underbrace{x}_{\text{identity path}} \; + \; \underbrace{f(\operatorname{LN}(x))}_{\text{normalized learned path}}.\]This is what “pre-norm” means: LayerNorm is before attention or the FFN on the learned branch, while the residual addition happens after that sublayer. After the last block, many language models apply one additional final normalization before the vocabulary projection:
\[h_{\mathrm{final}}=\operatorname{LN}_{f}(x^{(L)}),\qquad z=W_{\mathrm{vocab}}h_{\mathrm{final}}+b.\]The exact placement is an architecture choice. An older post-norm arrangement instead uses $\operatorname{LN}(x+f(x))$: first run the sublayer, add the residual, and then normalize the sum. Some architectures add more norms or replace LayerNorm with RMSNorm. Unless stated otherwise, the equations in this article use the pre-norm pattern above.
Why is normalization needed in the first place? Temporarily abstract either attention or the FFN as one learned function $f_\ell$. The residual stream is repeatedly modified as sublayers are stacked:
\[x^{(\ell+1)}=x^{(\ell)}+f_\ell(x^{(\ell)}).\]Nothing in this addition guarantees that the typical magnitude or common offset of the coordinates stays fixed. Consider a deliberately simple two-dimensional example. Let the initial residual vector for one token be
\[x^{(0)}=(1,-1),\]and suppose a learned branch happens to amplify its input according to $f_0(x)=2x$. The first residual update gives
\[x^{(1)}=x^{(0)}+f_0(x^{(0)})=(1,-1)+(2,-2)=(3,-3).\]The direction has not changed, but the magnitude is three times larger. If the next branch temporarily has the same gain, then
\[x^{(2)}=(3,-3)+(6,-6)=(9,-9).\]This toy example is intentionally exaggerated; it does not claim that every real layer computes $2x$. It demonstrates the missing constraint: residual addition alone contains no rule saying that $x^{(\ell+1)}$ must have the same scale as $x^{(\ell)}$. Learned matrices can amplify or shrink a vector, and many such updates accumulate through depth.
Why would the next layer care? Ignoring biases, $q=W^Qx$ and $k=W^Kx$ are linear in the hidden vector. Scaling the hidden vectors by 3 therefore scales both $q$ and $k$ by 3, so their dot product $q^\top k$ becomes 9 times larger. Suppose two attention scores before softmax are $(1,0)$. Their weights are approximately
\[\operatorname{softmax}(1,0)\approx(0.731,0.269).\]If larger hidden vectors make the corresponding score difference nine times larger, the scores become $(9,0)$ and
\[\operatorname{softmax}(9,0)\approx(0.9999,0.0001).\]The attention decision has changed from a soft mixture to an almost one-hot choice even though the score ordering is identical. A common offset can drift as well: adding an update $(10,10)$ to $(1,-1)$ produces $(11,9)$, whose mean is 10 instead of 0. A later linear projection generally reacts to both changes unless its weights happen to cancel them.
LayerNorm removes these two nuisances before the learned branch sees them. Ignoring $\epsilon$, all three vectors
\[(1,-1),\qquad (3,-3),\qquad (11,9)\]normalize to $(1,-1)$: the second differs only in scale, while the third differs only by a common offset. The next sublayer can therefore respond to the relative pattern of coordinates without first having to compensate for whichever magnitude and offset earlier layers happened to produce. The general formula below performs exactly this operation in $d$ dimensions.
What are $d$ and a “coordinate”? A token must be represented by numbers before a neural network can process it. If the model width—also called $d_{\text{model}}$ or the hidden size—is $d$, then one token is represented by a vector containing exactly $d$ numbers:
\[x_t=(x_{t,1},x_{t,2},\ldots,x_{t,d})\in\mathbb R^d.\]Each scalar $x_{t,i}$ is one coordinate, meaning the $i$-th component or slot of that vector. A coordinate is not another token. It is one axis in the model’s learned representation space, and it usually does not have a simple human-assigned meaning by itself. At the input, $d$ is the token-embedding dimension. The attention and FFN outputs are projected back to the same width, so inside later blocks $d$ is also the dimension of each token’s hidden-state or residual-stream vector.
For example, consider $T=3$ tokens and model width $d=4$:
\[X= \begin{bmatrix} 0.2 & -0.1 & 0.7 & 0.4\\ -0.3 & 0.8 & 0.5 & -0.2\\ 0.9 & 0.1 & -0.4 & 0.6 \end{bmatrix} \in\mathbb R^{3\times4}.\]The three rows correspond to three token positions. The four entries in one row are that token’s four coordinates. Thus $x_{2,3}=0.5$ means “coordinate 3 of token 2,” not “token 3.”
Which values are normalized together? LayerNorm treats every row independently. For token 2 above, it uses only $(-0.3,0.8,0.5,-0.2)$ and computes
\[\mu_2=\frac{-0.3+0.8+0.5-0.2}{4}=0.2, \qquad \sigma_2^2=\frac{(-0.5)^2+(0.6)^2+(0.3)^2+(-0.4)^2}{4}=0.215.\]It does not mix these numbers with the other two rows. More generally, for the vector $x_t=(x_{t,1},\ldots,x_{t,d})$ at position $t$, it computes
\[\mu_t=\frac{1}{d}\sum_{i=1}^{d}x_{t,i}, \qquad \sigma_t^2=\frac{1}{d}\sum_{i=1}^{d}(x_{t,i}-\mu_t)^2.\]It therefore does not average over the $T$ tokens, over other sequences, or over the batch. Every token supplies its own mean and variance by averaging the $d$ coordinates within its own row.
The wording is easy to misread: LayerNorm does normalize every token, but it normalizes each token separately. To normalize one $d$-dimensional token vector, it needs a summary of that vector’s location and scale, so it computes the mean and variance of the $d$ coordinates. Averaging coordinates does not assert that they have identical semantics. It treats the entire hidden vector as one geometric object; $\gamma_i$ and $\beta_i$ still remain different for every coordinate.
What if we averaged along the token axis instead? For each coordinate $i$, we would compute something like
\[\mu_i^{\mathrm{tokens}}=\frac{1}{T}\sum_{t=1}^{T}x_{t,i}.\]This is a valid but different operation: coordinate $i$ of token 2 would now depend on coordinate $i$ of every other token. Changing the sentence, its length, or its padding would change token 2’s normalized representation. More seriously, using all $T$ positions in a causal language model would allow an early token’s normalization statistics to depend on future tokens. Restricting the statistics to the prefix avoids leakage, but then the normalization of an already processed token changes whenever the prefix grows, which conflicts with caching its hidden state during generation. LayerNorm avoids all of this: a token’s normalized vector depends only on that token’s current $d$ coordinates.
Why is LayerNorm a line here, while BatchNorm is a plane? A line or plane is not an intrinsic property of either method. It simply counts how many tensor indices are allowed to vary inside one normalization group:
- LayerNorm fixes the batch index $b$ and token index $t$. Only the coordinate index $i=1,\ldots,d$ varies, so its $d$ values lie along one axis: a line.
- The sequence-style BatchNorm shown in the figure fixes coordinate $i$. Both $b=1,\ldots,B$ and $t=1,\ldots,T$ vary, so its $BT$ values span two axes: a plane.
This also explains why the geometry can change with the tensor layout. For an ordinary matrix $A\in\mathbb R^{B\times d}$, BatchNorm fixes $i$ and varies only $b$, so its group is a line, not a plane. A sequence implementation that normalizes each position separately across the batch would likewise fix $(t,i)$ and vary only $b$. The formulas—which indices are fixed and which are reduced—are the definition; the drawing is only their geometry in this particular three-axis tensor.
The normalized coordinates are then
\[\hat x_{t,i}=\frac{x_{t,i}-\mu_t}{\sqrt{\sigma_t^2+\epsilon}}, \qquad y_{t,i}=\gamma_i\hat x_{t,i}+\beta_i.\]These equations describe three steps. Subtracting $\mu_t$ centers the coordinates, so their mean becomes zero. Dividing by the standard deviation makes their variance approximately one; it is only approximate because the small constant $\epsilon$ prevents division by zero when the coordinates are almost identical. Finally, $\gamma,\beta\in\mathbb R^d$ apply a learned scale and offset to each feature coordinate. The symbol $\odot$ in the compact vector formula denotes coordinate-wise multiplication.
A small example makes the operation concrete. For $x=(1,2,3)$, the mean is $2$ and the variance is $2/3$. Ignoring $\epsilon$,
\[\hat x =\frac{(1,2,3)-2}{\sqrt{2/3}} \approx(-1.225,0,1.225).\]Now replace the input by $x’=100x+50=(150,250,350)$. Its mean and standard deviation change by the same shift and positive scale, so normalization produces the same $\hat x$. LayerNorm preserves the relative pattern “first coordinate below the token mean, second at the mean, third above it” while discarding that token’s overall offset and magnitude. It does not claim that the three coordinates mean the same thing, nor does it standardize one feature over the dataset.
Why introduce $\gamma$ and $\beta$ after deliberately removing scale and offset? Normalization removes a different, input-dependent mean and variance for every token. In contrast, $\gamma_i$ and $\beta_i$ are learned parameters for feature $i$, shared across tokens. They let the model decide that one feature should generally be amplified, suppressed, or shifted without forcing the next sublayer to absorb uncontrolled example-by-example scale. They cannot reconstruct the discarded mean and norm of each input token. In a pre-norm block, that information is not lost from the model as a whole because the unnormalized $x$ remains available on the residual identity path.
This distinction is visible in the pre-norm computation:
\[u^{(\ell)}=\operatorname{LayerNorm}(x^{(\ell)}), \qquad x^{(\ell+1)}=x^{(\ell)}+f_\ell(u^{(\ell)}).\]LayerNorm controls the input seen by $f_\ell$; it does not force the residual stream after the addition to have mean zero and variance one. The identity branch carries $x^{(\ell)}$ forward unchanged, while the learned branch operates on $u^{(\ell)}$ at a controlled scale. In backpropagation, this also leaves a direct identity contribution to the gradient, as Section 5.1 showed. LayerNorm therefore improves the conditioning of each learned branch, while the residual connection preserves a clean route through the stack. It moderates scale sensitivity; it does not by itself guarantee that activations or gradients can never grow.
What, then, is BatchNorm? For an ordinary activation matrix $A\in\mathbb R^{B\times d}$ containing $B$ examples, BatchNorm fixes a coordinate $i$ and estimates that coordinate’s mean and variance across the batch:
\[\mu_i^{\mathrm{BN}}=\frac{1}{B}\sum_{b=1}^{B}A_{b,i}, \qquad (\sigma_i^2)^{\mathrm{BN}}=\frac{1}{B}\sum_{b=1}^{B}(A_{b,i}-\mu_i^{\mathrm{BN}})^2.\]LayerNorm fixes an example and reduces across its coordinates; BatchNorm fixes a coordinate and reduces across examples. For sequence activations $X\in\mathbb R^{B\times T\times d}$, a common BatchNorm-style arrangement also pools token positions, giving $BT$ values for each fixed coordinate $i$:
\[\mu_i^{\mathrm{BN}}=\frac{1}{BT}\sum_{b=1}^{B}\sum_{t=1}^{T}X_{b,t,i}.\]The exact extra axes depend on the layer and tensor layout; the defining distinction is that BatchNorm estimates per-coordinate statistics from a collection of examples, whereas LayerNorm obtains per-token statistics from one vector.
A small matrix makes the distinction concrete. Ignore $\epsilon$, $\gamma$, and $\beta$, and let three examples have two coordinates:
\[A=\begin{bmatrix}1&10\\3&20\\5&30\end{bmatrix}.\]BatchNorm works down each column. The column means are $(3,20)$; after division by the corresponding standard deviations, its output is approximately
\[\operatorname{BN}(A)= \begin{bmatrix}-1.225&-1.225\\0&0\\1.225&1.225\end{bmatrix}.\]Each number now says where this example sits relative to the other examples for the same coordinate. LayerNorm instead works across each row. Both entries in every row are equally far below and above that row’s mean, so
\[\operatorname{LN}(A)= \begin{bmatrix}-1&1\\-1&1\\-1&1\end{bmatrix}.\]Each number now describes the relative pattern of coordinates within this example. BatchNorm compares examples coordinate by coordinate; LayerNorm compares coordinates inside one example.
During training, BatchNorm uses the current mini-batch statistics, so the output for one example can change when its batch companions change. A small batch gives a noisy estimate; changing batch size, distributing a batch across devices, or filling sequences with different amounts of padding can change the statistics. During inference, BatchNorm therefore usually substitutes running averages accumulated during training. This creates a second rule whose running estimates must represent the test data well. In a convolutional tensor $(N,C,H,W)$, BatchNorm usually fixes channel $C$ and pools over $(N,H,W)$; the many spatial samples often make those estimates stable, which is one reason BatchNorm remains highly useful in convolutional networks.
LayerNorm uses the current token’s coordinates in both training and inference, so its rule is unchanged and still works with a batch of one. In a causal language model, pooling across token positions would also require careful masking: padding should not enter the statistics, and future positions must not influence earlier ones. This batch independence, together with the absence of padding and future-token coupling, is why LayerNorm—or the closely related RMSNorm—is usually preferred in causal language models.
5.3 GELU
5.3 GELU
If the FFN contained only two linear projections, their composition would still be a single linear map. It could change coordinates but could not build genuinely nonlinear features. The activation $\phi$ between them prevents that collapse.
GELU, or the Gaussian Error Linear Unit, is a common choice:
\[\operatorname{GELU}(x)=x\,\Phi(x),\]where $\Phi(x)$ is the standard normal cumulative distribution function. Rather than using a hard threshold, GELU smoothly scales a value according to its magnitude: large positive values pass almost unchanged, large negative values are strongly suppressed, and values near zero are partially retained. Applied coordinate by coordinate, it lets the expanded FFN decide which intermediate features should influence the projection back to model width.
To implement this definition, use the identity
\[\Phi(x)=\frac12\left(1+\operatorname{erf}\!\left(\frac{x}{\sqrt2}\right)\right),\]where the error function is
\[\operatorname{erf}(u)=\frac{2}{\sqrt\pi}\int_0^u e^{-t^2}\,dt.\]Substitution gives the exact elementwise formula used by the implementation:
\[\boxed{\operatorname{GELU}(x) =\frac{x}{2}\left(1+\operatorname{erf}\!\left(\frac{x}{\sqrt2}\right)\right)}.\]In tensor code, torch.erf, multiplication, and addition all operate independently on every coordinate; no loop over coordinates is required. The interactive implementation below follows the exact formula and then uses it between the two FFN projections.
Many implementations also offer the inexpensive tanh approximation
\[\operatorname{GELU}(x)\approx \frac{x}{2}\left[1+\tanh\!\left( \sqrt{\frac{2}{\pi}}\left(x+0.044715x^3\right) \right)\right].\]Its implementation is available in the second tab of the snippet above.
PyTorch’s F.gelu(x, approximate="none") corresponds to the exact form, while F.gelu(x, approximate="tanh") selects the approximation. Autodifferentiation differentiates either sequence of tensor operations automatically; GELU does not require a custom backward pass.
5.4 Softmax
5.4 Softmax
The network’s last linear layer produces one real-valued score for every possible choice. At a language-model position these scores form
\[z=(z_1,\ldots,z_{\lvert\mathcal V\rvert})\in\mathbb R^{\lvert\mathcal V\rvert}.\]They are logits. A logit may be negative, and the logits need not sum to one, so they cannot yet be probabilities. We need a conversion with three properties: every output should be positive, the outputs should sum to one, and a larger score should yield a larger probability. Softmax supplies that conversion. For $m$ possible choices,
\[\operatorname{softmax}(z)_i=\frac{e^{z_i}}{\sum_{j=1}^{m}e^{z_j}}.\]Exponentiation makes every numerator positive; division by their sum makes the outputs add to one. But the more revealing identity is the ratio between two outputs:
\[\boxed{\frac{q_i}{q_k}=e^{z_i-z_k}}.\]Softmax therefore interprets logit differences as log probability ratios, or log-odds. If $z_i-z_k=\log 2$, choice $i$ receives twice the probability of choice $k$. If the two logits are equal, their probabilities are equal. For example,
\[z=(2,1,0) \quad\Longrightarrow\quad \operatorname{softmax}(z)\approx(0.665,0.245,0.090).\]The first choice is not assigned probability $2$; it is assigned $e^2$ units of positive mass, which are then normalized against the mass assigned to every other choice.
Is a logit a log-probability? Not by itself. Taking the logarithm of the softmax output gives
\[\boxed{\log q_i=z_i-\operatorname{logsumexp}(z)}, \qquad \operatorname{logsumexp}(z)=\log\sum_j e^{z_j}.\]Thus $z_i$ is an unnormalized log-probability. The second term is the shared normalizing constant that turns all logits into actual log-probabilities. This also explains the shift invariance
\[\operatorname{softmax}(z+c\mathbf 1)=\operatorname{softmax}(z).\]Adding the same constant to every logit changes no difference $z_i-z_k$, so it changes no probability. There are infinitely many logit vectors for the same distribution; softmax cares about relative evidence, not the absolute zero of the score scale.
Does softmax appear only at the final logits? No. In the architecture developed here it has two core jobs, and an optional architecture may use it for additional routing. The same formula is reused whenever scores must compete inside a specified set, but the meaning of that set changes.
1. Attention softmax: turning comparisons into a read operation. “Choosing where to read” is only shorthand. Softmax does not inspect the text and retrieve a token by itself. A complete attention head first constructs comparisons between token representations; softmax performs one precise step in the middle: it converts those comparison scores into mixing weights.
Start with just one example and one head. Suppose a layer receives token representations
\[x_1,x_2,\ldots,x_T,\qquad x_t\in\mathbb R^d.\]We focus on position $t$. Its current vector $x_t$ contains what the network has represented at that position so far, but updating it may require information stored at earlier positions. The head creates three learned views of the vectors:
\[q_t=W_Qx_t, \qquad k_s=W_Kx_s, \qquad v_s=W_Vx_s.\]Here $q_t\in\mathbb R^{d_k}$ is the query of the position being updated. Each $k_s\in\mathbb R^{d_k}$ is a key used to decide whether position $s$ is relevant to that query. Each $v_s\in\mathbb R^{d_v}$ is the value, the information position $s$ will contribute if it receives weight. “Query,” “key,” and “value” are learned vector roles—not literal words or manually assigned meanings.
The query is compared with every candidate key by a scaled dot product:
\[r_{t,s}=\frac{q_t^{\mathsf T}k_s}{\sqrt{d_k}}.\]A larger $r_{t,s}$ means that, according to this head’s learned geometry, key $s$ matches query $t$ more strongly. But these $r_{t,s}$ values are still arbitrary real scores: they may be negative, do not sum to one, and cannot yet serve as weights for a controlled average. The factor $1/\sqrt{d_k}$ keeps typical dot-product magnitudes from growing with key dimension; it does not perform the normalization over positions.
Next define which positions the query is allowed to use. In a causal language model, position $t$ may use its prefix but not the future:
\[\mathcal A_t=\{s: s\le t\text{ and }s\text{ is not padding}\}.\]Attention softmax holds the query $t$ fixed and normalizes only across candidates $s\in\mathcal A_t$:
\[\boxed{\alpha_{t,s} =\frac{e^{r_{t,s}}}{\sum_{u\in\mathcal A_t}e^{r_{t,u}}}}, \qquad s\in\mathcal A_t.\]All allowed weights are nonnegative and satisfy $\sum_{s\in\mathcal A_t}\alpha_{t,s}=1$. The head can now perform the actual read:
\[\boxed{o_t=\sum_{s\in\mathcal A_t}\alpha_{t,s}v_s}.\]This last equation is what “reading from the context” means. The head usually does not select one position. It builds a weighted mixture of several value vectors, and that mixture becomes the head’s proposed update for position $t$. The scores decide the relative weights; the values supply the content being mixed.
For example, suppose the query is at position 3 in a four-token sequence and its raw scores are
\[r_{3,:}=(1.2,\ 0.3,\ -0.4,\ 2.0).\]Position 4 has the largest raw score, but it is in the future and is therefore unavailable. After masking, shifting by the largest allowed score, exponentiating, and normalizing, we obtain
\[\begin{aligned} \text{masked scores}&=(1.2,\ 0.3,\ -0.4,\ -\infty),\\ \text{shifted scores}&=(0,\ -0.9,\ -1.6,\ -\infty),\\ \alpha_{3,:}&\approx(0.621,\ 0.253,\ 0.126,\ 0). \end{aligned}\]The read result is therefore
\[o_3\approx0.621v_1+0.253v_2+0.126v_3.\]Nothing from $v_4$ enters the result. Also notice that $0.621$ is not the probability that token 1 is “correct,” nor the probability that token 1 will be generated next. It is an internal routing weight: for this query, in this head, at this layer, it states how much of value vector $v_1$ enters the mixture.
Softmax makes all allowed positions compete. For two allowed positions,
\[\frac{\alpha_{t,s}}{\alpha_{t,u}}=e^{r_{t,s}-r_{t,u}}.\]Increasing one score does not merely increase its own weight; because the denominator is shared, it takes mass away from the other positions. This competition keeps total weight equal to one even when prefix length changes. Without such normalization, raw scores could be negative, the total scale of $o_t$ would vary with score magnitude and number of candidates, and “similarity” would not directly define a controlled mixture. The value vectors may still have different magnitudes, so attention does not force $o_t$ itself to have unit length.
Masked softmax is an equivalent implementation of the restricted set $\mathcal A_t$. Define
\[M_{t,s}=\begin{cases} 0,&s\in\mathcal A_t,\\ -\infty,&s\notin\mathcal A_t. \end{cases}\]Then
\[\alpha_{t,s} =\frac{e^{r_{t,s}+M_{t,s}}} {\sum_u e^{r_{t,u}+M_{t,u}}}.\]Because $e^{-\infty}=0$, a forbidden position gets zero weight and contributes nothing to the denominator. In other words, masking changes the support—the set that may receive weight—before softmax distributes one unit of mass over it.
Is masked softmax a triangular matrix? Masked softmax itself is an operation, so it has no single fixed shape. Its output inherits the shape and sparsity pattern of the mask. In full causal self-attention with $Q=K=T$, the mask is lower triangular:
\[M=\begin{bmatrix} 0&-\infty&-\infty\\ 0&0&-\infty\\ 0&0&0 \end{bmatrix},\]and row-wise masked softmax produces a lower-triangular attention-weight matrix
\[A=\begin{bmatrix} 1&0&0\\ \alpha_{2,1}&\alpha_{2,2}&0\\ \alpha_{3,1}&\alpha_{3,2}&\alpha_{3,3} \end{bmatrix}, \qquad \sum_s\alpha_{t,s}=1.\]Thus it is accurate to call this output lower triangular, but not to say every masked softmax is triangular. Masked softmax is a general operation; the mask supplied by the caller determines its pattern. A key padding mask $[B,K]$ is shared across queries and removes whole key columns, so by itself it does not impose $s\le t$. For example, a Transformer encoder normally uses bidirectional self-attention: every non-padding token may read both earlier and later tokens, and its padding mask is therefore not triangular. During cached one-token generation, the score slice may have shape $1\times K$ rather than being square. Sliding-window or other structural masks produce banded or different patterns. The general rule is simply that the output is zero wherever the mask forbids an entry.
Only after this single-query picture is clear do the full tensor indices help. With $B$ examples, $H$ heads, and $T$ positions, all scores have shape
\[S\in\mathbb R^{B\times H\times T\times T}, \qquad S_{b,h,t,s}=\frac{q_{b,h,t}^{\mathsf T}k_{b,h,s}}{\sqrt{d_k}}.\]For every fixed triple $(b,h,t)$, the slice $S_{b,h,t,:}$ is exactly the one-query score vector just studied. Softmax runs along its final key index $s$—axis=-1—and does not mix different examples, heads, or query positions. Consequently, each example contains $H\times T$ separate attention distributions per layer. Different heads can learn different comparisons and produce different value mixtures. The $H$ head outputs at position $t$ are then concatenated and passed through the output projection $W_O$; only that combined vector becomes the attention sublayer’s update to the residual stream. Efficient kernels may avoid storing the full $T\times T$ score matrix, but mathematically they still compute these same masked row-wise normalizations.
Softmax is the standard way to obtain this smooth, competitive read, although an architecture can define a different attention normalization. Its role should now be precise: query–key projections construct the scores, the mask determines which keys are legal, softmax turns the legal scores into relative weights, and the weighted sum reads the values.
2. Vocabulary softmax: choosing what token may come next. After the final hidden state $h_{b,t}\in\mathbb R^d$, the language-model head produces
\[z_{b,t,v}=h_{b,t}^{\mathsf T}w_v+b_v, \qquad z\in\mathbb R^{B\times T\times\lvert\mathcal V\rvert}.\]For fixed $(b,t)$, softmax is now taken over vocabulary index $v$:
\[q_\theta(v\mid x_{<t}) =\frac{e^{z_{b,t,v}}}{\sum_{u\in\mathcal V}e^{z_{b,t,u}}}.\]This time the normalized numbers really are the model’s categorical probabilities for the next token. Attention softmax competes over source positions; vocabulary softmax competes over token identities. Their denominators are unrelated even though their formulas look identical.
Stable softmax: why subtract the largest logit? The formula is mathematically simple but a literal floating-point implementation can fail. For example, $e^{1000}$ overflows, while $e^{-1000}$ may round to zero. If every numerator underflows, even the denominator becomes zero. Shift invariance gives a numerically stable but mathematically identical implementation:
\[c=\max_jz_j, \qquad q_i=\frac{e^{z_i-c}}{\sum_j e^{z_j-c}}.\]Now the largest exponent is $e^0=1$ and every other exponent lies in $(0,1]$. For instance, $(1000,999)$ is replaced by $(0,-1)$ and produces the same probabilities, approximately $(0.731,0.269)$, without ever computing $e^{1000}$. Stable softmax is not a new distribution or approximation: in exact arithmetic it is precisely ordinary softmax written in a safer form.
Stable and masked softmax are often combined. First add the mask, then find the maximum among the remaining allowed scores, subtract it, exponentiate, and normalize:
\[a_i=z_i+M_i, \qquad c=\max_{j:M_j=0}a_j, \qquad q_i=\frac{e^{a_i-c}}{\sum_j e^{a_j-c}}.\]Which one calls which? The mathematical relationship is most cleanly written as
\[\boxed{\operatorname{masked\_softmax}(z,M) =\operatorname{stable\_softmax}(z+M)}.\]In other words, masking specifies which entries are allowed, and the stable algorithm then computes softmax over those entries safely. The implementation below is interactive.
For training, implementations usually compute stable log_softmax directly rather than first forming probabilities and then taking their logarithms:
If the correct class is $k$, the next-token cross-entropy is
\[\boxed{L=-\log q_k=-z_k+\operatorname{logsumexp}(z)}.\]The first term rewards the correct token’s logit; the second compares it with all vocabulary logits. Raising $z_k$ helps only insofar as it raises the correct token relative to its competitors.
What signal does this send backward? Differentiating the loss with respect to every output logit gives the unusually simple result
\[\boxed{\frac{\partial L}{\partial z_i}=q_i-\mathbf 1[i=k]}.\]For the correct token, the derivative is $q_k-1\le 0$, so gradient descent increases its logit. For every incorrect token, the derivative is $q_i\ge0$, so gradient descent decreases its logit; an incorrect token that currently receives more probability is pushed down more strongly. When the prediction is already correct and confident, $q_k\approx1$ and all these gradients are small. When it is confidently wrong, the correct logit receives a derivative near $-1$ and the wrongly favored token a derivative near $+1$. For a soft target distribution $p$, the same calculation becomes $\partial L/\partial z_i=q_i-p_i$: training moves the predicted distribution toward the target distribution.
Notice that the logit gradients sum to zero. This is another expression of shift invariance: training can change relative logits, but a common shift of every logit has no effect on the loss.
What does temperature change? Dividing logits by a positive temperature $\tau$ changes how strongly their differences matter:
\[q_i(\tau)=\frac{e^{z_i/\tau}}{\sum_j e^{z_j/\tau}}.\]When $\tau<1$, differences are magnified and the distribution becomes sharper; as $\tau\to0^+$ it approaches a one-hot choice at the largest logit. When $\tau>1$, differences shrink and the distribution becomes flatter; as $\tau\to\infty$ it approaches the uniform distribution. Language models are normally trained with the model’s defined scale, while generation-time temperature modifies the sampling distribution without changing the stored model parameters. Temperature does not change which token has the largest logit—it changes how much probability the alternatives retain.
For two choices, softmax depends on only one difference:
\[q_1=\frac{e^{z_1}}{e^{z_0}+e^{z_1}}=\sigma(z_1-z_0).\]This is why binary classification may use either two logits with softmax or one logit with sigmoid: the one-logit convention fixes one reference score and learns their difference. A large vocabulary keeps one logit per token because there are many competing outcomes rather than one binary complement.
Finally, softmax does not guarantee that a model is correct or calibrated. It guarantees only that finite scores become a valid categorical distribution; every finite-logit choice receives strictly positive probability. The learned network and cross-entropy objective must determine whether those probabilities match the data.
5.5 A Complete Forward Pass
5.5 完整的 Forward Pass
We can now assemble the components without hiding the order in which they run. The function below keeps the signature of the running example, but reuses multi_head_attention from Section 2.5 instead of expanding its Q/K/V projections and scaled-attention kernel again.
Before reading the implementation, identify its task from the output path. It accepts token IDs with shape $[B,T]$, pools the $T$ final token representations into one vector, and returns logits with shape $[B,C]$. It is therefore a sequence classifier built from a language-model-style block, not an autoregressive next-token model. This distinction will matter at the final two steps.
The computation has five boundaries worth checking.
First, token lookup and position lookup produce tensors of the same shape, $[B,T,d]$, so they can be added coordinate by coordinate. Their sum is the initial residual stream. The assertion $d\bmod H=0$ records a requirement of the multi-head helper: its $d$ coordinates must be divisible into $H$ equal groups.
Second, the two learned sublayers use the same Pre-LN pattern but do not share states or normalization parameters:
\[\begin{aligned} x_1&=x_0+\operatorname{MHA}(\operatorname{LN}_1(x_0)),\\ x_2&=x_1+\operatorname{FFN}(\operatorname{LN}_2(x_1)). \end{aligned}\]The variables named residual save the unnormalized identity path. attn_input and ff_input belong only to the learned branches. Consequently, x = residual + attended is not interchangeable with x = x + layer_norm(attended, ...): the latter normalizes the branch output and implements neither the Pre-LN equation above nor the Post-LN equation $\operatorname{LN}(x+f(x))$. After the Pre-LN stack, a separate final LayerNorm controls the representation seen by the prediction head; its parameters ln_f_g and ln_f_b must be created during parameter initialization.
Third, the attention helper returns attended with the same $[B,T,d]$ shape as its input. Internally it performs the projections, head split, scaled dot-product attention, concatenation, and output projection already developed in Section 2.5. That matching output width is what makes the residual addition legal. The FFN likewise expands $d\to d_{\mathrm{ff}}$ and returns to $d$ before its residual addition.
Fourth, the mask participates in two different operations. Inside attention, it prevents padding positions from being used as keys. During pooling, mask[..., None] changes $[B,T]$ into $[B,T,1]`, allowing one validity bit to broadcast over all $d$ coordinates of a token. Once padded vectors are removed from the numerator, they must also be removed from the denominator:
Dividing by x.shape[1] would instead divide by the padded length $T$. A three-token sequence padded to length five would be represented as $(x_1+x_2+x_3)/5$ rather than the intended $(x_1+x_2+x_3)/3$, so appending meaningless padding would shrink the representation. The maximum only protects against division by zero; ordinary inputs should normally contain at least one real token.
Finally, head_w maps each pooled $d$-dimensional vector to $C$ class logits. These are raw scores, not probabilities, and a cross-entropy implementation can consume them directly. Setting return_attention=True additionally exposes weights with shape $[B,H,T,T]` for inspection without changing the logits.
To turn the trunk into an autoregressive language model, two changes are required together. First, do not pool: apply a vocabulary projection to every final token vector, producing
\[z=xW_{\mathrm{vocab}}+b,\qquad z\in\mathbb R^{B\times T\times|\mathcal V|}.\]Second, a padding mask alone is insufficient; attention must also enforce the causal condition that query position $t$ cannot read key position $s>t$. Removing pooling changes the output from one prediction per sequence to one next-token prediction per position, while causal masking prevents those predictions from using their own future targets.