Course contents

NLP / Transformers

Self-Attention

The self-attention mechanism, its matrix form, masking and role inside Transformer blocks.

Prerequisites

Linear algebra · Probability · Embeddings

Self-attention lets every token build a context-aware representation by collecting information from every other token in the same sequence. Unlike recurrence, the interaction is expressed as matrix operations and can be computed in parallel.

From embeddings to Q, K and V

For a sequence matrix X∈Rn×dX \in \mathbb{R}^{n \times d}, three learned projections produce queries, keys and values:

Q=XWQ,K=XWK,V=XWV.Q=XW_Q, \qquad K=XW_K, \qquad V=XW_V.

A query represents what a token is looking for; keys describe what each token offers; values contain the information to be aggregated.

Scaled dot-product attention

The core operation is

Attention⁡(Q,K,V)=softmax⁡ ⁣(QK⊤dk)V.\operatorname{Attention}(Q,K,V) = \operatorname{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V.

The scale factor dk\sqrt{d_k} keeps dot products from growing too large as the key dimension increases. Without it, the softmax can enter saturated regions and produce small gradients.

MatrixShapeMeaning
QQn×dkn \times d_kqueries
KKn×dkn \times d_kkeys
VVn×dvn \times d_vvalues
QK⊤QK^\topn×nn \times ntoken-to-token scores

Masks

Decoder-only language models use a causal mask so position ii cannot attend to a future position j>ij>i. Padding masks exclude synthetic padding tokens. In practice, invalid score locations receive a large negative value before softmax.

scores = (q @ k.transpose(-2, -1)) / sqrt(d_k)
scores = scores.masked_fill(~mask, float("-inf"))
weights = scores.softmax(dim=-1)
output = weights @ v

Multi-head attention

Multiple heads repeat the operation in smaller representation subspaces. Their outputs are concatenated and projected once more. Different heads can specialize in different relations, although an individual head should not automatically be interpreted as a clean human-readable concept.

Complexity

The attention matrix is n×nn \times n, so standard self-attention has quadratic time and memory cost in sequence length. Long-context systems therefore depend on careful caching, sparse or windowed patterns, efficient kernels, or architectural alternatives.