NLP / Transformers
Self-Attention
The self-attention mechanism, its matrix form, masking and role inside Transformer blocks.
Linear algebra · Probability · Embeddings
Self-attention lets every token build a context-aware representation by collecting information from every other token in the same sequence. Unlike recurrence, the interaction is expressed as matrix operations and can be computed in parallel.
From embeddings to Q, K and V
For a sequence matrix , three learned projections produce queries, keys and values:
A query represents what a token is looking for; keys describe what each token offers; values contain the information to be aggregated.
Scaled dot-product attention
The core operation is
The scale factor keeps dot products from growing too large as the key dimension increases. Without it, the softmax can enter saturated regions and produce small gradients.
| Matrix | Shape | Meaning |
|---|---|---|
| queries | ||
| keys | ||
| values | ||
| token-to-token scores |
Masks
Decoder-only language models use a causal mask so position cannot attend to a future position . Padding masks exclude synthetic padding tokens. In practice, invalid score locations receive a large negative value before softmax.
scores = (q @ k.transpose(-2, -1)) / sqrt(d_k)
scores = scores.masked_fill(~mask, float("-inf"))
weights = scores.softmax(dim=-1)
output = weights @ v
Multi-head attention
Multiple heads repeat the operation in smaller representation subspaces. Their outputs are concatenated and projected once more. Different heads can specialize in different relations, although an individual head should not automatically be interpreted as a clean human-readable concept.
Complexity
The attention matrix is , so standard self-attention has quadratic time and memory cost in sequence length. Long-context systems therefore depend on careful caching, sparse or windowed patterns, efficient kernels, or architectural alternatives.