NLP / Transformers
Transformer Architecture
A compact map of the residual, attention and feed-forward structure used by Transformer models.
Prerequisites
Self-Attention · Neural networks
A Transformer block alternates token mixing and channel mixing. Attention moves information between sequence positions; the feed-forward network transforms each position independently.
Pre-normalized block
A common decoder block can be written as
Residual paths preserve a direct route for representations and gradients. Normalization keeps activation scales controlled. Repeating the block produces progressively contextualized token states.
Encoder, decoder, encoder–decoder
- Encoder-only: bidirectional context; useful for representation and understanding tasks.
- Decoder-only: causal context; the dominant form for autoregressive language models.
- Encoder–decoder: source encoding plus conditional generation; natural for translation and transformation tasks.
Architecture names describe information flow, not a complete product. Tokenization, objectives, data, inference strategy and evaluation all remain consequential.