Course contents

NLP / Transformers

Transformer Architecture

A compact map of the residual, attention and feed-forward structure used by Transformer models.

Prerequisites

Self-Attention · Neural networks

A Transformer block alternates token mixing and channel mixing. Attention moves information between sequence positions; the feed-forward network transforms each position independently.

Pre-normalized block

A common decoder block can be written as

h′=h+Attention⁡(Norm⁡(h)),h' = h + \operatorname{Attention}(\operatorname{Norm}(h)), h′′=h′+MLP⁡(Norm⁡(h′)).h'' = h' + \operatorname{MLP}(\operatorname{Norm}(h')).

Residual paths preserve a direct route for representations and gradients. Normalization keeps activation scales controlled. Repeating the block produces progressively contextualized token states.

Encoder, decoder, encoder–decoder

  • Encoder-only: bidirectional context; useful for representation and understanding tasks.
  • Decoder-only: causal context; the dominant form for autoregressive language models.
  • Encoder–decoder: source encoding plus conditional generation; natural for translation and transformation tasks.

Architecture names describe information flow, not a complete product. Tokenization, objectives, data, inference strategy and evaluation all remain consequential.