Decoder
Decoders are a layer of LLM that generate the next token in autoregressive models. Unlike encoders, decoders can't see the full text; when calculating attention, they can only do it from start to end, i.e., left to right for English. In general, decoding is the piece of software that takes the encoded format and turns it into the original format. But in LLMs, decoders are the layer that generates one token at a time. The name decoder comes from the original paper where transformers were used to translate text from language A to language B encoder extracted features from language A, aka encoded it, and the decoder generated the language B, aka decoded the vectors. Decoder-only models are good for language generation tasks.
Examples
- GPT-2 / GPT-3 / GPT-4-class, LLaMA, Mistral: decoder-only; train with next-token prediction.
- Original Transformer / T5: decoder attends to its own past tokens (causal mask) and to encoder outputs (cross-attention).
- Generating
"The cat sat": when predictingsat, attention may useTheandcatonly—not future tokens.
Facts
- Causal (masked) self-attention: position i attends only to positions j ≤ i.
- Autoregressive loop: predict token → append → repeat until EOS or max length.
- Name comes from Attention Is All You Need (2017): decode target language from encoder representations of the source.
- Decoder-only models dominate open-ended generation; KV caching speeds repeated left-to-right steps at inference.
