Speculative Decoding
Speculative Decoding is an optimisation technique used to increase token generation speed and token throughput in LLM.
Context & Motivation
- Autoregressive Constraint: Due to the autoregressive nature of the transformer architecture, standard inference generates only one token at a time, meaning one forward pass results in one token.
- Resource Intensity: Bigger models require significantly more resources and time to generate new tokens.
- Efficiency Gap: In most cases, a smaller model could generate the exact same tokens with significantly fewer resources and in far less time.
Core Mechanism
- Draft Generation: A smaller model generates $X$ new candidate tokens along with their confidence scores.
- Parallel Validation: A bigger model validates all $X$ tokens in a single forward pass.
- Throughput Yield:
- Best Case: Yields up to $X$ tokens in a single pass of the larger model.
- Worst Case: Yields 1 token in a single pass (otherwise generates one token and repeats).
Example
- The smaller model generates $5$ tokens ($X_1, X_2, X_3, X_4, X_5$) with corresponding confidence scores ($S_1, S_2, S_3, S_4, S_5$).
- The bigger model verifies all $5$ tokens in a single pass, accepting anywhere from $0$ to $5$ tokens.
Implementation Approaches
- Two-Model Architecture: Using a separate smaller model alongside the main larger model.
- Single-Model Integration: Baking draft-generation capability directly into the model architecture.
Drawbacks & Trade-offs
- Wasted Compute: Across draft iterations, if a large chunk of draft tokens is rejected, compute is wasted on generating those $X$ draft tokens and subsequently rejecting them.
