Quick Contact

✉
fahimkhan20148@gmail.com
📱
+971 507 286 133
Back to Notes
July 31, 2026
llmspeculative-decodingoptimizationinferencethroughput

Speculative Decoding

Speculative Decoding is an optimisation technique used to increase token generation speed and token throughput in LLM.

Context & Motivation

  • Autoregressive Constraint: Due to the autoregressive nature of the transformer architecture, standard inference generates only one token at a time, meaning one forward pass results in one token.
  • Resource Intensity: Bigger models require significantly more resources and time to generate new tokens.
  • Efficiency Gap: In most cases, a smaller model could generate the exact same tokens with significantly fewer resources and in far less time.

Core Mechanism

  • Draft Generation: A smaller model generates $X$ new candidate tokens along with their confidence scores.
  • Parallel Validation: A bigger model validates all $X$ tokens in a single forward pass.
  • Throughput Yield:
    • Best Case: Yields up to $X$ tokens in a single pass of the larger model.
    • Worst Case: Yields 1 token in a single pass (otherwise generates one token and repeats).

Example

  1. The smaller model generates $5$ tokens ($X_1, X_2, X_3, X_4, X_5$) with corresponding confidence scores ($S_1, S_2, S_3, S_4, S_5$).
  2. The bigger model verifies all $5$ tokens in a single pass, accepting anywhere from $0$ to $5$ tokens.

Implementation Approaches

  • Two-Model Architecture: Using a separate smaller model alongside the main larger model.
  • Single-Model Integration: Baking draft-generation capability directly into the model architecture.

Drawbacks & Trade-offs

  • Wasted Compute: Across draft iterations, if a large chunk of draft tokens is rejected, compute is wasted on generating those $X$ draft tokens and subsequently rejecting them.