Quick Contact

✉
fahimkhan20148@gmail.com
📱
+971 507 286 133
Back to Notes
March 30, 2026
llmcachingperformanceinference

Token Caching

Token Caching (often referring to KV Cache or Prompt Caching) is a technique used to speed up LLM inference by storing the keys and values of previously processed tokens.

Key Concepts

  • KV Cache: Stores the Key and Value vectors for each token in the context window to avoid recomputing them for every new token generated.
  • Prompt Caching: Specifically caches the hidden states of common system prompts or long context headers.

Implementations

  • vLLM: Uses PagedAttention to efficiently manage KV caches.
  • LMCache: A specialized system for sharing caches across multiple requests or instances.

#TODO: Expand on LMCache vs standard KV caching.