vLLM
vLLM is a library for high-throughput and memory-efficient LLM serving and inference.
Key Innovations
- PagedAttention: A new attention algorithm that manages attention keys and values (KV cache) more efficiently, similar to virtual memory paging in operating systems.
- Dynamic Batching: Optimizes hardware utilization by grouping requests on the fly.
