Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

KV Cache Eviction

An optimization strategy that discards low-attention key-value states to stay within VRAM bounds.

Last reviewed: July 25, 2026

KV cache eviction refers to the strategies used to discard portions of an LLM’s key-value cache when available GPU memory runs out during inference, rather than allowing the cache to grow unbounded as generation continues. The KV cache stores the attention key and value tensors computed for every previous token in a sequence, and it grows linearly with both sequence length and the number of concurrent requests being served — for high-throughput serving with long conversations, it’s often the binding memory constraint, not the model weights themselves.

Why Eviction Is Necessary

A serving system handling many simultaneous users with long conversation histories can run out of GPU memory purely from KV cache growth, even if the model weights fit comfortably. Rather than rejecting new requests outright, eviction policies free up cache space by removing selected tokens’ cached key-value states, trading some potential accuracy loss for continued availability.

Common Eviction Strategies

Simple strategies evict the oldest tokens first (a sliding window approach), which is cheap to implement but can discard genuinely important early context, like system instructions or an important fact mentioned early in a conversation. More sophisticated approaches use the model’s own attention scores as a signal, evicting tokens that historically receive the least attention weight from more recent tokens — the reasoning being that tokens rarely attended to are less likely to be needed going forward. Some production serving systems combine this with explicitly protecting certain tokens (like system prompts) from eviction regardless of attention score.

Where It Fits

KV cache eviction is one of several techniques — alongside quantizing the cache to lower precision and memory-efficient allocation schemes like PagedAttention — that serving frameworks like vLLM use to maximize the number of concurrent requests a given amount of GPU memory can support.

Eviction vs. Quantization: Complementary Techniques

KV cache eviction and KV cache quantization address the same underlying memory pressure from different angles and are often used together rather than as alternatives: quantization reduces the memory footprint of every cached token by storing it at lower precision, while eviction reduces the total number of tokens being cached at all. A serving system under heavy memory pressure might apply both simultaneously — quantizing the cache to 8-bit precision and additionally evicting the least-attended older tokens — squeezing out memory savings from two independent, non-overlapping sources rather than relying on either technique alone to fully solve the problem.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.