In September 2023, researchers at UC Berkeley published a paper describing PagedAttention, a memory management technique that borrowed ideas from operating system virtual memory to solve a stubborn problem in large language model serving. Existing inference systems wasted 60% to 80% of their key-value cache memory through fragmentation and over-reservation, leaving only 20% to 38% available for actual token states. PagedAttention reduced that waste to under 4% by allocating KV cache in fixed-size blocks rather than contiguous chunks, enabling 2x to 4x higher throughput at the same latency. Within months, the vLLM serving engine built on PagedAttention became the de facto standard for multi-user LLM deployments.
Two years later, PagedAttention is no longer an optimization you turn on. It is the foundational substrate that modern serving systems are built on, adopted by vLLM, SGLang, TensorRT-LLM, and Hugging Face TGI. The technique also introduced new complexity, and the KV cache bottleneck remains unsolved in important ways. Here is what production data revealed about what PagedAttention fixed, what it cost, and what still needs solving.
The memory waste baseline
Before PagedAttention, inference engines allocated KV cache memory the way early programming languages allocated arrays: in large contiguous blocks sized for the maximum possible sequence length. For a model like OPT-13B, a single token requires 800 KB of KV cache (2 vectors × 5,120 hidden dimensions × 40 layers × 2 bytes per FP16 value). A full sequence could demand up to 1.6 GB for OPT-13B or 1.7 GB for LLaMA-13B, according to the original paper.
Because engines could not predict how long a conversation would run, they reserved memory for the worst case. Most sequences ended much shorter, leaving reserved memory idle. Internal fragmentation wasted space within each allocation; external fragmentation left gaps between allocations too small to reuse. The Berkeley team measured memory waste between 60% and 80% in production-like workloads, with utilization rates as low as 20.4% for some configurations.
The waste had direct economic consequences. GPU memory, not compute, became the binding constraint for multi-user serving. From the NVIDIA A100 to the H100 generation, FLOPS more than doubled while maximum memory remained capped at 80 GB. Serving providers either ran fewer concurrent users per GPU or paid for more hardware.
How paging changed the economics
PagedAttention applied the same insight that operating systems use for virtual memory: divide memory into fixed-size pages (blocks) and map logical addresses to physical locations through an indirection table. The vLLM implementation allocates KV cache in blocks rather than contiguous arrays, storing each block's physical location in a per-sequence block table. As a sequence generates tokens, the system allocates new blocks on demand. When a sequence finishes, its blocks return to a free pool for immediate reuse.
The result: memory waste dropped to under 4%, occurring only in the last partially filled block of each sequence, according to the vLLM project blog from June 2023. For OPT-13B, vLLM processed 2.2x more concurrent requests than the Orca system and 4.3x more than FasterTransformer at the same latency. Throughput improved by 2x to 4x across benchmarks.
LMSYS Chatbot Arena provided early production validation. Between April and May 2023, the service handled an average of 30,000 requests daily with peaks of 60,000, according to the vLLM team. After adopting vLLM, LMSYS cut the number of GPUs required for serving by 50%. The cost reduction was immediate and measurable.
PagedAttention also enabled memory sharing across requests. Parallel sampling (generating multiple responses to the same prompt) saved 6.1% to 9.8% of memory; beam search saved 37.6% to 55.2%. On the ShareGPT conversational dataset, savings reached 16.2% to 30.5% for parallel sampling and 44.3% to 66.3% for beam search, per the original paper.
Adoption and the new default
Within a year, PagedAttention moved from research prototype to production standard. vLLM became the reference implementation for high-throughput serving. NVIDIA integrated paged KV cache management into TensorRT-LLM. SGLang and Hugging Face Text Generation Inference adopted similar architectures. By 2025, paged memory management was the default substrate for any serving system targeting multi-user workloads, according to analysis from acingai.com and temperature2.com.
The technique proved valuable as context windows expanded. Longer contexts multiply the size of the KV cache, making memory waste more expensive. Models with 32K, 128K, or 1M token windows would be unservable at scale without paged allocation. PagedAttention gave providers a way to offer long-context APIs without prohibitive memory overhead.
Adoption was not universal. Single-user local inference engines often skip paged KV cache management entirely, as noted by zolotukhin.ai. Without batching or memory contention, the indirection overhead of block tables adds complexity for little gain. The technique solves a problem that only exists at scale.
What paging did not solve
PagedAttention eliminated memory waste but introduced new costs. Block-table indirection complicates kernel code. Every attention operation must look up physical block addresses, adding pointer chasing and reducing cache locality. For small batch sizes or short sequences, the overhead can outweigh the memory savings.
The technique also does not address the bandwidth bottleneck. KV cache access remains memory-bound. As context length, batch size, and concurrent users increase, the system must read and write more data per generated token. Paging reduces how much memory you need, but it does not reduce how much memory you touch. Techniques like grouped-query attention, multi-query attention, and cache quantization are still required to lower bandwidth demand, as outlined by techbloat.com.
Alternative approaches have emerged. Microsoft Research published vAttention in 2024, which argues for different trade-offs in how memory is managed and accessed. NVIDIA recently introduced KV Cache Transform Coding, a compression technique that applies JPEG-style transform coding to shrink the cache by up to 20x without modifying model weights. KVTC reduces memory footprint and speeds time-to-first-token by up to 8x by compressing the cache between inference phases, targeting long-context multi-turn scenarios where paging alone is insufficient.
The KV cache problem is not solved. It is managed. PagedAttention gave serving systems a way to avoid catastrophic memory waste, but the cache remains the dominant cost and latency driver for long-context, high-throughput workloads. The next generation of solutions will likely combine paging with compression, quantization, eviction policies, and hardware co-design.
What two years of production revealed
The most important lesson from two years of PagedAttention in production is that memory management is not a feature. It is infrastructure. The technique did not just improve performance; it changed what workloads were economically viable. Providers could serve more users per GPU, offer longer context windows, and reduce costs without sacrificing latency.
The gains were not evenly distributed. High-throughput multi-user serving saw dramatic improvements. Single-user and low-latency workloads saw less benefit or even regressions from indirection overhead. The technique is a solution to a specific problem: memory contention under batching. Where that problem does not exist, paging adds complexity without value.
The broader trend is clear. As models grow and context windows expand, memory will remain the binding constraint. Compute continues to scale faster than memory bandwidth and capacity. The next two years will likely bring more sophisticated cache management: hierarchical paging, predictive eviction, cross-request deduplication, and tighter integration with compression and quantization. PagedAttention was the first step, not the last.
Sources & further reading
- Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention," arXiv:2309.06180, September 2023 Read →
- vLLM project blog, "vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention," June 20, 2023 Read →
- Zolotukhin, "Paged KV cache is the serving fix a single-user local engine can mostly skip," 2026 Read →
- Acing AI, "KV Cache Engineering," acingai.com Read →
- Temperature2, "Did you know: KV cache," July 14, 2026 Read →
- Techbloat, "Overcoming the AI memory bottleneck," techbloat.com Read →
- VentureBeat, "Nvidia says it can shrink LLM memory 20x without changing model weights," venturebeat.com Read →