<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Memory Management on English AI Terms Dictionary</title><link>https://terms-en.ai-term-hub.com/en/tags/memory-management/</link><description>Recent content in Memory Management on English AI Terms Dictionary</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 18 Jul 2026 11:44:44 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/en/tags/memory-management/index.xml" rel="self" type="application/rss+xml"/><item><title>PagedAttention</title><link>https://terms-en.ai-term-hub.com/en/terms/pagedattention/</link><pubDate>Sat, 18 Jul 2026 10:10:06 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/pagedattention/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>PagedAttention is a technique introduced by the vLLM project to improve the efficiency of Large Language Model inference. It addresses the fragmentation and overhead issues in managing the KV cache, which stores attention states for generating tokens. By treating the KV cache like virtual memory pages, PagedAttention allows for dynamic allocation and sharing of memory blocks between sequences. This results in significant reductions in memory waste and enables higher batch sizes and throughput without requiring hardware changes.&lt;/p></description></item></channel></rss>