<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Serving on English AI Terms Dictionary</title><link>https://terms-en.ai-term-hub.com/en/tags/serving/</link><description>Recent content in Serving on English AI Terms Dictionary</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 18 Jul 2026 11:44:44 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/en/tags/serving/index.xml" rel="self" type="application/rss+xml"/><item><title>Vllm</title><link>https://terms-en.ai-term-hub.com/en/terms/vllm/</link><pubDate>Sat, 18 Jul 2026 10:19:24 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/vllm/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>vLLM (Virtual Large Language Model) is an open-source library designed to accelerate LLM serving. It introduces PagedAttention, a memory management technique inspired by operating system virtual memory, which eliminates memory fragmentation and allows for efficient handling of KV caches. This results in significantly higher throughput and lower latency compared to other serving frameworks like HuggingFace Transformers, making it ideal for production deployments requiring high concurrency.&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>vLLM is a high-throughput and memory-efficient inference engine for Large Language Models, utilizing PagedAttention to optimize GPU memory usage.&lt;/p></description></item></channel></rss>