<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Serving on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/serving/</link><description>Recent content in Serving on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/serving/index.xml" rel="self" type="application/rss+xml"/><item><title>Vllm</title><link>https://terms-en.ai-term-hub.com/zh/terms/vllm/</link><pubDate>Sat, 18 Jul 2026 11:37:39 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/vllm/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>vLLM（Virtual Large Language Model）是一个旨在加速 LLM 服务的开源库。它引入了 PagedAttention，这是一种受操作系统虚拟内存&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>vLLM 是一个高吞吐量且内存高效的 LLM 推理引擎，利用 PagedAttention 优化 GPU 内存使用。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>PagedAttention&lt;/li>
&lt;li>KV Cache 管理&lt;/li>
&lt;li>推理服务&lt;/li>
&lt;li>吞吐量优化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>高并发 API 服务&lt;/li>
&lt;li>批处理推理处理&lt;/li>
&lt;li>具有成本效益的 LLM 部署&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> vllm &lt;span style="color:#f92672">import&lt;/span> LLM, SamplingParams
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>llm &lt;span style="color:#f92672">=&lt;/span> LLM(model&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#34;facebook/opt-125m&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>prompts &lt;span style="color:#f92672">=&lt;/span> [&lt;span style="color:#e6db74">&amp;#34;Hello, my name is&amp;#34;&lt;/span>, &lt;span style="color:#e6db74">&amp;#34;The capital of France is&amp;#34;&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>sampling_params &lt;span style="color:#f92672">=&lt;/span> SamplingParams(temperature&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">0.8&lt;/span>, top_p&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">0.95&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>outputs &lt;span style="color:#f92672">=&lt;/span> llm&lt;span style="color:#f92672">.&lt;/span>generate(prompts, sampling_params)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tensorrt-nvidia-%E6%8E%A8%E7%90%86%E4%BC%98%E5%8C%96%E5%BA%93/">TensorRT (NVIDIA 推理优化库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tgi-%E6%96%87%E6%9C%AC%E7%94%9F%E6%88%90%E6%8E%A8%E7%90%86%E6%9C%8D%E5%8A%A1%E5%99%A8/">TGI (文本生成推理服务器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/pagedattention-%E5%88%86%E9%A1%B5%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6/">PagedAttention (分页注意力机制)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-serving-llm-%E6%9C%8D%E5%8A%A1%E5%8C%96/">LLM Serving (LLM 服务化)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>