<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Performance on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/performance/</link><description>Recent content in Performance on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/performance/index.xml" rel="self" type="application/rss+xml"/><item><title>追踪</title><link>https://terms-en.ai-term-hub.com/zh/terms/tracing/</link><pubDate>Sat, 18 Jul 2026 11:37:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/tracing/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在AI工程背景下，追踪涉及捕获数据在模型或应用程序中流动的详细信息日志，包括每一步的输入、输出、延迟和资源使用情况。这有助于全面监控系统的运行状况。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>追踪是一种记录程序或AI模型推理过程中的执行路径和中间状态的技术，旨在辅助调试和性能优化。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>执行路径&lt;/li>
&lt;li>延迟测量&lt;/li>
&lt;li>调试&lt;/li>
&lt;li>可观测性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>调试大语言模型（LLM）提示链&lt;/li>
&lt;li>优化推理延迟&lt;/li>
&lt;li>审计模型决策过程&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/logging-%E6%97%A5%E5%BF%97%E8%AE%B0%E5%BD%95/">Logging (日志记录)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/profiling-%E6%80%A7%E8%83%BD%E5%89%96%E6%9E%90/">Profiling (性能剖析)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/distributed-systems-%E5%88%86%E5%B8%83%E5%BC%8F%E7%B3%BB%E7%BB%9F/">Distributed Systems (分布式系统)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/opentelemetry-%E5%BC%80%E6%94%BE%E9%81%A5%E6%B5%8B%E6%8A%80%E6%9C%AF/">OpenTelemetry (开放遥测技术)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>吞吐量</title><link>https://terms-en.ai-term-hub.com/zh/terms/throughput/</link><pubDate>Sat, 18 Jul 2026 11:36:49 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/throughput/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能工程中，吞吐量是一个关键的性能指标，用于指示系统容量。对于大语言模型（LLMs），它通常以每秒令牌数（tokens per second）来衡量；对于计算机视觉模型，则以每秒图像数来衡量；对于查询处理，则可能以每秒查询数来表示。高吞吐量意味着系统能够高效地处理大量并发任务。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>吞吐量衡量人工智能系统在给定时间范围内成功处理的数据量或请求数量。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>性能指标&lt;/li>
&lt;li>可扩展性&lt;/li>
&lt;li>批处理&lt;/li>
&lt;li>延迟与吞吐量&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>大语言模型推理服务&lt;/li>
&lt;li>实时视频处理&lt;/li>
&lt;li>高并发API设计&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/latency-%E5%BB%B6%E8%BF%9F/">Latency (延迟)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/batch-size-%E6%89%B9%E5%A4%A7%E5%B0%8F/">Batch Size (批大小)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpu-utilization-gpu%E5%88%A9%E7%94%A8%E7%8E%87/">GPU Utilization (GPU利用率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/model-serving-%E6%A8%A1%E5%9E%8B%E6%9C%8D%E5%8A%A1/">Model Serving (模型服务)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>通义千问2代</title><link>https://terms-en.ai-term-hub.com/zh/terms/qwen2/</link><pubDate>Sat, 18 Jul 2026 11:31:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/qwen2/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>通义千问2代标志着通义千问模型家族的第二次重大升级，引入了架构增强和扩展的训练数据。该版本在多语言支持、长上下文理解以及复杂指令遵循方面提供了更优越的能力。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>通义千问2代（Qwen2）是通义千问大语言模型系列的第二个主要迭代版本，性能得到显著提升。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>模型迭代&lt;/li>
&lt;li>性能提升&lt;/li>
&lt;li>多语言支持&lt;/li>
&lt;li>指令微调&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>高精度问答系统&lt;/li>
&lt;li>多语言文档处理&lt;/li>
&lt;li>高级推理任务&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/qwen-%E9%80%9A%E4%B9%89%E5%8D%83%E9%97%AE/">Qwen (通义千问)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/qwen1-5-%E9%80%9A%E4%B9%89%E5%8D%83%E9%97%AE1-5%E4%BB%A3/">Qwen1.5 (通义千问1.5代)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-benchmarking-%E5%A4%A7%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95/">LLM Benchmarking (大语言模型基准测试)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-architecture-transformer%E6%9E%B6%E6%9E%84/">Transformer Architecture (Transformer架构)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>混合精度训练</title><link>https://terms-en.ai-term-hub.com/zh/terms/mixed_precision_training/</link><pubDate>Sat, 18 Jul 2026 11:26:13 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/mixed_precision_training/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>混合精度训练（MPT）在神经网络训练过程中结合使用半精度（FP16）和全精度（FP32）数据类型。通过使用FP16处理大多数操作，MPT减少了内存占用并提&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种使用16位和32位浮点数进行训练的技術，旨在加速计算并减少内存使用。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>FP16&lt;/li>
&lt;li>FP32&lt;/li>
&lt;li>张量核心&lt;/li>
&lt;li>数值稳定性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>大模型训练&lt;/li>
&lt;li>GPU加速&lt;/li>
&lt;li>内存受限环境&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch.cuda.amp &lt;span style="color:#66d9ef">as&lt;/span> amp
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Example snippet showing automatic mixed precision context&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">with&lt;/span> amp&lt;span style="color:#f92672">.&lt;/span>autocast():
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> output &lt;span style="color:#f92672">=&lt;/span> model(input)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> loss &lt;span style="color:#f92672">=&lt;/span> criterion(output, target)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gradient-scaling-%E6%A2%AF%E5%BA%A6%E7%BC%A9%E6%94%BE/">gradient scaling (梯度缩放)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/amp-%E8%87%AA%E5%8A%A8%E6%B7%B7%E5%90%88%E7%B2%BE%E5%BA%A6/">AMP (自动混合精度)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/half-precision-%E5%8D%8A%E7%B2%BE%E5%BA%A6/">half-precision (半精度)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/optimization-%E4%BC%98%E5%8C%96/">optimization (优化)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Kimi K25</title><link>https://terms-en.ai-term-hub.com/zh/terms/kimi_k25/</link><pubDate>Sat, 18 Jul 2026 11:23:05 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/kimi_k25/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Kimi K25 是月之暗面 Kimi 模型家族中的一个先进迭代版本。它在 Kimi K2 等先前版本的基础上，提供了推理速度等方面的改进。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Kimi K25 是月之暗面推出的后续大型语言模型变体，相比早期版本在性能和效率方面进行了进一步优化。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>模型迭代&lt;/li>
&lt;li>效率&lt;/li>
&lt;li>推理&lt;/li>
&lt;li>月之暗面&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>高并发 API 服务&lt;/li>
&lt;li>实时对话代理&lt;/li>
&lt;li>从长文本中提取数据&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/kimi-k2-kimi-k2%E6%A8%A1%E5%9E%8B/">Kimi K2 (Kimi K2模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/api-integration-api%E9%9B%86%E6%88%90/">API Integration (API集成)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/optimization-%E4%BC%98%E5%8C%96/">Optimization (优化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>压缩张量</title><link>https://terms-en.ai-term-hub.com/zh/terms/compressed_tensors/</link><pubDate>Sat, 18 Jul 2026 11:10:50 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/compressed_tensors/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>压缩张量是深度学习中使用的多维数组，其数值精度（例如从float32降至int8）或稀疏性已降低。这种技术被称为量化或剪枝。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>通过降低数据精度或大小以优化存储和计算效率的张量。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>量化&lt;/li>
&lt;li>稀疏性&lt;/li>
&lt;li>内存优化&lt;/li>
&lt;li>推理速度&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>移动AI应用部署&lt;/li>
&lt;li>边缘设备处理&lt;/li>
&lt;li>大型语言模型服务优化&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Example of converting a tensor to half precision&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>x &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>randn(&lt;span style="color:#ae81ff">10&lt;/span>, &lt;span style="color:#ae81ff">10&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>x_compressed &lt;span style="color:#f92672">=&lt;/span> x&lt;span style="color:#f92672">.&lt;/span>half()
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/quantization-%E9%87%8F%E5%8C%96/">Quantization (量化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/pruning-%E5%89%AA%E6%9E%9D/">Pruning (剪枝)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/model-distillation-%E6%A8%A1%E5%9E%8B%E8%92%B8%E9%A6%8F/">Model Distillation (模型蒸馏)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/float16-%E5%8D%8A%E7%B2%BE%E5%BA%A6%E6%B5%AE%E7%82%B9%E6%95%B0/">Float16 (半精度浮点数)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>缓存</title><link>https://terms-en.ai-term-hub.com/zh/terms/caching/</link><pubDate>Sat, 18 Jul 2026 11:09:42 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/caching/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在AI工程中，缓存通过将最近或频繁的查询结果、模型预测或中间计算保留在快速内存（如RAM）中来优化性能。这减少了昂贵的主数据存储访问需求，从而显著提升系统响应速度。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>缓存是一种将频繁访问的数据存储在临时高速存储层中的技术，旨在降低延迟并减少主数据源的负载。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>延迟降低&lt;/li>
&lt;li>内存优化&lt;/li>
&lt;li>驱逐策略&lt;/li>
&lt;li>命中率&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>存储模型推理结果&lt;/li>
&lt;li>缓存数据库查询输出&lt;/li>
&lt;li>预计算特征嵌入&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> redis
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Simple caching example&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>r &lt;span style="color:#f92672">=&lt;/span> redis&lt;span style="color:#f92672">.&lt;/span>Redis(host&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;localhost&amp;#39;&lt;/span>, port&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">6379&lt;/span>, db&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">0&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">get_prediction&lt;/span>(model_id, input_data):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> cache_key &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#e6db74">f&lt;/span>&lt;span style="color:#e6db74">&amp;#34;pred_&lt;/span>&lt;span style="color:#e6db74">{&lt;/span>model_id&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74">_&lt;/span>&lt;span style="color:#e6db74">{&lt;/span>hash(str(input_data))&lt;span style="color:#e6db74">}&lt;/span>&lt;span style="color:#e6db74">&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> result &lt;span style="color:#f92672">=&lt;/span> r&lt;span style="color:#f92672">.&lt;/span>get(cache_key)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">if&lt;/span> result:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">return&lt;/span> result&lt;span style="color:#f92672">.&lt;/span>decode(&lt;span style="color:#e6db74">&amp;#39;utf-8&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e"># Compute if not cached&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> prediction &lt;span style="color:#f92672">=&lt;/span> model&lt;span style="color:#f92672">.&lt;/span>predict(input_data)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> r&lt;span style="color:#f92672">.&lt;/span>setex(cache_key, &lt;span style="color:#ae81ff">3600&lt;/span>, str(prediction))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">return&lt;/span> prediction
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/redis-%E5%86%85%E5%AD%98%E6%95%B0%E6%8D%AE%E7%BB%93%E6%9E%84%E5%AD%98%E5%82%A8/">Redis (内存数据结构存储)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/memcached-%E5%88%86%E5%B8%83%E5%BC%8F%E5%86%85%E5%AD%98%E7%BC%93%E5%AD%98%E7%B3%BB%E7%BB%9F/">memcached (分布式内存缓存系统)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/performance-tuning-%E6%80%A7%E8%83%BD%E8%B0%83%E4%BC%98/">performance tuning (性能调优)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/database-indexing-%E6%95%B0%E6%8D%AE%E5%BA%93%E7%B4%A2%E5%BC%95/">database indexing (数据库索引)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>异步处理</title><link>https://terms-en.ai-term-hub.com/zh/terms/async_processing/</link><pubDate>Sat, 18 Jul 2026 11:07:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/async_processing/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>异步处理允许软件执行长时间运行的任务（如I/O操作或复杂计算），而不会冻结主应用程序界面或阻塞其他进程。通过事件循环和回调机制，程序可以在等待任务完成的同时继续响应其他请求，从而提高系统的并发能力和响应速度。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种编程范式，任务在主执行线程之外独立运行，允许进行非阻塞操作。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>非阻塞I/O&lt;/li>
&lt;li>事件循环&lt;/li>
&lt;li>并发&lt;/li>
&lt;li>线程&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>实时视频流处理&lt;/li>
&lt;li>同时处理多个API请求&lt;/li>
&lt;li>后台模型训练任务&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> asyncio
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">async&lt;/span> &lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">fetch_data&lt;/span>():
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">await&lt;/span> asyncio&lt;span style="color:#f92672">.&lt;/span>sleep(&lt;span style="color:#ae81ff">1&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">return&lt;/span> &lt;span style="color:#e6db74">&amp;#39;Data&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>asyncio&lt;span style="color:#f92672">.&lt;/span>run(fetch_data())
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multithreading-%E5%A4%9A%E7%BA%BF%E7%A8%8B/">Multithreading (多线程)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/callbacks-%E5%9B%9E%E8%B0%83/">Callbacks (回调)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/promises-%E6%89%BF%E8%AF%BA-%E5%BC%82%E6%AD%A5%E5%AF%B9%E8%B1%A1/">Promises (承诺/异步对象)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/microservices-%E5%BE%AE%E6%9C%8D%E5%8A%A1/">Microservices (微服务)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Accelerated Linear Algebra</title><link>https://terms-en.ai-term-hub.com/zh/terms/accelerated_linear_algebra/</link><pubDate>Sat, 18 Jul 2026 11:04:01 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/accelerated_linear_algebra/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>该领域专注于加速基本的线性代数计算，这些计算是机器学习和科学模拟的核心。通过利用 GPU、TPU 和其他并行处理能力的优势，显著提高了大规模矩阵运算的速度和效率。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>加速线性代数涉及使用 GPU 和 TPU 等硬件加速器优化矩阵运算，以实现高性能计算。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>GPU 计算&lt;/li>
&lt;li>矩阵乘法&lt;/li>
&lt;li>并行处理&lt;/li>
&lt;li>CUDA&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>深度学习模型训练&lt;/li>
&lt;li>科学模拟&lt;/li>
&lt;li>实时图形渲染&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/blas-%E5%9F%BA%E7%A1%80%E7%BA%BF%E6%80%A7%E4%BB%A3%E6%95%B0%E5%AD%90%E7%A8%8B%E5%BA%8F/">BLAS (基础线性代数子程序)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/cudnn-cuda-%E6%B7%B1%E5%BA%A6%E7%A5%9E%E7%BB%8F%E7%BD%91%E7%BB%9C%E5%BA%93/">cuDNN (CUDA 深度神经网络库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tensor-cores-%E5%BC%A0%E9%87%8F%E6%A0%B8%E5%BF%83/">Tensor Cores (张量核心)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>量化</title><link>https://terms-en.ai-term-hub.com/zh/terms/quantization/</link><pubDate>Sat, 18 Jul 2026 11:01:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/quantization/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>量化将高精度浮点数（如 FP32）转换为低精度格式（如 INT8 或 FP16）。这种转换减少了模型的内存使用和计算需求，从而加速推理过程并降低硬件要求。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种模型优化技术，通过降低神经网络计算中数字的精度来减小模型体积并提高速度。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>精度降低&lt;/li>
&lt;li>推理速度&lt;/li>
&lt;li>内存优化&lt;/li>
&lt;li>INT8/FP16&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>边缘设备部署&lt;/li>
&lt;li>移动 AI 应用&lt;/li>
&lt;li>实时推理&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch.quantization &lt;span style="color:#66d9ef">as&lt;/span> quant
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Example of converting a model to quantized format&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>model&lt;span style="color:#f92672">.&lt;/span>eval()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>model&lt;span style="color:#f92672">.&lt;/span>qconfig &lt;span style="color:#f92672">=&lt;/span> quant&lt;span style="color:#f92672">.&lt;/span>get_default_qconfig(&lt;span style="color:#e6db74">&amp;#39;fbgemm&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>quantized_model &lt;span style="color:#f92672">=&lt;/span> quant&lt;span style="color:#f92672">.&lt;/span>prepare(model, inplace&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">False&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>quantized_model &lt;span style="color:#f92672">=&lt;/span> quant&lt;span style="color:#f92672">.&lt;/span>convert(quantized_model, inplace&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#66d9ef">False&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%89%AA%E6%9E%9D-pruning/">剪枝 (Pruning)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%9F%A5%E8%AF%86%E8%92%B8%E9%A6%8F-knowledge-distillation/">知识蒸馏 (Knowledge Distillation)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%B7%B7%E5%90%88%E7%B2%BE%E5%BA%A6%E8%AE%AD%E7%BB%83-mixed-precision-training/">混合精度训练 (Mixed Precision Training)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/onnx/">ONNX&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>延迟</title><link>https://terms-en.ai-term-hub.com/zh/terms/latency/</link><pubDate>Sat, 18 Jul 2026 11:00:35 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/latency/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>延迟衡量AI服务的响应速度，通常以毫秒为单位表示。它包括推理时间、网络传输延迟和处理开销。低延迟对于实时应用至关重要。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>AI系统中请求发起与响应开始之间的时间延迟。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>推理时间&lt;/li>
&lt;li>响应时间&lt;/li>
&lt;li>实时处理&lt;/li>
&lt;li>优化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>语音识别系统&lt;/li>
&lt;li>自动驾驶车辆控制&lt;/li>
&lt;li>实时翻译服务&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/throughput-%E5%90%9E%E5%90%90%E9%87%8F/">throughput (吞吐量)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/inference-%E6%8E%A8%E7%90%86/">inference (推理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/optimization-%E4%BC%98%E5%8C%96/">optimization (优化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/real_time-%E5%AE%9E%E6%97%B6/">real_time (实时)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>分布式训练</title><link>https://terms-en.ai-term-hub.com/zh/terms/distributed_training/</link><pubDate>Sat, 18 Jul 2026 10:59:40 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/distributed_training/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>分布式训练通过在多个 GPU 或节点上并行化计算来加速模型收敛。主要技术包括数据并行（每个工作进程处理数据子集）和模型并行（将模型的不同部分分布在不同设备上），以处理大规模数据集和巨型模型。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>通过将数据或计算任务拆分到多个设备或服务器上，从而训练机器学习模型的方法。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>数据并行&lt;/li>
&lt;li>模型并行&lt;/li>
&lt;li>GPU 集群&lt;/li>
&lt;li>梯度同步&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练大型语言模型&lt;/li>
&lt;li>加速计算机视觉数据集的处理&lt;/li>
&lt;li>减少复杂神经网络的训练时间&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/parallel-computing-%E5%B9%B6%E8%A1%8C%E8%AE%A1%E7%AE%97/">Parallel Computing (并行计算)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpu-%E5%9B%BE%E5%BD%A2%E5%A4%84%E7%90%86%E5%99%A8/">GPU (图形处理器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/horovod-%E5%88%86%E5%B8%83%E5%BC%8F%E8%AE%AD%E7%BB%83%E6%A1%86%E6%9E%B6/">Horovod (分布式训练框架)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/pytorch-ddp-pytorch-%E5%88%86%E5%B8%83%E5%BC%8F%E6%95%B0%E6%8D%AE%E5%B9%B6%E8%A1%8C/">PyTorch DDP (PyTorch 分布式数据并行)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>实时</title><link>https://terms-en.ai-term-hub.com/zh/terms/real_time/</link><pubDate>Sat, 18 Jul 2026 10:57:22 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/real_time/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能领域，实时指的是系统以极低的延迟（通常为毫秒级）处理输入并生成输出的能力。这对于那些需要即时响应、延迟不可接受的应用场景至关重要。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>实时处理指系统在接收到输入后，在严格且保证的时间限制内计算并交付结果。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>延迟&lt;/li>
&lt;li>吞吐量&lt;/li>
&lt;li>推理优化&lt;/li>
&lt;li>确定性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自动驾驶车辆导航&lt;/li>
&lt;li>银行实时欺诈检测&lt;/li>
&lt;li>实时语音翻译&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/latency-%E5%BB%B6%E8%BF%9F/">latency (延迟)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/inference-%E6%8E%A8%E7%90%86/">inference (推理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/edge_computing-%E8%BE%B9%E7%BC%98%E8%AE%A1%E7%AE%97/">edge_computing (边缘计算)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/streaming-%E6%B5%81%E5%BC%8F%E4%BC%A0%E8%BE%93/">streaming (流式传输)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>鲁棒性</title><link>https://terms-en.ai-term-hub.com/zh/terms/robust/</link><pubDate>Sat, 18 Jul 2026 10:54:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/robust/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能中，鲁棒性指模型对抗攻击、数据分布偏移或噪声输入的韧性。一个具有鲁棒性的算法即使在面对干扰时也能继续正确运行。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>描述AI模型或系统在面临噪声、错误或意外输入时保持性能的能力。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>对抗韧性&lt;/li>
&lt;li>泛化&lt;/li>
&lt;li>噪声容忍度&lt;/li>
&lt;li>稳定性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>恶劣天气下的自动驾驶&lt;/li>
&lt;li>带有噪声数据的欺诈检测&lt;/li>
&lt;li>记录不全的医疗诊断&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/overfitting-%E8%BF%87%E6%8B%9F%E5%90%88/">Overfitting (过拟合)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/regularization-%E6%AD%A3%E5%88%99%E5%8C%96/">Regularization (正则化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/adversarial-attacks-%E5%AF%B9%E6%8A%97%E6%94%BB%E5%87%BB/">Adversarial Attacks (对抗攻击)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/generalization-%E6%B3%9B%E5%8C%96/">Generalization (泛化)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>率</title><link>https://terms-en.ai-term-hub.com/zh/terms/rate/</link><pubDate>Sat, 18 Jul 2026 10:54:02 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/rate/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在AI中，“率”最常指学习率，这是一个超参数，控制每次更新模型权重时，根据估计误差对模型进行调整的幅度。一个适……&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>频率或速度的度量，通常指优化中的学习率或令牌生成速度。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>学习率&lt;/li>
&lt;li>优化&lt;/li>
&lt;li>吞吐量&lt;/li>
&lt;li>超参数&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>调整梯度下降优化&lt;/li>
&lt;li>监控API使用限制&lt;/li>
&lt;li>测量推理延迟&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>optimizer &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>optim&lt;span style="color:#f92672">.&lt;/span>SGD(model&lt;span style="color:#f92672">.&lt;/span>parameters(), lr&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">0.01&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/optimizer-%E4%BC%98%E5%8C%96%E5%99%A8/">Optimizer (优化器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/convergence-%E6%94%B6%E6%95%9B/">Convergence (收敛)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speed-%E9%80%9F%E5%BA%A6/">Speed (速度)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/latency-%E5%BB%B6%E8%BF%9F/">Latency (延迟)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>快速</title><link>https://terms-en.ai-term-hub.com/zh/terms/fast/</link><pubDate>Sat, 18 Jul 2026 10:51:19 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/fast/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>“快速”一词描述了人工智能模型中的计算效率，强调快速的推理时间和高效的数据处理能力。这对于实时应用至关重要。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>在人工智能领域，“快速”指为低延迟和高吞吐量处理任务而优化的系统或算法。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>延迟&lt;/li>
&lt;li>吞吐量&lt;/li>
&lt;li>实时处理&lt;/li>
&lt;li>优化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自动驾驶导航&lt;/li>
&lt;li>直播视频流分析&lt;/li>
&lt;li>高频交易算法&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/efficiency-%E6%95%88%E7%8E%87/">Efficiency (效率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/latency-%E5%BB%B6%E8%BF%9F/">Latency (延迟)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/performance-%E6%80%A7%E8%83%BD/">Performance (性能)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>高效</title><link>https://terms-en.ai-term-hub.com/zh/terms/efficient/</link><pubDate>Sat, 18 Jul 2026 10:50:55 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/efficient/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>效率是人工智能中的一项关键指标，用于衡量模型或算法利用可用资源的程度。它包括计算效率（推理和训练的速度）以及资源利用率。高效的AI系统能够在保持高精度的同时，显著降低硬件需求和能耗，这对于大规模部署和实时应用至关重要。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>在人工智能领域，效率是指以最少的时间、内存或计算资源消耗来实现最佳性能。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>计算复杂度&lt;/li>
&lt;li>资源优化&lt;/li>
&lt;li>延迟降低&lt;/li>
&lt;li>模型压缩&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为移动端部署优化大型语言模型&lt;/li>
&lt;li>降低深度学习网络的训练成本&lt;/li>
&lt;li>提高自主系统中的实时推理速度&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/optimization-%E4%BC%98%E5%8C%96/">Optimization (优化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/scalability-%E5%8F%AF%E6%89%A9%E5%B1%95%E6%80%A7/">Scalability (可扩展性)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/latency-%E5%BB%B6%E8%BF%9F/">Latency (延迟)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/throughput-%E5%90%9E%E5%90%90%E9%87%8F/">Throughput (吞吐量)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>推理</title><link>https://terms-en.ai-term-hub.com/zh/terms/inference/</link><pubDate>Sat, 18 Jul 2026 07:44:46 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/inference/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>推理指的是部署阶段，在此阶段使用最终确定的模型对未见过的数据进行决策或预测。与更新权重的训练不同，推理消耗计算资源以产生结果。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>训练好的模型处理新数据以生成预测或输出的阶段。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>预测&lt;/li>
&lt;li>延迟&lt;/li>
&lt;li>吞吐量&lt;/li>
&lt;li>部署&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>银行交易中的实时欺诈检测&lt;/li>
&lt;li>实时聊天交互中生成响应&lt;/li>
&lt;li>自动驾驶系统中的图像分类&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>model&lt;span style="color:#f92672">.&lt;/span>eval()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">with&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>no_grad():
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> output &lt;span style="color:#f92672">=&lt;/span> model(input_tensor)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> prediction &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>argmax(output, dim&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AE%AD%E7%BB%83/">训练&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%BB%B6%E8%BF%9F%E4%BC%98%E5%8C%96/">延迟优化&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%89%B9%E5%A4%84%E7%90%86/">批处理&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%A8%A1%E5%9E%8B%E6%9C%8D%E5%8A%A1/">模型服务&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>