<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Metrics on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/metrics/</link><description>Recent content in Metrics on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/metrics/index.xml" rel="self" type="application/rss+xml"/><item><title>句子相似度</title><link>https://terms-en.ai-term-hub.com/zh/terms/sentence_similarity/</link><pubDate>Sat, 18 Jul 2026 11:33:07 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/sentence_similarity/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>句子相似度衡量两个不同句子之间的语义重叠程度。它超越了词汇匹配，旨在理解含义、上下文和意图。这通常通过计算向量之间的距离来实现。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种量化两个句子在语义上相似程度的指标或任务，通常表示为数值分数。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>余弦相似度&lt;/li>
&lt;li>语义等价&lt;/li>
&lt;li>向量距离&lt;/li>
&lt;li>意义表示&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>论坛中的重复问题检测&lt;/li>
&lt;li>同义句识别&lt;/li>
&lt;li>信息检索和文档聚类&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic-search-%E8%AF%AD%E4%B9%89%E6%90%9C%E7%B4%A2/">Semantic search (语义搜索)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embeddings-%E5%B5%8C%E5%85%A5/">Embeddings (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-inference-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E6%8E%A8%E7%90%86/">Natural Language Inference (自然语言推理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/word-embeddings-%E8%AF%8D%E5%B5%8C%E5%85%A5/">Word Embeddings (词嵌入)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>MAUVE</title><link>https://terms-en.ai-term-hub.com/zh/terms/mauve/</link><pubDate>Sat, 18 Jul 2026 11:24:58 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/mauve/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>MAUVE 是一种统计度量，旨在评估生成语言模型的输出在多大程度上类似于人类语言使用习惯。与简单的困惑度分数不同，MAUVE 使用虚拟嵌入&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>MAUVE（基于虚拟嵌入的测量对齐）是一种用于自然语言处理的指标，用于评估生成文本分布与人类写作文本分布之间的对齐程度。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本生成评估&lt;/li>
&lt;li>分布匹配&lt;/li>
&lt;li>虚拟嵌入&lt;/li>
&lt;li>语言自然度&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估类GPT模型的输出&lt;/li>
&lt;li>微调语言模型以生成类人文本&lt;/li>
&lt;li>基准测试生成式AI的性能&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/perplexity-%E5%9B%B0%E6%83%91%E5%BA%A6/">Perplexity (困惑度)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bleu-score-bleu%E5%88%86%E6%95%B0/">BLEU Score (BLEU分数)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/language-modeling-%E8%AF%AD%E8%A8%80%E5%BB%BA%E6%A8%A1/">Language Modeling (语言建模)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/generative-ai-%E7%94%9F%E6%88%90%E5%BC%8Fai/">Generative AI (生成式AI)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Inception Score</title><link>https://terms-en.ai-term-hub.com/zh/terms/inception_score/</link><pubDate>Sat, 18 Jul 2026 11:22:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/inception_score/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Inception Score（IS）是一种引入用于评估生成对抗网络（GANs）及其他生成模型性能的统计度量。它结合了两个因素：图像质量（清晰度）和多样性。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种用于评估生成图像质量的指标，通过衡量图像的清晰度和多样性来实现。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>生成模型&lt;/li>
&lt;li>图像质量&lt;/li>
&lt;li>多样性度量&lt;/li>
&lt;li>GAN评估&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估GAN性能&lt;/li>
&lt;li>比较生成模型架构&lt;/li>
&lt;li>基准测试图像合成质量&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/fr%C3%A9chet-inception-distance-%E5%BC%97%E9%9B%B7%E6%AD%87-inception-%E8%B7%9D%E7%A6%BB/">Fréchet Inception Distance (弗雷歇 Inception 距离)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/precision-and-recall-%E7%B2%BE%E7%A1%AE%E7%8E%87%E5%92%8C%E5%8F%AC%E5%9B%9E%E7%8E%87/">Precision and Recall (精确率和召回率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/generative-adversarial-networks-%E7%94%9F%E6%88%90%E5%AF%B9%E6%8A%97%E7%BD%91%E7%BB%9C/">Generative Adversarial Networks (生成对抗网络)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/image-synthesis-%E5%9B%BE%E5%83%8F%E5%90%88%E6%88%90/">Image Synthesis (图像合成)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>二分类器评估</title><link>https://terms-en.ai-term-hub.com/zh/terms/evaluation_of_binary_classifiers/</link><pubDate>Sat, 18 Jul 2026 11:16:37 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/evaluation_of_binary_classifiers/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>该领域涉及分析准确率、精确率、召回率、F1分数以及接收者操作特征曲线下面积（AUC-ROC）等指标。它有助于确定模型在区分正负样本方面的表现，并揭示模型在不同阈值下的权衡情况。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>评估预测两种可能结果之一的机器学习模型性能的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>混淆矩阵&lt;/li>
&lt;li>精确率-召回率权衡&lt;/li>
&lt;li>ROC曲线&lt;/li>
&lt;li>F1分数&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>医学疾病筛查&lt;/li>
&lt;li>垃圾邮件过滤&lt;/li>
&lt;li>信用风险评估&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> sklearn.metrics &lt;span style="color:#f92672">import&lt;/span> classification_report
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(classification_report(y_true, y_pred))
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/confusion_matrix-%E6%B7%B7%E6%B7%86%E7%9F%A9%E9%98%B5/">confusion_matrix (混淆矩阵)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/roc_auc-roc%E6%9B%B2%E7%BA%BF%E4%B8%8B%E9%9D%A2%E7%A7%AF/">roc_auc (ROC曲线下面积)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/precision_recall-%E7%B2%BE%E7%A1%AE%E7%8E%87%E4%B8%8E%E5%8F%AC%E5%9B%9E%E7%8E%87/">precision_recall (精确率与召回率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/cross_validation-%E4%BA%A4%E5%8F%89%E9%AA%8C%E8%AF%81/">cross_validation (交叉验证)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>等几率公平</title><link>https://terms-en.ai-term-hub.com/zh/terms/equalized_odds/</link><pubDate>Sat, 18 Jul 2026 11:16:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/equalized_odds/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>等几率公平是算法公平性中使用的一种统计parity约束，旨在确保模型对所有受保护群体具有同等良好的表现。具体而言，它要求模型在不同群体中的真正例率（TPR）和假正例率（FPR）保持一致。这意味着无论个体的敏感属性如何，模型做出正确预测的概率以及错误预测的概率应当是均衡的，从而减少系统性偏见。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种公平性指标，要求不同人口统计群体的真正例率和假正例率相等。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>算法公平性&lt;/li>
&lt;li>真正例率&lt;/li>
&lt;li>假正例率&lt;/li>
&lt;li>偏见缓解&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>公平借贷算法&lt;/li>
&lt;li>招聘推荐系统&lt;/li>
&lt;li>刑事司法风险评估&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/demographic-parity-%E4%BA%BA%E5%8F%A3%E7%BB%9F%E8%AE%A1%E5%9D%87%E7%AD%89/">Demographic Parity (人口统计均等)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/individual-fairness-%E4%B8%AA%E4%BD%93%E5%85%AC%E5%B9%B3/">Individual Fairness (个体公平)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bias-audit-%E5%81%8F%E8%A7%81%E5%AE%A1%E8%AE%A1/">Bias Audit (偏见审计)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/fairness-constraints-%E5%85%AC%E5%B9%B3%E6%80%A7%E7%BA%A6%E6%9D%9F/">Fairness Constraints (公平性约束)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>混淆矩阵</title><link>https://terms-en.ai-term-hub.com/zh/terms/confusion_matrix/</link><pubDate>Sat, 18 Jul 2026 11:11:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/confusion_matrix/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>混淆矩阵是一种特定的表格布局，用于可视化算法（通常是监督学习算法）的性能。它显示了真阳性、真阴性、假阳性和假阴性的计数。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>用于描述分类模型在测试数据集上性能的表格。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>真阳性&lt;/li>
&lt;li>假阴性&lt;/li>
&lt;li>精确率&lt;/li>
&lt;li>召回率&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估二分类器&lt;/li>
&lt;li>分析多分类性能&lt;/li>
&lt;li>调试不平衡数据集中的模型偏差&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> sklearn.metrics &lt;span style="color:#f92672">import&lt;/span> confusion_matrix
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>y_true &lt;span style="color:#f92672">=&lt;/span> [&lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#ae81ff">0&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#ae81ff">0&lt;/span>, &lt;span style="color:#ae81ff">1&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>y_pred &lt;span style="color:#f92672">=&lt;/span> [&lt;span style="color:#ae81ff">0&lt;/span>, &lt;span style="color:#ae81ff">0&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>, &lt;span style="color:#ae81ff">0&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(confusion_matrix(y_true, y_pred))
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/precision-%E7%B2%BE%E7%A1%AE%E7%8E%87/">precision (精确率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/recall-%E5%8F%AC%E5%9B%9E%E7%8E%87/">recall (召回率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/f1-score-f1%E5%88%86%E6%95%B0/">F1 score (F1分数)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/roc-curve-roc%E6%9B%B2%E7%BA%BF/">ROC curve (ROC曲线)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>类别效用</title><link>https://terms-en.ai-term-hub.com/zh/terms/category_utility/</link><pubDate>Sat, 18 Jul 2026 11:09:42 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/category_utility/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>该指标量化了一组类别在多大程度上允许人们预测这些类别内属性的值。它在类别大小与其内容同质性之间取得平衡，是概念学习和聚类评估中的重要指标。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>类别效用是一种数学度量，用于根据分类方案提供的信息增益来评估分类方案的有效性。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>信息增益&lt;/li>
&lt;li>聚类评估&lt;/li>
&lt;li>预测准确性&lt;/li>
&lt;li>概念学习&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估聚类质量&lt;/li>
&lt;li>概念获取研究&lt;/li>
&lt;li>优化特征选择&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/entropy-%E7%86%B5/">entropy (熵)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/mutual-information-%E4%BA%92%E4%BF%A1%E6%81%AF/">mutual information (互信息)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/clustering-%E8%81%9A%E7%B1%BB/">clustering (聚类)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/concept-formation-%E6%A6%82%E5%BF%B5%E5%BD%A2%E6%88%90/">concept formation (概念形成)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>贝叶斯遗憾</title><link>https://terms-en.ai-term-hub.com/zh/terms/bayesian_regret/</link><pubDate>Sat, 18 Jul 2026 11:09:03 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/bayesian_regret/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>贝叶斯遗憾量化了在拥有完美信息时可实现的最佳奖励与智能体在不确定性下行动所获得的预期奖励之间的差异。它是通过对所有可能的世界状态进行积分计算得出的，反映了决策者在信息不完全情况下的性能损失。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>决策理论中衡量因对世界真实状态不确定而导致的预期损失的指标。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>期望效用&lt;/li>
&lt;li>先验分布&lt;/li>
&lt;li>决策理论&lt;/li>
&lt;li>遗憾最小化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>强化学习评估&lt;/li>
&lt;li>多臂老虎机问题&lt;/li>
&lt;li>策略博弈分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/frequentist-regret-%E9%A2%91%E7%8E%87%E5%AD%A6%E6%B4%BE%E9%81%97%E6%86%BE/">Frequentist regret (频率学派遗憾)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/pareto-optimality-%E5%B8%95%E7%B4%AF%E6%89%98%E6%9C%80%E4%BC%98/">Pareto optimality (帕累托最优)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/expected-value-%E6%9C%9F%E6%9C%9B%E5%80%BC/">Expected value (期望值)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/information-gain-%E4%BF%A1%E6%81%AF%E5%A2%9E%E7%9B%8A/">Information gain (信息增益)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>ASR-complete</title><link>https://terms-en.ai-term-hub.com/zh/terms/asr_complete/</link><pubDate>Sat, 18 Jul 2026 11:04:01 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/asr_complete/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>术语“ASR-complete”表示自动语音识别系统在特定且定义明确的任务和数据集上，其性能已达到与人类转录员相当的水平。这是一个重要的里程碑，标志着系统在特定领域内的识别精度已满足实际应用的高标准要求。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>ASR-complete 描述在标准化基准数据集上达到人类水平准确率的语音识别系统。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>语音识别&lt;/li>
&lt;li>人类水平准确率&lt;/li>
&lt;li>错误率&lt;/li>
&lt;li>基准测试&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估 ASR 模型性能&lt;/li>
&lt;li>制定行业标准&lt;/li>
&lt;li>比较不同的声学模型&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/automatic-speech-recognition-%E8%87%AA%E5%8A%A8%E8%AF%AD%E9%9F%B3%E8%AF%86%E5%88%AB/">Automatic Speech Recognition (自动语音识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wer-%E8%AF%8D%E9%94%99%E8%AF%AF%E7%8E%87/">WER (词错误率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>延迟</title><link>https://terms-en.ai-term-hub.com/zh/terms/latency/</link><pubDate>Sat, 18 Jul 2026 11:00:35 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/latency/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>延迟衡量AI服务的响应速度，通常以毫秒为单位表示。它包括推理时间、网络传输延迟和处理开销。低延迟对于实时应用至关重要。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>AI系统中请求发起与响应开始之间的时间延迟。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>推理时间&lt;/li>
&lt;li>响应时间&lt;/li>
&lt;li>实时处理&lt;/li>
&lt;li>优化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>语音识别系统&lt;/li>
&lt;li>自动驾驶车辆控制&lt;/li>
&lt;li>实时翻译服务&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/throughput-%E5%90%9E%E5%90%90%E9%87%8F/">throughput (吞吐量)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/inference-%E6%8E%A8%E7%90%86/">inference (推理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/optimization-%E4%BC%98%E5%8C%96/">optimization (优化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/real_time-%E5%AE%9E%E6%97%B6/">real_time (实时)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Wasserstein（瓦瑟斯坦）</title><link>https://terms-en.ai-term-hub.com/zh/terms/wasserstein/</link><pubDate>Sat, 18 Jul 2026 10:56:05 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/wasserstein/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Wasserstein距离，也称为地球移动距离（Earth Mover&amp;rsquo;s Distance），通过计算将质量从一个分布移动到另一个分布所需的“最小工作量”来量化两个概率分布之间的差异。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种基于将一个概率分布转换为另一个所需最小成本来衡量概率分布之间距离的指标。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>地球移动距离&lt;/li>
&lt;li>概率分布&lt;/li>
&lt;li>最优传输&lt;/li>
&lt;li>梯度稳定性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练稳定的生成对抗网络&lt;/li>
&lt;li>领域自适应&lt;/li>
&lt;li>测量分布相似性&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/kl%E6%95%A3%E5%BA%A6-kullback-leibler-divergence/">KL散度 (Kullback-Leibler Divergence)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%94%9F%E6%88%90%E5%AF%B9%E6%8A%97%E7%BD%91%E7%BB%9C-generative-adversarial-networks/">生成对抗网络 (Generative Adversarial Networks)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%9C%80%E4%BC%98%E4%BC%A0%E8%BE%93-optimal-transport/">最优传输 (Optimal Transport)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%88%86%E5%B8%83%E5%8C%B9%E9%85%8D-distribution-matching/">分布匹配 (Distribution Matching)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>得分</title><link>https://terms-en.ai-term-hub.com/zh/terms/score/</link><pubDate>Sat, 18 Jul 2026 10:54:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/score/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>得分量化了机器学习模型针对特定指标（如准确率、精确度或奖励）的表现。在强化学习中，得分表示累积奖励，而在分类任务中，得&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>得分是表示模型预测或解决方案质量、置信度或适应度的数值。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>指标评估&lt;/li>
&lt;li>置信水平&lt;/li>
&lt;li>奖励信号&lt;/li>
&lt;li>优化目标&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估分类准确率&lt;/li>
&lt;li>跟踪强化学习智能体的进展&lt;/li>
&lt;li>对搜索结果进行排名&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/metric-%E6%8C%87%E6%A0%87/">Metric (指标)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/accuracy-%E5%87%86%E7%A1%AE%E7%8E%87/">Accuracy (准确率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/reward-%E5%A5%96%E5%8A%B1/">Reward (奖励)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/evaluation-%E8%AF%84%E4%BC%B0/">Evaluation (评估)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>总体</title><link>https://terms-en.ai-term-hub.com/zh/terms/overall/</link><pubDate>Sat, 18 Jul 2026 10:53:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/overall/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在评估 AI 模型时，“总体”指标提供了系统性能的全面视图，而不是仅关注孤立的部分。这包括总体准确率、平均精度均值（mAP）或总体的业务影响评估。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>总体指 AI 系统在所有测试用例或操作场景下的聚合性能、准确性或影响。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>聚合指标&lt;/li>
&lt;li>整体评估&lt;/li>
&lt;li>性能总结&lt;/li>
&lt;li>系统性影响&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>报告模型基准测试结果&lt;/li>
&lt;li>评估 AI 部署的业务投资回报率（ROI）&lt;/li>
&lt;li>比较不同的算法架构&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/accuracy-%E5%87%86%E7%A1%AE%E7%8E%87/">Accuracy (准确率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/benchmark-%E5%9F%BA%E5%87%86/">Benchmark (基准)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/kpi-%E5%85%B3%E9%94%AE%E7%BB%A9%E6%95%88%E6%8C%87%E6%A0%87/">KPI (关键绩效指标)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/evaluation-%E8%AF%84%E4%BC%B0/">Evaluation (评估)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>评估</title><link>https://terms-en.ai-term-hub.com/zh/terms/evaluation/</link><pubDate>Sat, 18 Jul 2026 10:50:55 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/evaluation/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>评估涉及使用定量指标（如准确率、F1分数、BLEU分数）和定性分析，系统地测量AI模型在特定任务上的表现。它包括验证集测试、交叉验证以及对模型泛化能力、公平性和偏差的审计。有效的评估流程对于确保模型在实际部署中的可靠性、安全性和合规性至关重要。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>评估是根据预定义的指标和数据集，对AI模型的性能、准确性和鲁棒性进行评估的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>指标&lt;/li>
&lt;li>验证集&lt;/li>
&lt;li>泛化&lt;/li>
&lt;li>基准测试&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>在超参数调优期间比较模型版本&lt;/li>
&lt;li>审计模型的公平性和偏见&lt;/li>
&lt;li>认证AI系统以符合监管要求&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/accuracy-%E5%87%86%E7%A1%AE%E7%8E%87/">Accuracy (准确率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/overfitting-%E8%BF%87%E6%8B%9F%E5%90%88/">Overfitting (过拟合)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/benchmark-%E5%9F%BA%E5%87%86/">Benchmark (基准)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/loss-function-%E6%8D%9F%E5%A4%B1%E5%87%BD%E6%95%B0/">Loss Function (损失函数)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>基准</title><link>https://terms-en.ai-term-hub.com/zh/terms/benchmark/</link><pubDate>Sat, 18 Jul 2026 10:49:30 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/benchmark/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能中，基准是一套标准化的测试套件或数据集，旨在衡量机器学习模型的能力。它提供了一个一致的框架，用于比较不同模型的性能表现。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>用于评估AI模型性能相对于既定基线的标准参考点或指标。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>标准化评估&lt;/li>
&lt;li>性能指标&lt;/li>
&lt;li>对比分析&lt;/li>
&lt;li>基线比较&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>将新模型架构与现有的最先进方法进行对比评估。&lt;/li>
&lt;li>在不同条件或数据集下评估模型的鲁棒性。&lt;/li>
&lt;li>在学术研究中报告结果以确保公平的比较。&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/baseline-%E5%9F%BA%E7%BA%BF/">baseline (基线)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/evaluation_metrics-%E8%AF%84%E4%BC%B0%E6%8C%87%E6%A0%87/">evaluation_metrics (评估指标)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dataset-%E6%95%B0%E6%8D%AE%E9%9B%86/">dataset (数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/state_of_the_art-%E6%9C%80%E5%85%88%E8%BF%9B%E6%8A%80%E6%9C%AF/">state_of_the_art (最先进技术)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>基准</title><link>https://terms-en.ai-term-hub.com/zh/terms/bench/</link><pubDate>Sat, 18 Jul 2026 10:49:18 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/bench/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>基准测试为比较不同AI模型或算法的能力提供了标准化的参考点。它通常涉及精心策划的数据集以及特定的评估指标，如准确率、召回率等。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Benchmark（基准）的缩写，用于评估AI模型性能的标准测试集或指标。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>评估指标&lt;/li>
&lt;li>标准化测试&lt;/li>
&lt;li>性能比较&lt;/li>
&lt;li>数据集策展&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>NLP领域的GLUE基准&lt;/li>
&lt;li>计算机视觉领域的ImageNet&lt;/li>
&lt;li>生产环境中的模型选择&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/metric-%E6%8C%87%E6%A0%87/">Metric (指标)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/validation-%E9%AA%8C%E8%AF%81/">Validation (验证)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dataset-%E6%95%B0%E6%8D%AE%E9%9B%86/">Dataset (数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/evaluation-%E8%AF%84%E4%BC%B0/">Evaluation (评估)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>