<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Benchmark on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/benchmark/</link><description>Recent content in Benchmark on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/benchmark/index.xml" rel="self" type="application/rss+xml"/><item><title>登山车问题</title><link>https://terms-en.ai-term-hub.com/zh/terms/mountain_car_problem/</link><pubDate>Sat, 18 Jul 2026 11:26:37 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/mountain_car_problem/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>登山车问题是强化学习研究中的标准基准。目标是将一辆动力不足的车控制到陡坡顶部。由于车辆无法直接爬上山坡，智能体需要利用惯性来回摆动以积累动能。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个经典强化学习任务，智能体必须仅使用加速控制将车开上陡峭的山坡。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>稀疏奖励&lt;/li>
&lt;li>延迟后果&lt;/li>
&lt;li>连续控制&lt;/li>
&lt;li>基准测试&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>测试新的强化学习算法&lt;/li>
&lt;li>展示价值函数近似技术&lt;/li>
&lt;li>强化学习概念的教学示例&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/openai-gym-openai-%E7%8E%AF%E5%A2%83%E5%BA%93/">OpenAI Gym (OpenAI 环境库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0-reinforcement-learning/">强化学习 (Reinforcement Learning)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%8D%95%E6%91%86%E4%B8%8A%E6%91%86-pendulum-swing-up/">单摆上摆 (Pendulum Swing-Up)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%80%BC%E8%BF%AD%E4%BB%A3-value-iteration/">值迭代 (Value Iteration)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>FrontierMath</title><link>https://terms-en.ai-term-hub.com/zh/terms/frontiermath/</link><pubDate>Sat, 18 Jul 2026 11:17:53 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/frontiermath/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>FrontierMath是一个专门的评估套件，用于测试大型语言模型在复杂数学问题解决方面的极限。与标准的算术基准不同，它侧重于高水平（high-scoring）的数学推理能力评估。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个旨在评估最先进AI模型高级数学推理能力的基准数据集。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>数学推理&lt;/li>
&lt;li>基准评估&lt;/li>
&lt;li>思维链&lt;/li>
&lt;li>最先进水平&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估大型语言模型在复杂数学问题上的表现&lt;/li>
&lt;li>研究模型推理能力的改进&lt;/li>
&lt;li>比较不同模型架构的定量技能&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/math-benchmark-math%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95/">MATH Benchmark (MATH基准测试)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/chain-of-thought-%E6%80%9D%E7%BB%B4%E9%93%BE/">Chain-of-Thought (思维链)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/reasoning-models-%E6%8E%A8%E7%90%86%E6%A8%A1%E5%9E%8B/">Reasoning Models (推理模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ai-evaluation-ai%E8%AF%84%E4%BC%B0/">AI Evaluation (AI评估)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：Snli</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetsnli/</link><pubDate>Sat, 18 Jul 2026 11:13:26 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetsnli/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>SNLI 是一个基准数据集，包含超过 50 万个标注的句子对，分为三类：蕴含（entailment）、矛盾（contradiction）和中性（neutral）。它旨在推动自然语言推理研究。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>斯坦福自然语言推理语料库，一个包含英语句子对及人工编写的文本蕴含标签的大型数据集。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自然语言推理&lt;/li>
&lt;li>文本蕴含&lt;/li>
&lt;li>句子对&lt;/li>
&lt;li>基准测试&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估语义相似度模型&lt;/li>
&lt;li>训练NLI分类器&lt;/li>
&lt;li>研究NLP中的逻辑推理&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multinli-%E5%A4%9A%E9%A2%86%E5%9F%9F%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E6%8E%A8%E7%90%86%E8%AF%AD%E6%96%99%E5%BA%93/">MultiNLI (多领域自然语言推理语料库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rte-%E6%96%87%E6%9C%AC%E8%AF%86%E5%88%AB%E8%95%B4%E5%90%AB%E4%BB%BB%E5%8A%A1/">RTE (文本识别蕴含任务)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BAtransformer/">BERT (双向编码器表示Transformer)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic-role-labeling-%E8%AF%AD%E4%B9%89%E8%A7%92%E8%89%B2%E6%A0%87%E6%B3%A8/">Semantic Role Labeling (语义角色标注)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集:Ms Marco</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetms_marco/</link><pubDate>Sat, 18 Jul 2026 11:13:13 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetms_marco/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>MS MARCO（Microsoft Machine Reading Comprehension）是自然语言处理中广泛使用的数据集，特别适用于信息检索和问答任务。它由匿名化的搜索查询组成。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>微软机器阅读理解数据集，是一个大规模的真实搜索查询和相关文档片段集合，用于训练信息检索系统。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>信息检索&lt;/li>
&lt;li>段落排序&lt;/li>
&lt;li>搜索查询&lt;/li>
&lt;li>机器阅读理解&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练搜索引擎&lt;/li>
&lt;li>开发问答系统&lt;/li>
&lt;li>基准测试检索模型&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AF%86%E9%9B%86%E6%A3%80%E7%B4%A2-dense-retrieval/">密集检索 (Dense Retrieval)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert/">Bert&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%90%9C%E7%B4%A2%E5%BC%95%E6%93%8E%E4%BC%98%E5%8C%96-search-engine-optimization/">搜索引擎优化 (Search Engine Optimization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95-nlp-benchmarks/">NLP 基准测试 (NLP Benchmarks)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集:Multi Nli</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetmulti_nli/</link><pubDate>Sat, 18 Jul 2026 11:13:13 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetmulti_nli/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>MultiNLI 是通过 GLUE 基准测试提供的众包语料库，旨在评估口语和书面语各种体裁下的自然语言推理（NLI）能力。它提供了前提-假设对。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>多体裁自然语言推理语料库，是一个包含数百万句人工撰写的英语句子的大型数据集，拥有用于文本蕴含关系的黄金人工标注。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自然语言推理&lt;/li>
&lt;li>文本蕴含&lt;/li>
&lt;li>GLUE 基准测试&lt;/li>
&lt;li>语义相似度&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估语义理解能力&lt;/li>
&lt;li>训练 NLI 模型&lt;/li>
&lt;li>跨体裁泛化测试&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/snli/">SNLI&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/glue/">GLUE&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%95%B4%E5%90%AB-entailment/">蕴含 (Entailment)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%9F%9B%E7%9B%BE-contradiction/">矛盾 (Contradiction)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>