<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Retrieval on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/retrieval/</link><description>Recent content in Retrieval on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/retrieval/index.xml" rel="self" type="application/rss+xml"/><item><title>混合搜索</title><link>https://terms-en.ai-term-hub.com/zh/terms/hybrid_search/</link><pubDate>Sat, 18 Jul 2026 11:21:34 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/hybrid_search/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>混合搜索整合了两种不同的检索方法：捕捉语义含义和上下文的稠密向量搜索，以及匹配确切术语的稀疏向量（关键词）搜索。通过利用这两种方法的互补优势，混合搜索能够显著提升检索结果的质量。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种检索策略，将语义向量搜索与传统基于关键词的索引相结合，以提高准确性和相关性。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>向量搜索&lt;/li>
&lt;li>关键词匹配&lt;/li>
&lt;li>重排序&lt;/li>
&lt;li>倒数排名融合&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>企业文档检索&lt;/li>
&lt;li>电子商务产品搜索&lt;/li>
&lt;li>高级检索增强生成（RAG）管道&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic_search-%E8%AF%AD%E4%B9%89%E6%90%9C%E7%B4%A2/">semantic_search (语义搜索)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sparse_vectors-%E7%A8%80%E7%96%8F%E5%90%91%E9%87%8F/">sparse_vectors (稀疏向量)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dense_vectors-%E7%A8%A0%E5%AF%86%E5%90%91%E9%87%8F/">dense_vectors (稠密向量)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vector_database-%E5%90%91%E9%87%8F%E6%95%B0%E6%8D%AE%E5%BA%93/">vector_database (向量数据库)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集:Ms Marco</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetms_marco/</link><pubDate>Sat, 18 Jul 2026 11:13:13 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetms_marco/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>MS MARCO（Microsoft Machine Reading Comprehension）是自然语言处理中广泛使用的数据集，特别适用于信息检索和问答任务。它由匿名化的搜索查询组成。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>微软机器阅读理解数据集，是一个大规模的真实搜索查询和相关文档片段集合，用于训练信息检索系统。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>信息检索&lt;/li>
&lt;li>段落排序&lt;/li>
&lt;li>搜索查询&lt;/li>
&lt;li>机器阅读理解&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练搜索引擎&lt;/li>
&lt;li>开发问答系统&lt;/li>
&lt;li>基准测试检索模型&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AF%86%E9%9B%86%E6%A3%80%E7%B4%A2-dense-retrieval/">密集检索 (Dense Retrieval)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert/">Bert&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%90%9C%E7%B4%A2%E5%BC%95%E6%93%8E%E4%BC%98%E5%8C%96-search-engine-optimization/">搜索引擎优化 (Search Engine Optimization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E5%9F%BA%E5%87%86%E6%B5%8B%E8%AF%95-nlp-benchmarks/">NLP 基准测试 (NLP Benchmarks)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集:Natural Questions</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetnatural_questions/</link><pubDate>Sat, 18 Jul 2026 11:13:13 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetnatural_questions/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Natural Questions (NQ) 是由 Google 推出的基准数据集，旨在推动开放域问答研究的发展。它将 Google 的真实匿名搜索查询映射到长篇幅的答案上。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个用于开放域问答的大规模数据集，包含来自 Google 搜索的真实用户查询，并配以标注好的维基百科段落作为答案。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>开放域问答&lt;/li>
&lt;li>维基百科&lt;/li>
&lt;li>Google 搜索&lt;/li>
&lt;li>长篇幅答案&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>构建基于知识的问答系统&lt;/li>
&lt;li>检索增强生成（RAG）的训练&lt;/li>
&lt;li>评估事实准确性&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/squad/">SQuAD&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dpr/">DPR&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%9F%A5%E8%AF%86%E6%A3%80%E7%B4%A2-knowledge-retrieval/">知识检索 (Knowledge Retrieval)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E4%BA%8B%E5%AE%9E%E9%AA%8C%E8%AF%81-fact-verification/">事实验证 (Fact Verification)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：Gooaq</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetgooaq/</link><pubDate>Sat, 18 Jul 2026 11:13:01 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetgooaq/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>GooAQ 是从 Google Answers 服务编译而成的数据集，拥有海量用户提交的问题以及详细的付费回答。它是训练&amp;hellip;的宝贵资源&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个包含 Google Answers 查询和响应的大规模数据集，用于训练信息检索和问答模型。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>开放域问答&lt;/li>
&lt;li>信息检索&lt;/li>
&lt;li>Google Answers&lt;/li>
&lt;li>大规模数据集&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>构建问答系统&lt;/li>
&lt;li>检索增强生成 (RAG)&lt;/li>
&lt;li>语义搜索评估&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/trec-qa-trec%E9%97%AE%E7%AD%94%E4%BB%BB%E5%8A%A1/">TREC QA (TREC问答任务)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-questions-%E8%87%AA%E7%84%B6%E9%97%AE%E9%A2%98%E6%95%B0%E6%8D%AE%E9%9B%86/">Natural Questions (自然问题数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/retrieval-models-%E6%A3%80%E7%B4%A2%E6%A8%A1%E5%9E%8B/">Retrieval Models (检索模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：嵌入数据/PAQ 对</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_datapaq_pairs/</link><pubDate>Sat, 18 Jul 2026 11:12:55 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_datapaq_pairs/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>PAQ（伪答案质量）数据集包含从维基百科中提取的数百万个自动生成的问答对。它是专门为提供训练稠密检索器所需的数据而设计的。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个源自维基百科的大规模问答对数据集，旨在用于稠密段落检索的训练。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>稠密段落检索&lt;/li>
&lt;li>负采样&lt;/li>
&lt;li>开放域问答&lt;/li>
&lt;li>维基百科提取&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练稠密检索器&lt;/li>
&lt;li>开放域问答&lt;/li>
&lt;li>评估检索性能&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dpr-dense-passage-retrieval/">DPR (Dense Passage Retrieval)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%8F%8C%E7%BC%96%E7%A0%81%E5%99%A8-bi-encoder/">双编码器 (Bi-Encoder)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%9F%A5%E8%AF%86%E5%BA%93-knowledge-base/">知识库 (Knowledge Base)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>问答</title><link>https://terms-en.ai-term-hub.com/zh/terms/question_answering/</link><pubDate>Sat, 18 Jul 2026 11:01:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/question_answering/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>问答（QA）涉及从给定上下文或知识库中检索或生成对用户查询的准确响应。它包括依赖特定文档的封闭领域问答，以及基于通用知识的开放领域问答。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一项自然语言处理任务，系统自动提供用自然语言提出的问题的精确答案。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>信息检索&lt;/li>
&lt;li>语义理解&lt;/li>
&lt;li>上下文提取&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>虚拟助手&lt;/li>
&lt;li>搜索引擎&lt;/li>
&lt;li>客户支持自动化&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%A3%80%E7%B4%A2%E5%A2%9E%E5%BC%BA%E7%94%9F%E6%88%90-retrieval-augmented-generation/">检索增强生成 (Retrieval-Augmented Generation)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%91%BD%E5%90%8D%E5%AE%9E%E4%BD%93%E8%AF%86%E5%88%AB-named-entity-recognition/">命名实体识别 (Named Entity Recognition)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%96%87%E6%9C%AC%E6%91%98%E8%A6%81-text-summarization/">文本摘要 (Text Summarization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%84%8F%E5%9B%BE%E5%88%86%E7%B1%BB-intent-classification/">意图分类 (Intent Classification)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>匹配</title><link>https://terms-en.ai-term-hub.com/zh/terms/matching/</link><pubDate>Sat, 18 Jul 2026 10:53:03 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/matching/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>匹配是机器学习中用于建立不同数据实体之间关系的关键技术。在计算机视觉中，特征匹配用于识别图像间的对应点；在推荐系统中，则用于匹配用户偏好与物品特征。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>匹配涉及将两组数据点或特征对齐，以识别它们之间的对应关系、相似性或最优配对。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>特征对应&lt;/li>
&lt;li>相似度度量&lt;/li>
&lt;li>嵌入空间&lt;/li>
&lt;li>最近邻搜索&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>视频中的目标跟踪&lt;/li>
&lt;li>推荐引擎&lt;/li>
&lt;li>数据库中的重复项检测&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/cosine-similarity-%E4%BD%99%E5%BC%A6%E7%9B%B8%E4%BC%BC%E5%BA%A6/">Cosine Similarity (余弦相似度)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/retrieval-%E6%A3%80%E7%B4%A2/">Retrieval (检索)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/clustering-%E8%81%9A%E7%B1%BB/">Clustering (聚类)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5/">Embedding (嵌入)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>