<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>NLP on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/nlp/</link><description>Recent content in NLP on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/nlp/index.xml" rel="self" type="application/rss+xml"/><item><title>XLM-RoBERTa</title><link>https://terms-en.ai-term-hub.com/zh/terms/xlm_roberta/</link><pubDate>Sat, 18 Jul 2026 11:38:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/xlm_roberta/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>XLM-RoBERTa（跨语言语言模型 RoBERTa）是由 Meta AI 开发的大规模多语言模型。它通过在一个涵盖 100 多种语言的多样化数据集上进行预训练，扩展了 RoBERTa 的架构，从而实现了强大的跨语言理解和生成能力，&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种基于 RoBERTa 的多语言 Transformer 模型，在超过 100 种语言的庞大文本数据集上进行预训练。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>多语言自然语言处理&lt;/li>
&lt;li>Transformer 架构&lt;/li>
&lt;li>跨语言迁移&lt;/li>
&lt;li>预训练&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>多语言情感分析&lt;/li>
&lt;li>跨语言问答&lt;/li>
&lt;li>机器翻译&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BAtransformer/">BERT (双向编码器表示Transformer)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/roberta-%E4%BC%98%E5%8C%96%E7%89%88%E7%9A%84bert/">RoBERTa (优化版的BERT)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/m-bert-%E5%A4%9A%E8%AF%AD%E8%A8%80bert/">M-BERT (多语言BERT)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentence-transformers-%E5%8F%A5%E5%AD%90transformer/">Sentence Transformers (句子Transformer)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>零样本提示</title><link>https://terms-en.ai-term-hub.com/zh/terms/zero_shot_prompting/</link><pubDate>Sat, 18 Jul 2026 11:38:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/zero_shot_prompting/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>零样本提示涉及直接通过文本提示要求预训练语言模型完成任务，而不提供任何少样本示例或进行额外的微调。该模型利用其在大规模数据上学习到的通用知识和指令遵循能力来生成响应，&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种技术，大型语言模型无需先前示例或微调，仅依靠自然语言指令即可执行任务。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>提示工程&lt;/li>
&lt;li>涌现能力&lt;/li>
&lt;li>上下文学习&lt;/li>
&lt;li>指令微调&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>AI 应用的快速原型设计&lt;/li>
&lt;li>动态任务切换&lt;/li>
&lt;li>降低数据标注成本&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/few-shot-prompting-%E5%B0%91%E6%A0%B7%E6%9C%AC%E6%8F%90%E7%A4%BA/">Few-Shot Prompting (少样本提示)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/chain-of-thought-%E6%80%9D%E7%BB%B4%E9%93%BE/">Chain-of-Thought (思维链)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/fine-tuning-%E5%BE%AE%E8%B0%83/">Fine-Tuning (微调)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E5%9E%8B%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大型语言模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>WordPiece</title><link>https://terms-en.ai-term-hub.com/zh/terms/wordpiece/</link><pubDate>Sat, 18 Jul 2026 11:38:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/wordpiece/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>WordPiece 是一种广泛应用于 BERT 和 ALBERT 等自然语言处理模型的分词方法。它将单词分解为更小的子词单元，以应对形态学丰富性并减少词汇表大小，从而更好地处理未见过的单词。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种子词分词算法，通过递归合并最频繁出现的字符对来处理未登录词（OOV）。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>子词分词&lt;/li>
&lt;li>词汇扩展&lt;/li>
&lt;li>未登录词处理&lt;/li>
&lt;li>形态分析&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为 BERT 模型预处理文本&lt;/li>
&lt;li>处理低资源语言&lt;/li>
&lt;li>减小嵌入矩阵大小&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> BertTokenizer
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokenizer &lt;span style="color:#f92672">=&lt;/span> BertTokenizer&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;bert-base-uncased&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokens &lt;span style="color:#f92672">=&lt;/span> tokenizer&lt;span style="color:#f92672">.&lt;/span>tokenize(&lt;span style="color:#e6db74">&amp;#39;unhappiness&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(tokens)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AD%97%E8%8A%82%E5%AF%B9%E7%BC%96%E7%A0%81-byte-pair-encoding/">字节对编码 (Byte-Pair Encoding)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentencepiece/">SentencePiece&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%88%86%E8%AF%8D-tokenization/">分词 (Tokenization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E9%A2%84%E5%A4%84%E7%90%86-nlp-preprocessing/">NLP 预处理 (NLP preprocessing)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>毒性</title><link>https://terms-en.ai-term-hub.com/zh/terms/toxicity/</link><pubDate>Sat, 18 Jul 2026 11:37:02 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/toxicity/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>AI中的毒性是指生成或传播不尊重他人、可能导致用户离开讨论或针对特定身份群体的内容。它涵盖了一个从轻微的不当言论到严重的仇恨言论和暴力威胁的光谱。在自然语言处理中，毒性通常被视为一种需要被识别和过滤的安全对齐问题，以确保AI系统不会成为网络欺凌或极端主义思想的传播渠道。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>AI生成文本中存在有害、冒犯性或虐待性内容的现象，包括仇恨言论、骚扰和威胁。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>仇恨言论&lt;/li>
&lt;li>骚扰&lt;/li>
&lt;li>内容审核&lt;/li>
&lt;li>伦理AI&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>社交媒体平台安全&lt;/li>
&lt;li>客服机器人过滤&lt;/li>
&lt;li>社区准则执行&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%81%8F%E8%A7%81-bias/">偏见 (Bias)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AE%89%E5%85%A8%E5%AF%B9%E9%BD%90-safety-alignment/">安全对齐 (Safety Alignment)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%86%85%E5%AE%B9%E8%BF%87%E6%BB%A4-content-filtering/">内容过滤 (Content Filtering)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86-natural-language-processing/">自然语言处理 (Natural Language Processing)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>毒性检测</title><link>https://terms-en.ai-term-hub.com/zh/terms/toxicity_detection/</link><pubDate>Sat, 18 Jul 2026 11:37:02 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/toxicity_detection/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>毒性检测采用自然语言处理技术分析文本输入，并分配一个概率分数以指示有害内容的可能性。这些系统通常使用监督学习训练，基于标注好的数据集（如包含仇恨言论或骚扰的评论）来学习识别有毒模式的特征。它们被广泛应用于实时内容审核，帮助平台自动标记或删除违反社区准则的帖子，从而维护在线环境的安全与健康。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>利用机器学习模型自动识别和分类文本中有害或虐待性语言的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本分类&lt;/li>
&lt;li>机器学习&lt;/li>
&lt;li>审核&lt;/li>
&lt;li>NLP&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自动化内容审核&lt;/li>
&lt;li>实时聊天过滤&lt;/li>
&lt;li>品牌声誉监控&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%83%85%E6%84%9F%E5%88%86%E6%9E%90-sentiment-analysis/">情感分析 (Sentiment Analysis)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%9E%83%E5%9C%BE%E9%82%AE%E4%BB%B6%E6%A3%80%E6%B5%8B-spam-detection/">垃圾邮件检测 (Spam Detection)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%9C%89%E5%AE%B3%E5%86%85%E5%AE%B9-harmful-content/">有害内容 (Harmful Content)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ai%E5%AE%89%E5%85%A8-ai-safety/">AI安全 (AI Safety)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>文本生成</title><link>https://terms-en.ai-term-hub.com/zh/terms/text_generation/</link><pubDate>Sat, 18 Jul 2026 11:36:07 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/text_generation/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>文本生成是自然语言处理中的一种基本应用范式，人工智能模型在此过程中创建新的文本内容。通过预测序列中下一个最可能的标记，模型能够连贯地生成文章、对话或代码等多样化文本形式。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种AI能力，模型根据提供的提示或上下文，逐个标记地生成类似人类文本的序列。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自回归建模&lt;/li>
&lt;li>标记预测&lt;/li>
&lt;li>提示工程&lt;/li>
&lt;li>采样策略&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自动化内容创作&lt;/li>
&lt;li>对话式聊天机器人&lt;/li>
&lt;li>代码补全工具&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E5%9E%8B%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">llm (大型语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompting-%E6%8F%90%E7%A4%BA%E6%8A%80%E6%9C%AF/">prompting (提示技术)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">nlp (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/autoregressive-%E8%87%AA%E5%9B%9E%E5%BD%92/">autoregressive (自回归)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>文本分类</title><link>https://terms-en.ai-term-hub.com/zh/terms/text_classification/</link><pubDate>Sat, 18 Jul 2026 11:35:54 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/text_classification/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>文本分类是一种监督学习任务，算法为无结构的文本数据分配预定义的类别。常用技术包括朴素贝叶斯、支持向量机和深度学习模型。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>根据内容或语义含义将文本归类到不同组织组的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>监督学习&lt;/li>
&lt;li>标注&lt;/li>
&lt;li>特征提取&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>情感分析&lt;/li>
&lt;li>垃圾邮件过滤&lt;/li>
&lt;li>主题建模&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> pipeline
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>classifier &lt;span style="color:#f92672">=&lt;/span> pipeline(&lt;span style="color:#e6db74">&amp;#34;sentiment-analysis&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/named-entity-recognition-%E5%91%BD%E5%90%8D%E5%AE%9E%E4%BD%93%E8%AF%86%E5%88%AB/">Named Entity Recognition (命名实体识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentiment-analysis-%E6%83%85%E6%84%9F%E5%88%86%E6%9E%90/">Sentiment Analysis (情感分析)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-models-transformer%E6%A8%A1%E5%9E%8B/">Transformer Models (Transformer模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>序列标注</title><link>https://terms-en.ai-term-hub.com/zh/terms/sequence_labeling/</link><pubDate>Sat, 18 Jul 2026 11:33:18 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/sequence_labeling/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>序列标注涉及为给定输入序列中的每个标记预测分类标签，例如句子中的单词或字符串中的字符。常见的应用包括词性标注、命名实体识别和句法分块。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种自然语言处理任务，为输入序列中的每个元素分配一个标签。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>标记分类&lt;/li>
&lt;li>上下文依赖&lt;/li>
&lt;li>双向编码&lt;/li>
&lt;li>条件随机场层&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>命名实体识别&lt;/li>
&lt;li>词性标注&lt;/li>
&lt;li>句法分块&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer/">Transformer&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bilstm-%E5%8F%8C%E5%90%91%E9%95%BF%E7%9F%AD%E6%9C%9F%E8%AE%B0%E5%BF%86%E7%BD%91%E7%BB%9C/">BiLSTM (双向长短期记忆网络)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/crf-%E6%9D%A1%E4%BB%B6%E9%9A%8F%E6%9C%BA%E5%9C%BA/">CRF (条件随机场)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>句子变换器</title><link>https://terms-en.ai-term-hub.com/zh/terms/sentence_transformers/</link><pubDate>Sat, 18 Jul 2026 11:33:07 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/sentence_transformers/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>句子变换器是传统变换器模型（如BERT）的扩展，经过微调以产生整个句子的有意义稠密向量表示。与标准的基于标记的模型不同，它们直接输出句子级别的嵌入。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>专门设计用于为任意文本句子生成固定大小向量嵌入的神经网络架构。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>池化层&lt;/li>
&lt;li>对比学习&lt;/li>
&lt;li>稠密嵌入&lt;/li>
&lt;li>变换器架构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>语义搜索引擎&lt;/li>
&lt;li>文本数据聚类&lt;/li>
&lt;li>检索增强生成 (RAG) 管道&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BA/">BERT (双向编码器表示)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embeddings-%E5%B5%8C%E5%85%A5/">Embeddings (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentence-similarity-%E5%8F%A5%E5%AD%90%E7%9B%B8%E4%BC%BC%E5%BA%A6/">Sentence Similarity (句子相似度)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/contrastive-loss-%E5%AF%B9%E6%AF%94%E6%8D%9F%E5%A4%B1/">Contrastive Loss (对比损失)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>句子相似度</title><link>https://terms-en.ai-term-hub.com/zh/terms/sentence_similarity/</link><pubDate>Sat, 18 Jul 2026 11:33:07 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/sentence_similarity/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>句子相似度衡量两个不同句子之间的语义重叠程度。它超越了词汇匹配，旨在理解含义、上下文和意图。这通常通过计算向量之间的距离来实现。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种量化两个句子在语义上相似程度的指标或任务，通常表示为数值分数。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>余弦相似度&lt;/li>
&lt;li>语义等价&lt;/li>
&lt;li>向量距离&lt;/li>
&lt;li>意义表示&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>论坛中的重复问题检测&lt;/li>
&lt;li>同义句识别&lt;/li>
&lt;li>信息检索和文档聚类&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic-search-%E8%AF%AD%E4%B9%89%E6%90%9C%E7%B4%A2/">Semantic search (语义搜索)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embeddings-%E5%B5%8C%E5%85%A5/">Embeddings (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-inference-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E6%8E%A8%E7%90%86/">Natural Language Inference (自然语言推理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/word-embeddings-%E8%AF%8D%E5%B5%8C%E5%85%A5/">Word Embeddings (词嵌入)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>语义折叠</title><link>https://terms-en.ai-term-hub.com/zh/terms/semantic_folding/</link><pubDate>Sat, 18 Jul 2026 11:33:07 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/semantic_folding/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>语义折叠是指将复杂的高维向量嵌入压缩为更易管理的低维表示的过程，且不会造成语义意义的显著丢失。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种将高维语义表示映射到低维空间同时保留关系结构的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>降维&lt;/li>
&lt;li>向量嵌入&lt;/li>
&lt;li>语义保持&lt;/li>
&lt;li>压缩&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>高效的大规模向量数据库索引&lt;/li>
&lt;li>减少自然语言处理模型的内存占用&lt;/li>
&lt;li>优化实时语义搜索系统&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/pca-%E4%B8%BB%E6%88%90%E5%88%86%E5%88%86%E6%9E%90/">PCA (主成分分析)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/autoencoders-%E8%87%AA%E7%BC%96%E7%A0%81%E5%99%A8/">Autoencoders (自编码器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-compression-%E5%B5%8C%E5%85%A5%E5%8E%8B%E7%BC%A9/">Embedding compression (嵌入压缩)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/latent-space-%E6%BD%9C%E5%9C%A8%E7%A9%BA%E9%97%B4/">Latent space (潜在空间)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>语义分析</title><link>https://terms-en.ai-term-hub.com/zh/terms/semantic_analysis/</link><pubDate>Sat, 18 Jul 2026 11:32:54 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/semantic_analysis/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>它超越了句法结构，旨在解释语言输入的实际意图和重要性。这包括根据上下文消除词义歧义、识别实体以及理解&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>语义分析是通过理解自然语言处理中单词之间及其上下文的关系统，从文本中提取意义的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>上下文理解&lt;/li>
&lt;li>词义消歧&lt;/li>
&lt;li>意图识别&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>情感分析&lt;/li>
&lt;li>聊天机器人意图检测&lt;/li>
&lt;li>信息检索系统&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%8F%A5%E6%B3%95%E8%A7%A3%E6%9E%90-syntax-parsing-%E5%88%86%E6%9E%90%E5%8F%A5%E5%AD%90%E8%AF%AD%E6%B3%95%E7%BB%93%E6%9E%84%E7%9A%84%E8%BF%87%E7%A8%8B/">句法解析 (Syntax parsing，分析句子语法结构的过程)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%91%BD%E5%90%8D%E5%AE%9E%E4%BD%93%E8%AF%86%E5%88%AB-named-entity-recognition-%E8%AF%86%E5%88%AB%E6%96%87%E6%9C%AC%E4%B8%AD%E4%BA%BA%E5%90%8D-%E5%9C%B0%E5%90%8D%E7%AD%89%E5%AE%9E%E4%BD%93%E7%9A%84%E6%8A%80%E6%9C%AF/">命名实体识别 (Named entity recognition，识别文本中人名、地名等实体的技术)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%B5%8C%E5%85%A5-embeddings-%E5%B0%86%E7%A6%BB%E6%95%A3%E6%95%B0%E6%8D%AE%E6%98%A0%E5%B0%84%E5%88%B0%E8%BF%9E%E7%BB%AD%E5%90%91%E9%87%8F%E7%A9%BA%E9%97%B4%E7%9A%84%E6%8A%80%E6%9C%AF/">嵌入 (Embeddings，将离散数据映射到连续向量空间的技术)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E7%90%86%E8%A7%A3-natural-language-understanding-%E4%BD%BF%E8%AE%A1%E7%AE%97%E6%9C%BA%E8%83%BD%E5%A4%9F%E7%90%86%E8%A7%A3%E4%BA%BA%E7%B1%BB%E8%AF%AD%E8%A8%80%E7%9A%84%E6%8A%80%E6%9C%AF/">自然语言理解 (Natural language understanding，使计算机能够理解人类语言的技术)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Qwen3.5</title><link>https://terms-en.ai-term-hub.com/zh/terms/qwen35/</link><pubDate>Sat, 18 Jul 2026 11:31:37 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/qwen35/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Qwen3.5 代表由阿里云开发的 Qwen 谱系中的特定发布版本。这一迭代通常建立在先前版本的基础上，通过改进逻辑推理能力、编程熟练度和自然语言理解来增强模型性能，旨在提供更准确和更通用的智能助手体验。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Qwen 大语言模型系列的迭代版本，侧重于增强的推理能力和多语言支持。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>迭代改进&lt;/li>
&lt;li>多语言支持&lt;/li>
&lt;li>推理增强&lt;/li>
&lt;li>阿里云&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>复杂问题解决&lt;/li>
&lt;li>代码生成与调试&lt;/li>
&lt;li>跨语言翻译&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E5%9E%8B%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大型语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/open-source-models-%E5%BC%80%E6%BA%90%E6%A8%A1%E5%9E%8B/">Open Source Models (开源模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>通义千问</title><link>https://terms-en.ai-term-hub.com/zh/terms/qwen/</link><pubDate>Sat, 18 Jul 2026 11:31:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/qwen/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>通义千问代表了由阿里巴巴集团旗下通义实验室创建的一系列先进大语言模型。它涵盖了针对不同任务优化的各个版本，包括自然语言理解、代码生成、数学计算及视觉分析等能力。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>通义千问（Qwen）是由阿里巴巴集团旗下通义实验室自主研发的大语言模型系列。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>大语言模型&lt;/li>
&lt;li>阿里巴巴通义实验室&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;li>基础模型&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>通用对话式AI助手&lt;/li>
&lt;li>内容创作与摘要生成&lt;/li>
&lt;li>复杂逻辑推理任务&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E5%9E%8B%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大型语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tongyi-qianwen-%E9%80%9A%E4%B9%89%E5%8D%83%E9%97%AE/">Tongyi Qianwen (通义千问)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/qwen2-%E9%80%9A%E4%B9%89%E5%8D%83%E9%97%AE%E7%AC%AC%E4%BA%8C%E4%BB%A3/">Qwen2 (通义千问第二代)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/qwen-coder-%E9%80%9A%E4%B9%89%E5%8D%83%E9%97%AE%E4%BB%A3%E7%A0%81%E4%B8%93%E7%94%A8%E7%89%88/">Qwen-Coder (通义千问代码专用版)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Pythia</title><link>https://terms-en.ai-term-hub.com/zh/terms/pythia/</link><pubDate>Sat, 18 Jul 2026 11:31:05 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/pythia/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Pythia 是由 EleutherAI 创建的一系列开源大型语言模型（LLM），旨在促进对神经网络可解释性和行为的研究。该套件包括模&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Pythia 是由 EleutherAI 开发的一系列仅解码器架构的大型语言模型，参数量从 7000 万到 120 亿不等。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>大型语言模型&lt;/li>
&lt;li>可解释性研究&lt;/li>
&lt;li>GPT 架构&lt;/li>
&lt;li>开源人工智能&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>研究 LLM 的缩放行为&lt;/li>
&lt;li>研究模型可解释性&lt;/li>
&lt;li>自然语言处理的教学用途&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/eleutherai/">EleutherAI&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpt-2/">GPT-2&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/the-pile-%E6%95%B0%E6%8D%AE%E9%9B%86-the-pile-dataset/">The Pile 数据集 (The Pile Dataset)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E6%A8%A1%E5%9E%8B-transformer-models/">Transformer 模型 (Transformer Models)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>OpenAI的产品与应用</title><link>https://terms-en.ai-term-hub.com/zh/terms/products_and_applications_of_openai/</link><pubDate>Sat, 18 Jul 2026 11:30:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/products_and_applications_of_openai/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>该术语涵盖了领先的人工智能研究实验室OpenAI创建的商业和研究产品。主要提供包括生成式预训练Transformer（GPT）系列模型用于自然语言处理，DALL-E用于图像生成，以及ChatGPT等对话式AI应用。这些产品广泛应用于内容创作、代码辅助、客户服务等领域。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>指OpenAI开发的一系列AI工具、API和研究成果，包括GPT模型和DALL-E。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>大语言模型&lt;/li>
&lt;li>生成式AI&lt;/li>
&lt;li>API服务&lt;/li>
&lt;li>商业部署&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>通过GitHub Copilot进行自动化代码辅助&lt;/li>
&lt;li>用于营销的创意内容生成&lt;/li>
&lt;li>用于客户服务的对话代理&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpt_models-gpt%E6%A8%A1%E5%9E%8B/">gpt_models (GPT模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dall_e-dall-e%E5%9B%BE%E5%83%8F%E7%94%9F%E6%88%90%E6%A8%A1%E5%9E%8B/">dall_e (DALL-E图像生成模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/chatgpt-chatgpt%E5%AF%B9%E8%AF%9D%E5%8A%A9%E6%89%8B/">chatgpt (ChatGPT对话助手)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/generative_pretrained_transformer-%E7%94%9F%E6%88%90%E5%BC%8F%E9%A2%84%E8%AE%AD%E7%BB%83transformer/">generative_pretrained_transformer (生成式预训练Transformer)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>教学智能体</title><link>https://terms-en.ai-term-hub.com/zh/terms/pedagogical_agent/</link><pubDate>Sat, 18 Jul 2026 11:29:41 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/pedagogical_agent/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>教学智能体是一种软件组件，通常表现为虚拟角色，在教育环境中充当教师或导师。这些智能体利用自然语言处理&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>旨在通过提供指导、反馈和引导来促进学习的人工智能实体。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>智能辅导&lt;/li>
&lt;li>虚拟人&lt;/li>
&lt;li>自适应学习&lt;/li>
&lt;li>人机交互&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>在线课程助手&lt;/li>
&lt;li>语言学习应用&lt;/li>
&lt;li>企业培训模拟&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/intelligent-tutoring-system-%E6%99%BA%E8%83%BD%E8%BE%85%E5%AF%BC%E7%B3%BB%E7%BB%9F/">Intelligent Tutoring System (智能辅导系统)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/chatbot-%E8%81%8A%E5%A4%A9%E6%9C%BA%E5%99%A8%E4%BA%BA/">Chatbot (聊天机器人)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/edtech-%E6%95%99%E8%82%B2%E7%A7%91%E6%8A%80/">EdTech (教育科技)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>文本改写</title><link>https://terms-en.ai-term-hub.com/zh/terms/paraphrasing/</link><pubDate>Sat, 18 Jul 2026 11:29:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/paraphrasing/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在自然语言处理中，文本改写涉及为给定的输入文本生成替代表达，同时保留其原始语义含义。这对于减少抄袭、增强数据多样性以及提高内容可读性至关重要，通常依赖于预训练的语言模型来实现高质量的语义等价转换。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>自然语言处理任务，旨在用不同的词汇或句式重写文本，同时保持原意不变。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>语义保留&lt;/li>
&lt;li>文本重写&lt;/li>
&lt;li>自然语言生成&lt;/li>
&lt;li>同义性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>抄袭检测与规避工具&lt;/li>
&lt;li>用于训练NLP模型的数据增强&lt;/li>
&lt;li>提高技术文档的可读性&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%96%87%E6%9C%AC%E6%91%98%E8%A6%81-text-summarization/">文本摘要 (Text Summarization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%9C%BA%E5%99%A8%E7%BF%BB%E8%AF%91-machine-translation/">机器翻译 (Machine Translation)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86-nlp/">自然语言处理 (NLP)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%BA%8F%E5%88%97%E5%88%B0%E5%BA%8F%E5%88%97%E6%A8%A1%E5%9E%8B-sequence-to-sequence-models/">序列到序列模型 (Sequence-to-Sequence Models)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>P-Tuning</title><link>https://terms-en.ai-term-hub.com/zh/terms/p_tuning/</link><pubDate>Sat, 18 Jul 2026 11:29:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/p_tuning/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>P-Tuning（提示微调）是一种旨在以最低计算成本将大型预训练语言模型适配到特定下游任务的技术。与微调所有模型参数不同，P-Tuning 仅优化少量可学习的提示向量，同时保持预训练模型的权重冻结，从而显著降低资源消耗并防止灾难性遗忘。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>P-Tuning 是一种参数高效微调方法，它通过优化连续的提示嵌入（prompt embeddings）来适应任务，而不是更新整个预训练模型的权重。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>参数高效微调&lt;/li>
&lt;li>虚拟令牌&lt;/li>
&lt;li>冻结权重&lt;/li>
&lt;li>嵌入优化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>少样本学习适配&lt;/li>
&lt;li>资源受限环境&lt;/li>
&lt;li>大语言模型应用的快速原型开发&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/lora-%E4%BD%8E%E7%A7%A9%E9%80%82%E5%BA%94/">LoRA (低秩适应)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/adapter-modules-%E9%80%82%E9%85%8D%E5%99%A8%E6%A8%A1%E5%9D%97/">Adapter Modules (适配器模块)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompt-engineering-%E6%8F%90%E7%A4%BA%E5%B7%A5%E7%A8%8B/">Prompt Engineering (提示工程)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transfer-learning-%E8%BF%81%E7%A7%BB%E5%AD%A6%E4%B9%A0/">Transfer Learning (迁移学习)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>光学字符识别</title><link>https://terms-en.ai-term-hub.com/zh/terms/ocr/</link><pubDate>Sat, 18 Jul 2026 11:28:49 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/ocr/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>光学字符识别（OCR）利用图像处理与模式识别算法来识别数字图像中的文本。它将打印或手写的字符转换为机器编码的数据，从而实现信息的数字化提取。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>OCR 是一项将扫描纸质文档或图像等不同类型的文档转换为可编辑和可搜索数据的技術。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本检测&lt;/li>
&lt;li>字符识别&lt;/li>
&lt;li>图像处理&lt;/li>
&lt;li>数字化&lt;/li>
&lt;li>模式匹配&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>历史档案数字化&lt;/li>
&lt;li>自动化发票处理&lt;/li>
&lt;li>从截图中提取文本&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/computer_vision-%E8%AE%A1%E7%AE%97%E6%9C%BA%E8%A7%86%E8%A7%89/">computer_vision (计算机视觉)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/text_recognition-%E6%96%87%E6%9C%AC%E8%AF%86%E5%88%AB/">text_recognition (文本识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/document_processing-%E6%96%87%E6%A1%A3%E5%A4%84%E7%90%86/">document_processing (文档处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/image_classification-%E5%9B%BE%E5%83%8F%E5%88%86%E7%B1%BB/">image_classification (图像分类)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>母语识别</title><link>https://terms-en.ai-term-hub.com/zh/terms/native_language_identification/</link><pubDate>Sat, 18 Jul 2026 11:28:00 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/native_language_identification/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>母语识别（NLI）是自然语言处理的一个子领域，专注于识别说话者学习的第一语言。与一般的语言检测不同，NLI 分析说话者在发音、语法结构和词汇选择上无意识的母语特征。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>自动从说话者的语音或文本样本中确定其母语的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>说话人画像&lt;/li>
&lt;li>语言特征&lt;/li>
&lt;li>口音识别&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>生物特征安全认证&lt;/li>
&lt;li>个性化客户服务交互&lt;/li>
&lt;li>社会语言学人口统计分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%AD%E8%A8%80%E6%A3%80%E6%B5%8B-language-detection/">语言检测 (Language Detection)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%B4%E8%AF%9D%E4%BA%BA%E5%88%86%E7%A6%BB-speaker-diarization/">说话人分离 (Speaker Diarization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%8F%A3%E9%9F%B3%E8%AF%86%E5%88%AB-accent-identification/">口音识别 (Accent Identification)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%B3%95%E5%8C%BB%E8%AF%AD%E8%A8%80%E5%AD%A6-forensic-linguistics/">法医语言学 (Forensic Linguistics)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>多语言</title><link>https://terms-en.ai-term-hub.com/zh/terms/multilingual/</link><pubDate>Sat, 18 Jul 2026 11:26:49 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/multilingual/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>多语言模型旨在无需为每种语言单独构建模型的情况下处理多样化的语言输入。这些系统通常利用共享的词嵌入或跨语言对齐技术，以实现不同语言间的知识迁移和统一处理。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>AI中的多语言指能够处理、理解或生成多种自然语言内容的模型。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>跨语言迁移&lt;/li>
&lt;li>共享词汇表&lt;/li>
&lt;li>语言识别&lt;/li>
&lt;li>零样本翻译&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>机器翻译服务&lt;/li>
&lt;li>全球客户支持聊天机器人&lt;/li>
&lt;li>跨语言搜索引擎&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/translation-%E7%BF%BB%E8%AF%91/">Translation (翻译)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embeddings-%E5%B5%8C%E5%85%A5/">Embeddings (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大语言模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>掩码生成</title><link>https://terms-en.ai-term-hub.com/zh/terms/mask_generation/</link><pubDate>Sat, 18 Jul 2026 11:25:36 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/mask_generation/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>掩码生成涉及产生空间或时间掩码，以确定在特定操作期间数据集的哪些元素是可见或激活的。在计算机视觉中，它用于对象分割、图像修复等任务；在自然语言处理中，则用于控制注意力机制。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>创建二进制或概率掩码的过程，用于在模型处理期间选择性地隐藏或强调输入数据的某些部分。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>二值掩码&lt;/li>
&lt;li>注意力掩码&lt;/li>
&lt;li>图像修复&lt;/li>
&lt;li>特征选择&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>图像修复与重建&lt;/li>
&lt;li>Transformer注意力机制&lt;/li>
&lt;li>目标检测与分割&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/attention-mechanism-%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6/">Attention mechanism (注意力机制)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic-segmentation-%E8%AF%AD%E4%B9%89%E5%88%86%E5%89%B2/">Semantic segmentation (语义分割)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/inpainting-%E5%9B%BE%E5%83%8F%E4%BF%AE%E5%A4%8D/">Inpainting (图像修复)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/feature-masking-%E7%89%B9%E5%BE%81%E6%8E%A9%E7%A0%81/">Feature masking (特征掩码)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>机器学习与知识提取</title><link>https://terms-en.ai-term-hub.com/zh/terms/machine_learning_and_knowledge_extraction/</link><pubDate>Sat, 18 Jul 2026 11:25:10 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/machine_learning_and_knowledge_extraction/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>该领域将机器学习技术与自然语言处理和数据挖掘相结合，旨在将原始数据转化为可操作的知识。它涉及训练模型以识别实体、关系等关键要素，从而从复杂的数据中提取有价值的洞察。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>利用机器学习算法自动从大规模非结构化数据集中识别模式并推导结构化信息的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>模式识别&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;li>数据挖掘&lt;/li>
&lt;li>特征工程&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>从社交媒体帖子中提取客户情感&lt;/li>
&lt;li>从临床笔记中识别疾病状况&lt;/li>
&lt;li>在法律事务所自动化文档分类&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/data-mining-%E6%95%B0%E6%8D%AE%E6%8C%96%E6%8E%98/">Data Mining (数据挖掘)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/information-retrieval-%E4%BF%A1%E6%81%AF%E6%A3%80%E7%B4%A2/">Information Retrieval (信息检索)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Lyra</title><link>https://terms-en.ai-term-hub.com/zh/terms/lyra/</link><pubDate>Sat, 18 Jul 2026 11:24:58 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/lyra/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在现代AI术语的背景下，Lyra通常指专注于通过自然语言处理增强用户交互的专用AI系统。它可能指代一个开源的大型语言模型开发&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Lyra 指代各种人工智能倡议或模型，最著名的是开源大型语言模型，或旨在增强信息检索的特定AI驱动搜索和发现工具。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>大型语言模型&lt;/li>
&lt;li>信息检索&lt;/li>
&lt;li>开源AI&lt;/li>
&lt;li>语义搜索&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为对话代理提供动力&lt;/li>
&lt;li>提高搜索引擎结果的相关性&lt;/li>
&lt;li>为开发人员提供开源NLP功能&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E5%9E%8B%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大型语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/search-engine-optimization-%E6%90%9C%E7%B4%A2%E5%BC%95%E6%93%8E%E4%BC%98%E5%8C%96/">Search Engine Optimization (搜索引擎优化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/open-source-%E5%BC%80%E6%BA%90/">Open Source (开源)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>MAUVE</title><link>https://terms-en.ai-term-hub.com/zh/terms/mauve/</link><pubDate>Sat, 18 Jul 2026 11:24:58 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/mauve/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>MAUVE 是一种统计度量，旨在评估生成语言模型的输出在多大程度上类似于人类语言使用习惯。与简单的困惑度分数不同，MAUVE 使用虚拟嵌入&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>MAUVE（基于虚拟嵌入的测量对齐）是一种用于自然语言处理的指标，用于评估生成文本分布与人类写作文本分布之间的对齐程度。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本生成评估&lt;/li>
&lt;li>分布匹配&lt;/li>
&lt;li>虚拟嵌入&lt;/li>
&lt;li>语言自然度&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估类GPT模型的输出&lt;/li>
&lt;li>微调语言模型以生成类人文本&lt;/li>
&lt;li>基准测试生成式AI的性能&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/perplexity-%E5%9B%B0%E6%83%91%E5%BA%A6/">Perplexity (困惑度)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bleu-score-bleu%E5%88%86%E6%95%B0/">BLEU Score (BLEU分数)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/language-modeling-%E8%AF%AD%E8%A8%80%E5%BB%BA%E6%A8%A1/">Language Modeling (语言建模)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/generative-ai-%E7%94%9F%E6%88%90%E5%BC%8Fai/">Generative AI (生成式AI)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>长上下文</title><link>https://terms-en.ai-term-hub.com/zh/terms/long_context/</link><pubDate>Sat, 18 Jul 2026 11:24:46 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/long_context/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>长上下文指的是基于 Transformer 的模型处理极长输入长度的能力，通常超过标准的 2k 或 4k 标记限制。这种能力使模型能够分析完整的文档、代码库或长文本，保持全局一致性和细节记忆。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>语言模型处理并保留包含数千或数百万个标记（token）的输入序列信息的能力。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>上下文窗口&lt;/li>
&lt;li>标记限制&lt;/li>
&lt;li>注意力机制&lt;/li>
&lt;li>位置编码&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>总结完整的法律合同&lt;/li>
&lt;li>分析完整的源代码仓库&lt;/li>
&lt;li>处理长篇音频转录文本&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/context-window-%E4%B8%8A%E4%B8%8B%E6%96%87%E7%AA%97%E5%8F%A3/">Context Window (上下文窗口)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-architecture-transformer-%E6%9E%B6%E6%9E%84/">Transformer Architecture (Transformer 架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rope-rotary-positional-embeddings-%E6%97%8B%E8%BD%AC%E4%BD%8D%E7%BD%AE%E7%BC%96%E7%A0%81/">RoPE (Rotary Positional Embeddings, 旋转位置编码)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/kv-cache-%E9%94%AE%E5%80%BC%E7%BC%93%E5%AD%98/">KV Cache (键值缓存)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>语言/行动视角</title><link>https://terms-en.ai-term-hub.com/zh/terms/languageaction_perspective/</link><pubDate>Sat, 18 Jul 2026 11:23:41 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/languageaction_perspective/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>植根于言语行为理论和语用学，这一视角强调话语如何执行请求、承诺或命令等功能。在自然语言处理中，它指导了意图理解和对话生成的设计。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种理论框架，主要将语言视为一种社会行为形式，而不仅仅是描述现实的系统。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>言语行为理论&lt;/li>
&lt;li>语用学&lt;/li>
&lt;li>意图识别&lt;/li>
&lt;li>对话系统&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>设计对话式AI代理&lt;/li>
&lt;li>分析客户服务交互&lt;/li>
&lt;li>理解文本中的隐含意义&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/pragmatics-%E8%AF%AD%E7%94%A8%E5%AD%A6/">pragmatics (语用学)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/intent_classification-%E6%84%8F%E5%9B%BE%E5%88%86%E7%B1%BB/">intent_classification (意图分类)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dialogue_management-%E5%AF%B9%E8%AF%9D%E7%AE%A1%E7%90%86/">dialogue_management (对话管理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speech_act-%E8%A8%80%E8%AF%AD%E8%A1%8C%E4%B8%BA/">speech_act (言语行为)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>大模型作为裁判</title><link>https://terms-en.ai-term-hub.com/zh/terms/llm_as_a_judge/</link><pubDate>Sat, 18 Jul 2026 11:23:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/llm_as_a_judge/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>“大模型作为裁判”（LLM-as-a-Judge）是一种评估范式，其中大语言模型充当其他模型输出质量的自动化评估者。这种方法旨在减少对人工标注员或严格规则匹配的依赖，通过提示工程让LLM根据特定标准（如相关性、安全性、创造性等）对生成内容进行打分或排序，从而提高评估效率和一致性。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种通过使用另一个大语言模型根据标准对响应进行评分或排名，从而评估大语言模型输出的方法。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自动化评估&lt;/li>
&lt;li>提示工程&lt;/li>
&lt;li>模型对齐&lt;/li>
&lt;li>质量指标&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>RLHF模型的基准测试&lt;/li>
&lt;li>创意写作评估&lt;/li>
&lt;li>安全性和偏见检测&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rlhf-%E5%9F%BA%E4%BA%8E%E4%BA%BA%E7%B1%BB%E5%8F%8D%E9%A6%88%E7%9A%84%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/">rlhf (基于人类反馈的强化学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/evaluation_metrics-%E8%AF%84%E4%BC%B0%E6%8C%87%E6%A0%87/">evaluation_metrics (评估指标)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompt_engineering-%E6%8F%90%E7%A4%BA%E5%B7%A5%E7%A8%8B/">prompt_engineering (提示工程)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/model_alignment-%E6%A8%A1%E5%9E%8B%E5%AF%B9%E9%BD%90/">model_alignment (模型对齐)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>知识图谱嵌入</title><link>https://terms-en.ai-term-hub.com/zh/terms/knowledge_graph_embedding/</link><pubDate>Sat, 18 Jul 2026 11:23:17 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/knowledge_graph_embedding/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>知识图谱嵌入方法（如 TransE 或 DistMult）将离散的图结构转换为低维稠密向量。这使得机器学习模型能够执行数学运算，从而进行链接预测等任务。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种将知识图谱中的实体和关系映射到连续向量空间，同时保留结构语义的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>向量表示&lt;/li>
&lt;li>链接预测&lt;/li>
&lt;li>语义保留&lt;/li>
&lt;li>平移模型&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>利用关系数据的推荐系统&lt;/li>
&lt;li>基于结构化数据库的问答系统&lt;/li>
&lt;li>实体解析与匹配&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/word2vec-%E8%AF%8D%E5%B5%8C%E5%85%A5%E6%A8%A1%E5%9E%8B/">Word2Vec (词嵌入模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/graph-neural-networks-%E5%9B%BE%E7%A5%9E%E7%BB%8F%E7%BD%91%E7%BB%9C/">Graph Neural Networks (图神经网络)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transe-%E5%B9%B3%E7%A7%BB%E5%B5%8C%E5%85%A5%E7%AE%97%E6%B3%95/">TransE (平移嵌入算法)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/node-embedding-%E8%8A%82%E7%82%B9%E5%B5%8C%E5%85%A5/">Node Embedding (节点嵌入)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>智能文字识别</title><link>https://terms-en.ai-term-hub.com/zh/terms/intelligent_word_recognition/</link><pubDate>Sat, 18 Jul 2026 11:22:43 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/intelligent_word_recognition/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>智能文字识别是指由神经网络驱动的高级光学字符识别（OCR）技术。它超越了简单的模式匹配，通过理解上下文并处理n&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>利用人工智能算法（特别是深度学习）准确识别和解读来自图像或手写来源的文本。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>光学字符识别&lt;/li>
&lt;li>深度学习&lt;/li>
&lt;li>计算机视觉&lt;/li>
&lt;li>文本提取&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>历史档案数字化&lt;/li>
&lt;li>发票自动化处理&lt;/li>
&lt;li>实时翻译应用&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ocr-%E5%85%89%E5%AD%A6%E5%AD%97%E7%AC%A6%E8%AF%86%E5%88%AB/">OCR (光学字符识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/image-classification-%E5%9B%BE%E5%83%8F%E5%88%86%E7%B1%BB/">Image Classification (图像分类)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/handwriting-recognition-%E6%89%8B%E5%86%99%E8%AF%86%E5%88%AB/">Handwriting Recognition (手写识别)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>GPT-2</title><link>https://terms-en.ai-term-hub.com/zh/terms/gpt2/</link><pubDate>Sat, 18 Jul 2026 11:19:45 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/gpt2/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>生成式预训练 Transformer 2（GPT-2）是一个自回归语言模型，它利用 Transformer 架构来生成类人文本。它在海量的互联网文本数据集上进行了训练。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>GPT-2 是由 OpenAI 开发的一个基于 Transformer 架构的大规模语言模型，用于文本生成和理解。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>Transformer 架构&lt;/li>
&lt;li>自回归建模&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;li>少样本学习&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>文本补全与生成&lt;/li>
&lt;li>机器翻译&lt;/li>
&lt;li>对话智能体&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E5%9E%8B%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大型语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert/">BERT&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6/">注意力机制&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%88%86%E8%AF%8D/">分词&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>GPT-5.6</title><link>https://terms-en.ai-term-hub.com/zh/terms/gpt_56/</link><pubDate>Sat, 18 Jul 2026 11:18:04 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/gpt_56/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>GPT-5.6指的是OpenAI大型语言模型谱系中一个推测性的或即将推出的版本。虽然具体细节可能因开发时间线的不同而有所差异，但此类迭代通常体现了技术的进一步演进。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>OpenAI的生成式预训练Transformer系列的一个假设性或未来迭代版本，代表了超越当前GPT模型的进步。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>大型语言模型&lt;/li>
&lt;li>Transformer架构&lt;/li>
&lt;li>模型迭代&lt;/li>
&lt;li>通用人工智能&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>高级自然语言处理&lt;/li>
&lt;li>复杂推理任务&lt;/li>
&lt;li>多模态内容生成&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpt-4/">GPT-4&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E5%9E%8B%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大型语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/openai/">OpenAI&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%A5%9E%E7%BB%8F%E7%BD%91%E7%BB%9C/">神经网络&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>掩码填充</title><link>https://terms-en.ai-term-hub.com/zh/terms/fill_mask/</link><pubDate>Sat, 18 Jul 2026 11:17:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/fill_mask/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>掩码填充是 BERT 等基于 Transformer 模型中使用的一种基本预训练目标。该过程涉及掩盖文本序列中的随机标记，并训练模型预测被掩盖的原始标记，从而学习语言的深层语义表示。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种自然语言处理任务，模型根据上下文预测句子中缺失的标记。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>掩码语言建模&lt;/li>
&lt;li>上下文理解&lt;/li>
&lt;li>自监督学习&lt;/li>
&lt;li>标记预测&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>文本补全&lt;/li>
&lt;li>语义角色标注&lt;/li>
&lt;li>预训练基础&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BA/">BERT (双向编码器表示)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/masked-language-model-%E6%8E%A9%E7%A0%81%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">Masked Language Model (掩码语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E8%BD%AC%E6%8D%A2%E5%99%A8%E6%9E%B6%E6%9E%84/">Transformer (转换器架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenization-%E5%88%86%E8%AF%8D/">Tokenization (分词)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>ExBERT</title><link>https://terms-en.ai-term-hub.com/zh/terms/exbert/</link><pubDate>Sat, 18 Jul 2026 11:16:37 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/exbert/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>ExBERT通过分析不同层中单个注意力头的重要性，为BERT Transformer模型提供可解释性。它使用基于梯度的归因或其他技术来量化每个组件对最终预测的贡献，从而帮助理解模型的决策过程。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种通过识别对特定输出贡献最大的注意力头和层来解释BERT预测的方法。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>注意力头&lt;/li>
&lt;li>模型可解释性&lt;/li>
&lt;li>梯度归因&lt;/li>
&lt;li>Transformer架构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>调试NLP模型&lt;/li>
&lt;li>理解句法与语义处理&lt;/li>
&lt;li>医疗文本分析中的可解释AI&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-bert%E6%A8%A1%E5%9E%8B/">bert (BERT模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/attention_mechanism-%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6/">attention_mechanism (注意力机制)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/explainable_ai-%E5%8F%AF%E8%A7%A3%E9%87%8A%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD/">explainable_ai (可解释人工智能)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer_models-transformer%E6%A8%A1%E5%9E%8B/">transformer_models (Transformer模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>ELMo</title><link>https://terms-en.ai-term-hub.com/zh/terms/elmo/</link><pubDate>Sat, 18 Jul 2026 11:15:33 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/elmo/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>ELMo通过将输入文本通过在大型语料库上训练的双向LSTM进行处理，生成上下文敏感的词嵌入。与Word2Vec等静态嵌入不同，ELMo通过产生不同的向量表示来捕捉一词多义现象。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>来自语言模型的嵌入，一种使用双向LSTM的深度上下文词表示方法。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>上下文嵌入&lt;/li>
&lt;li>双向LSTM&lt;/li>
&lt;li>预训练&lt;/li>
&lt;li>一词多义&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>情感分析&lt;/li>
&lt;li>命名实体识别&lt;/li>
&lt;li>共指消解&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/word2vec-word2vec/">Word2Vec (Word2Vec)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-bert/">BERT (BERT)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-transformer/">Transformer (Transformer)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/language-modeling-%E8%AF%AD%E8%A8%80%E5%BB%BA%E6%A8%A1/">Language Modeling (语言建模)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>文档分类</title><link>https://terms-en.ai-term-hub.com/zh/terms/document_classification/</link><pubDate>Sat, 18 Jul 2026 11:15:33 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/document_classification/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>文档分类是一项基本的自然语言处理任务，算法在此过程中为无结构文本数据分配标签。它涉及从文档中提取特征，并将它们映射到特定的类别中。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>根据内容将文本文档归类到预定义组的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本预处理&lt;/li>
&lt;li>特征提取&lt;/li>
&lt;li>监督学习&lt;/li>
&lt;li>标注&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>电子邮件服务中的垃圾邮件过滤&lt;/li>
&lt;li>新闻自动标记&lt;/li>
&lt;li>法律文档分类&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/named-entity-recognition-%E5%91%BD%E5%90%8D%E5%AE%9E%E4%BD%93%E8%AF%86%E5%88%AB/">Named Entity Recognition (命名实体识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/text-mining-%E6%96%87%E6%9C%AC%E6%8C%96%E6%8E%98/">Text Mining (文本挖掘)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/naive-bayes-%E6%9C%B4%E7%B4%A0%E8%B4%9D%E5%8F%B6%E6%96%AF/">Naive Bayes (朴素贝叶斯)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：TriviaQA</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasettrivia_qa/</link><pubDate>Sat, 18 Jul 2026 11:13:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasettrivia_qa/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>TriviaQA 是一个专为开放域问答设计的数据集，包含超过一百万个问题及其对应的答案。该数据集旨在通过要求模型进行多步推理和知识整合来挑战现有模型的性能。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个用于开放域问答的大规模数据集，包含数百万个问题及其答案，涵盖各种 trivia（冷知识）领域。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>开放域问答&lt;/li>
&lt;li>多跳推理&lt;/li>
&lt;li>知识整合&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练检索增强生成模型&lt;/li>
&lt;li>评估大语言模型的事实一致性&lt;/li>
&lt;li>基准测试基于知识的问答系统&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/squad-%E6%96%AF%E5%9D%A6%E7%A6%8F%E9%97%AE%E7%AD%94%E6%95%B0%E6%8D%AE%E9%9B%86/">SQuAD (斯坦福问答数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/hotpotqa-%E7%83%AD%E7%82%B9%E9%97%AE%E7%AD%94%E6%95%B0%E6%8D%AE%E9%9B%86/">HotpotQA (热点问答数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-questions-%E8%87%AA%E7%84%B6%E9%97%AE%E9%A2%98%E6%95%B0%E6%8D%AE%E9%9B%86/">Natural Questions (自然问题数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/retrieval-augmented-generation-%E6%A3%80%E7%B4%A2%E5%A2%9E%E5%BC%BA%E7%94%9F%E6%88%90/">Retrieval-Augmented Generation (检索增强生成)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：WikiHow</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetwikihow/</link><pubDate>Sat, 18 Jul 2026 11:13:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetwikihow/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>WikiHow 数据集包含从 WikiHow 网站收集的约 60,000 篇操作指南文章。它广泛应用于自然语言处理研究中，用于抽象式文本摘要、步骤提取等任务。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个由 WikiHow 上的操作指南文章组成的大规模数据集，主要用于文本摘要和指令生成任务。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本摘要&lt;/li>
&lt;li>程序性文本&lt;/li>
&lt;li>指令提取&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>抽象式摘要模型训练&lt;/li>
&lt;li>分步指令生成&lt;/li>
&lt;li>理解程序性语言结构&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/cnn-dailymail-cnn-%E6%AF%8F%E6%97%A5%E9%82%AE%E6%8A%A5%E6%95%B0%E6%8D%AE%E9%9B%86/">CNN/DailyMail (CNN/每日邮报数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/xsum-%E6%9E%81%E7%AB%AF%E6%91%98%E8%A6%81%E6%95%B0%E6%8D%AE%E9%9B%86/">XSum (极端摘要数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/text-summarization-%E6%96%87%E6%9C%AC%E6%91%98%E8%A6%81/">Text Summarization (文本摘要)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/instruction-tuning-%E6%8C%87%E4%BB%A4%E5%BE%AE%E8%B0%83/">Instruction Tuning (指令微调)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：Wikipedia</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetwikipedia/</link><pubDate>Sat, 18 Jul 2026 11:13:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetwikipedia/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>维基百科是可用文本格式的人类知识最大且最全面的集合之一。在人工智能领域，它是预训练大型语言模型的主要来源，提供了丰富的语言模式和事实知识。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>来自维基百科的海量文本集合，作为预训练语言模型和知识密集型自然语言处理任务的基础语料库。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>预训练语料库&lt;/li>
&lt;li>知识库&lt;/li>
&lt;li>语言多样性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>预训练基础语言模型&lt;/li>
&lt;li>实体链接与消歧&lt;/li>
&lt;li>事实核查与知识检索&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/common-crawl-%E9%80%9A%E7%94%A8%E7%BD%91%E9%A1%B5%E7%88%AC%E5%8F%96%E6%95%B0%E6%8D%AE/">Common Crawl (通用网页爬取数据)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bookcorpus-%E4%B9%A6%E7%B1%8D%E8%AF%AD%E6%96%99%E5%BA%93/">BookCorpus (书籍语料库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BAtransformer/">BERT (双向编码器表示Transformer)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/knowledge-graph-%E7%9F%A5%E8%AF%86%E5%9B%BE%E8%B0%B1/">Knowledge Graph (知识图谱)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：Yahoo Answers Topics</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetyahoo_answers_topics/</link><pubDate>Sat, 18 Jul 2026 11:13:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetyahoo_answers_topics/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Yahoo Answers Topics 数据集是更大的雅虎问答档案的子集，专注于组织成不同主题类别的问题和答案。它常用于文本分类、语义相似性分析和非正式语言模式研究。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>源自雅虎问答的数据集，包含按特定主题分类的用户生成问题和答案，用于语义匹配和分类任务。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本分类&lt;/li>
&lt;li>语义相似度&lt;/li>
&lt;li>用户生成内容&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练问题分类模型&lt;/li>
&lt;li>语义文本相似度基准测试&lt;/li>
&lt;li>分析非正式语言模式&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ag-news-ag%E6%96%B0%E9%97%BB%E6%95%B0%E6%8D%AE%E9%9B%86/">AG News (AG新闻数据集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/snli-%E6%96%AF%E5%9D%A6%E7%A6%8F%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E6%8E%A8%E7%90%86%E8%AF%AD%E6%96%99%E5%BA%93/">SNLI (斯坦福自然语言推理语料库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/text-classification-%E6%96%87%E6%9C%AC%E5%88%86%E7%B1%BB/">Text Classification (文本分类)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic-search-%E8%AF%AD%E4%B9%89%E6%90%9C%E7%B4%A2/">Semantic Search (语义搜索)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：S2ORC</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasets2orc/</link><pubDate>Sat, 18 Jul 2026 11:13:26 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasets2orc/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>S2ORC 是从 Semantic Scholar 派生的学术文章综合语料库。它包括数百万篇跨各个科学领域的论文的全文本内容、元数据和引用关系。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Semantic Scholar开放研究语料库，一个包含结构化元数据和引用网络的大规模学术论文数据集。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>学术自然语言处理&lt;/li>
&lt;li>引用网络&lt;/li>
&lt;li>元数据&lt;/li>
&lt;li>学术数据&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>构建引用推荐系统&lt;/li>
&lt;li>科学文本分类&lt;/li>
&lt;li>从研究论文中提取实体&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantic-scholar-%E8%AF%AD%E4%B9%89%E5%AD%A6%E8%80%85%E6%90%9C%E7%B4%A2%E5%BC%95%E6%93%8E/">Semantic Scholar (语义学者搜索引擎)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/acl-anthology-%E8%AE%A1%E7%AE%97%E8%AF%AD%E8%A8%80%E5%AD%A6%E5%8D%8F%E4%BC%9A%E8%AE%BA%E6%96%87%E9%9B%86/">ACL Anthology (计算语言学协会论文集)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/arxiv-%E9%A2%84%E5%8D%B0%E6%9C%AC%E6%9C%8D%E5%8A%A1%E5%99%A8/">ArXiv (预印本服务器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/citation-prediction-%E5%BC%95%E7%94%A8%E9%A2%84%E6%B5%8B/">Citation Prediction (引用预测)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：嵌入数据/Altlex</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_dataaltlex/</link><pubDate>Sat, 18 Jul 2026 11:12:55 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_dataaltlex/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Altlex 数据集由共享相同底层含义但使用不同词汇或句法结构的句子对组成。它主要用于训练嵌入模型，以捕捉语义上的等价关系。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>包含用于训练模型进行语义等价和同义句检测的替代词汇形式的语料库。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>语义等价&lt;/li>
&lt;li>同义句检测&lt;/li>
&lt;li>词汇变化&lt;/li>
&lt;li>向量相似度&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练语义搜索引擎&lt;/li>
&lt;li>改进问答系统&lt;/li>
&lt;li>增强文本相似度度量&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%AD%E4%B9%89%E5%B5%8C%E5%85%A5-semantic-embeddings/">语义嵌入 (Semantic Embeddings)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%90%8C%E4%B9%89%E5%8F%A5%E8%AF%86%E5%88%AB-paraphrase-identification/">同义句识别 (Paraphrase Identification)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%8D%E4%B9%89%E6%B6%88%E6%AD%A7-word-sense-disambiguation/">词义消歧 (Word Sense Disambiguation)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：嵌入数据/QQP</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_dataqqp/</link><pubDate>Sat, 18 Jul 2026 11:12:55 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_dataqqp/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Quora 问题对（QQP）是一个二元分类数据集，包含来自 Quora 平台的超过 40 万个问题对。其任务是确定两个问题是否具有相同的意图或含义。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Quora 问题对数据集，用于训练模型检测问题之间的语义相似度。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>语义文本相似度&lt;/li>
&lt;li>二元分类&lt;/li>
&lt;li>意图匹配&lt;/li>
&lt;li>句子嵌入&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>重复问题检测&lt;/li>
&lt;li>微调句子转换器&lt;/li>
&lt;li>改进 FAQ 机器人&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sts-benchmark/">STS Benchmark&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AD%AA%E7%94%9F%E7%BD%91%E7%BB%9C-siamese-networks/">孪生网络 (Siamese Networks)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%96%87%E6%9C%AC%E8%95%B4%E5%90%AB-textual-entailment/">文本蕴含 (Textual Entailment)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：嵌入数据/句子压缩</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_datasentence_compression/</link><pubDate>Sat, 18 Jul 2026 11:12:55 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetembedding_datasentence_compression/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>句子压缩数据集由成对的句子组成，其中目标句子是源句子的缩短版本，在去除冗余信息的同时保留核心含义。这些数据集常用于训练能够理解信息密度的模型。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>包含原始句子及其压缩版本的数据集，用于训练模型在保留信息方面的能力。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>信息密度&lt;/li>
&lt;li>结构简化&lt;/li>
&lt;li>摘要&lt;/li>
&lt;li>语义保留&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自动文本摘要&lt;/li>
&lt;li>训练感知压缩的嵌入模型&lt;/li>
&lt;li>改进可读性指标&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%96%87%E6%9C%AC%E6%91%98%E8%A6%81-text-summarization/">文本摘要 (Text Summarization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%8A%BD%E8%B1%A1%E5%BC%8F%E5%8E%8B%E7%BC%A9-abstractive-compression/">抽象式压缩 (Abstractive Compression)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%AD%E8%A8%80%E6%95%88%E7%8E%87-linguistic-efficiency/">语言效率 (Linguistic Efficiency)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据集：BookCorpus</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetbookcorpus/</link><pubDate>Sat, 18 Jul 2026 11:12:42 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetbookcorpus/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>BookCorpus 是从互联网上爬取的来自 10,000 多本未出版书籍的文本集合。它是训练和评估自然语言处理（NLP）模型的基础资源。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一个包含超过 10,000 本未出版书籍的大规模数据集，广泛用于自然语言处理模型的预训练。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>NLP 预训练&lt;/li>
&lt;li>文本语料库&lt;/li>
&lt;li>语言模型&lt;/li>
&lt;li>未出版书籍&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>预训练 Transformer 模型&lt;/li>
&lt;li>评估语言流畅度&lt;/li>
&lt;li>文学文本分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wikitext/">WikiText&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/common-crawl/">Common Crawl&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert/">BERT&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpt/">GPT&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>条件随机场</title><link>https://terms-en.ai-term-hub.com/zh/terms/conditional_random_field/</link><pubDate>Sat, 18 Jul 2026 11:11:03 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/conditional_random_field/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>条件随机场（CRF）是一类判别式模型，常用于自然语言处理和生物信息学。与生成模型不同，CRF直接对给定输入序列下标签序列的条件概率进行建模。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>条件随机场是一种判别式概率模型，用于序列标注等结构化预测任务。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>判别式模型&lt;/li>
&lt;li>结构化预测&lt;/li>
&lt;li>序列标注&lt;/li>
&lt;li>全局归一化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>命名实体识别&lt;/li>
&lt;li>词性标注&lt;/li>
&lt;li>生物信息学序列分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/hidden-markov-model-%E9%9A%90%E9%A9%AC%E5%B0%94%E5%8F%AF%E5%A4%AB%E6%A8%A1%E5%9E%8B/">Hidden Markov Model (隐马尔可夫模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sequence-modeling-%E5%BA%8F%E5%88%97%E5%BB%BA%E6%A8%A1/">Sequence Modeling (序列建模)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/log-linear-model-%E5%AF%B9%E6%95%B0%E7%BA%BF%E6%80%A7%E6%A8%A1%E5%9E%8B/">Log-linear Model (对数线性模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>计算幽默</title><link>https://terms-en.ai-term-hub.com/zh/terms/computational_humor/</link><pubDate>Sat, 18 Jul 2026 11:10:50 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/computational_humor/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>计算幽默研究机器如何产生或解读笑话、双关语和机智言论。它通常依赖自然语言处理来检测不协调、语义转换或未预期的结果。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>专注于通过计算方法生成、理解和欣赏幽默内容的AI子领域。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>不协调理论&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;li>语义转换&lt;/li>
&lt;li>生成模型&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>聊天机器人个性增强&lt;/li>
&lt;li>自动化笑话生成&lt;/li>
&lt;li>创意写作辅助工具&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentiment-analysis-%E6%83%85%E6%84%9F%E5%88%86%E6%9E%90/">Sentiment Analysis (情感分析)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/creativity-in-ai-ai%E5%88%9B%E9%80%A0%E5%8A%9B/">Creativity in AI (AI创造力)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/linguistics-%E8%AF%AD%E8%A8%80%E5%AD%A6/">Linguistics (语言学)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>常识知识</title><link>https://terms-en.ai-term-hub.com/zh/terms/commonsense_knowledge/</link><pubDate>Sat, 18 Jul 2026 11:10:38 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/commonsense_knowledge/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>常识知识指的是人类自然获得的关于日常生活、物理学、社会规范和因果关系的庞大隐性信息库。在人工智能领域，获取这种知识是实现真正智能的关键挑战，因为机器往往难以像人类一样理解隐含的情境和现实世界的逻辑。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>人类拥有但机器通常缺乏的关于物理和社会世界的背景信息与直觉理解。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>隐性知识&lt;/li>
&lt;li>推理&lt;/li>
&lt;li>语境理解&lt;/li>
&lt;li>类人直觉&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自然语言处理中的消歧&lt;/li>
&lt;li>机器人导航与交互&lt;/li>
&lt;li>问答系统&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%9F%A5%E8%AF%86%E5%9B%BE%E8%B0%B1-knowledge-graphs/">知识图谱 (Knowledge Graphs)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%8E%A8%E7%90%86-reasoning/">推理 (Reasoning)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86-nlp/">自然语言处理 (NLP)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%85%B7%E8%BA%AB%E6%99%BA%E8%83%BD-embodied-ai/">具身智能 (Embodied AI)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>字符计算</title><link>https://terms-en.ai-term-hub.com/zh/terms/character_computing/</link><pubDate>Sat, 18 Jul 2026 11:09:54 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/character_computing/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>这一概念侧重于文本操作，其中计算的基本单位是单个字符。它通常用于需要细粒度文本分析的任务，例如拼写检查。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>字符计算涉及在单个字符层面而非单词或句子层面处理、生成或分析文本。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>分词&lt;/li>
&lt;li>文本生成&lt;/li>
&lt;li>细粒度分析&lt;/li>
&lt;li>子词单元&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>手写识别（OCR）&lt;/li>
&lt;li>低资源语言建模&lt;/li>
&lt;li>密文分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/token-%E8%AF%8D%E5%85%83/">Token (词元)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rnn-%E5%BE%AA%E7%8E%AF%E7%A5%9E%E7%BB%8F%E7%BD%91%E7%BB%9C/">RNN (循环神经网络)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E5%8F%98%E6%8D%A2%E5%99%A8%E6%9E%B6%E6%9E%84/">Transformer (变换器架构)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Bloom</title><link>https://terms-en.ai-term-hub.com/zh/terms/bloom/</link><pubDate>Sat, 18 Jul 2026 11:09:30 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/bloom/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>虽然历史上指本杰明·布鲁姆的教育分类法，但在现代人工智能语境中，它通常指由BigScience开发的Bloom文本嵌入模型。该模型生成高质量的……&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>在机器学习中，“Bloom”通常指应用于人工智能教育的布鲁姆分类法，或特定的嵌入模型（如Bloom文本嵌入模型）。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本嵌入&lt;/li>
&lt;li>语义搜索&lt;/li>
&lt;li>布隆过滤器&lt;/li>
&lt;li>概率数据结构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自然语言处理 (NLP)&lt;/li>
&lt;li>数据库索引优化&lt;/li>
&lt;li>内容推荐系统&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5/">Embedding (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vector-database-%E5%90%91%E9%87%8F%E6%95%B0%E6%8D%AE%E5%BA%93/">Vector Database (向量数据库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/hash-function-%E5%93%88%E5%B8%8C%E5%87%BD%E6%95%B0/">Hash Function (哈希函数)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>BERT</title><link>https://terms-en.ai-term-hub.com/zh/terms/bert/</link><pubDate>Sat, 18 Jul 2026 11:09:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/bert/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>BERT（Bidirectional Encoder Representations from Transformers）是由Google开发的一种基于Transformer的机器学习技术，用于自然语言处理（NLP）的预训练。它利用掩码语言建模和下一句预测来学习双向表示。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>基于Transformer的双向编码器表示是一种预训练的自然语言处理模型。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>Transformer架构&lt;/li>
&lt;li>掩码语言建模&lt;/li>
&lt;li>预训练&lt;/li>
&lt;li>微调&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>问答系统&lt;/li>
&lt;li>情感分析&lt;/li>
&lt;li>文本分类&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpt-%E7%94%9F%E6%88%90%E5%BC%8F%E9%A2%84%E8%AE%AD%E7%BB%83%E5%8F%98%E6%8D%A2%E5%99%A8/">GPT (生成式预训练变换器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E5%8F%98%E6%8D%A2%E5%99%A8%E6%9E%B6%E6%9E%84/">Transformer (变换器架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embeddings-%E8%AF%8D%E5%B5%8C%E5%85%A5/">Embeddings (词嵌入)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>词袋模型</title><link>https://terms-en.ai-term-hub.com/zh/terms/bag_of_words_model/</link><pubDate>Sat, 18 Jul 2026 11:08:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/bag_of_words_model/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>这种自然语言处理技术将文本表示为单词的多重集， disregarding 句法和序列。它根据词频或存在性将文档转换为数值向量。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>词袋模型是一种简化的文本表示方法，描述文档中单词的出现情况，忽略语法和词序。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>分词&lt;/li>
&lt;li>频率统计&lt;/li>
&lt;li>向量空间&lt;/li>
&lt;li>特征提取&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>文本分类&lt;/li>
&lt;li>垃圾邮件过滤&lt;/li>
&lt;li>信息检索&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> sklearn.feature_extraction.text &lt;span style="color:#f92672">import&lt;/span> CountVectorizer
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>corpus &lt;span style="color:#f92672">=&lt;/span> [&lt;span style="color:#e6db74">&amp;#34;Hello world&amp;#34;&lt;/span>, &lt;span style="color:#e6db74">&amp;#34;World hello&amp;#34;&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>vectorizer &lt;span style="color:#f92672">=&lt;/span> CountVectorizer()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>X &lt;span style="color:#f92672">=&lt;/span> vectorizer&lt;span style="color:#f92672">.&lt;/span>fit_transform(corpus)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tf-idf-%E8%AF%8D%E9%A2%91-%E9%80%86%E6%96%87%E6%A1%A3%E9%A2%91%E7%8E%87/">TF-IDF (词频-逆文档频率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/n-grams-n%E5%85%83%E8%AF%AD%E6%B3%95/">N-grams (N元语法)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/word-embeddings-%E8%AF%8D%E5%B5%8C%E5%85%A5/">Word Embeddings (词嵌入)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>自动化医疗文书录入员</title><link>https://terms-en.ai-term-hub.com/zh/terms/automated_medical_scribe/</link><pubDate>Sat, 18 Jul 2026 11:08:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/automated_medical_scribe/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>自动化医疗文书录入员利用自然语言处理和语音识别技术，聆听医生与患者的对话，并创建结构化的电子健康记录。该技术旨在减轻医护人员的行政负担，提高数据记录的准确性和效率。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种由人工智能驱动的系统，能够自动生成基于医患互动的临床文档。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自然语言处理&lt;/li>
&lt;li>临床文档&lt;/li>
&lt;li>语音识别&lt;/li>
&lt;li>电子健康记录&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>患者就诊时的实时文档记录&lt;/li>
&lt;li>减少医生职业倦怠&lt;/li>
&lt;li>提高计费编码的准确性&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/clinical_nlp-%E4%B8%B4%E5%BA%8A%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">clinical_nlp (临床自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ehr_integration-%E7%94%B5%E5%AD%90%E5%81%A5%E5%BA%B7%E8%AE%B0%E5%BD%95%E9%9B%86%E6%88%90/">ehr_integration (电子健康记录集成)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/voice_assistants-%E8%AF%AD%E9%9F%B3%E5%8A%A9%E6%89%8B/">voice_assistants (语音助手)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/healthcare_ai-%E5%8C%BB%E7%96%97%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD/">healthcare_ai (医疗人工智能)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>ASR-complete</title><link>https://terms-en.ai-term-hub.com/zh/terms/asr_complete/</link><pubDate>Sat, 18 Jul 2026 11:04:01 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/asr_complete/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>术语“ASR-complete”表示自动语音识别系统在特定且定义明确的任务和数据集上，其性能已达到与人类转录员相当的水平。这是一个重要的里程碑，标志着系统在特定领域内的识别精度已满足实际应用的高标准要求。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>ASR-complete 描述在标准化基准数据集上达到人类水平准确率的语音识别系统。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>语音识别&lt;/li>
&lt;li>人类水平准确率&lt;/li>
&lt;li>错误率&lt;/li>
&lt;li>基准测试&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估 ASR 模型性能&lt;/li>
&lt;li>制定行业标准&lt;/li>
&lt;li>比较不同的声学模型&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/automatic-speech-recognition-%E8%87%AA%E5%8A%A8%E8%AF%AD%E9%9F%B3%E8%AF%86%E5%88%AB/">Automatic Speech Recognition (自动语音识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wer-%E8%AF%8D%E9%94%99%E8%AF%AF%E7%8E%87/">WER (词错误率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>视觉语言模型</title><link>https://terms-en.ai-term-hub.com/zh/terms/vision_language/</link><pubDate>Sat, 18 Jul 2026 11:02:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/vision_language/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>视觉语言模型通常被称为多模态大语言模型（MLLMs），它们整合了计算机视觉和自然语言处理技术。这些模型使AI能够理解图像并生成相应的文本描述或回答。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>视觉语言模型是处理并将视觉数据与文本信息相关联以理解多模态上下文的AI系统。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>多模态&lt;/li>
&lt;li>跨模态对齐&lt;/li>
&lt;li>图像描述生成&lt;/li>
&lt;li>视觉问答&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自动化图像描述生成&lt;/li>
&lt;li>视觉问答系统&lt;/li>
&lt;li>结合上下文的内容审核&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/clip-%E5%AF%B9%E6%AF%94%E8%AF%AD%E8%A8%80-%E5%9B%BE%E5%83%8F%E9%A2%84%E8%AE%AD%E7%BB%83%E6%A8%A1%E5%9E%8B/">CLIP (对比语言-图像预训练模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multimodal-ai-%E5%A4%9A%E6%A8%A1%E6%80%81%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD/">Multimodal AI (多模态人工智能)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-transformer%E6%9E%B6%E6%9E%84/">Transformer (Transformer架构)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Translation</title><link>https://terms-en.ai-term-hub.com/zh/terms/translation/</link><pubDate>Sat, 18 Jul 2026 11:02:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/translation/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>AI中的翻译指的是神经机器翻译，其中深度学习模型映射语言之间的语义表示。与基于规则的系统不同，现代方法学习上下文细微差别&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>将文本从源自然语言转换为目标自然语言，同时保留含义的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>神经机器翻译&lt;/li>
&lt;li>源语言&lt;/li>
&lt;li>目标语言&lt;/li>
&lt;li>语义保留&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>全球内容本地化&lt;/li>
&lt;li>实时聊天口译&lt;/li>
&lt;li>文档多语言化&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nmt-%E7%A5%9E%E7%BB%8F%E6%9C%BA%E5%99%A8%E7%BF%BB%E8%AF%91/">nmt (神经机器翻译)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/localization-%E6%9C%AC%E5%9C%B0%E5%8C%96/">localization (本地化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multilingual_models-%E5%A4%9A%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">multilingual_models (多语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/cross_lingual_transfer-%E8%B7%A8%E8%AF%AD%E8%A8%80%E8%BF%81%E7%A7%BB/">cross_lingual_transfer (跨语言迁移)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>文本摘要</title><link>https://terms-en.ai-term-hub.com/zh/terms/summarization/</link><pubDate>Sat, 18 Jul 2026 11:02:04 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/summarization/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>文本摘要将大量文本缩减为较短版本，而不丢失核心含义。它可以是抽取式的，即从源文本中选择重要句子；也可以是抽象式的，即生成新的概括性语句。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一项自然语言处理任务，生成较长文本的简洁连贯摘要，同时保留其关键信息。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>抽取式摘要&lt;/li>
&lt;li>抽象式摘要&lt;/li>
&lt;li>信息密度&lt;/li>
&lt;li>连贯性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>新闻文章精简&lt;/li>
&lt;li>会议纪要生成&lt;/li>
&lt;li>法律文档审查&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> pipeline
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>summarizer &lt;span style="color:#f92672">=&lt;/span> pipeline(&lt;span style="color:#e6db74">&amp;#34;summarization&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>result &lt;span style="color:#f92672">=&lt;/span> summarizer(&lt;span style="color:#e6db74">&amp;#34;AI is transforming industries...&amp;#34;&lt;/span>, max_length&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">50&lt;/span>, min_length&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">10&lt;/span>)[&lt;span style="color:#ae81ff">0&lt;/span>][&lt;span style="color:#e6db74">&amp;#39;summary_text&amp;#39;&lt;/span>]
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86-nlp/">自然语言处理 (NLP)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E6%A8%A1%E5%9E%8B/">Transformer 模型&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert/">BERT&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/t5/">T5&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>语义搜索</title><link>https://terms-en.ai-term-hub.com/zh/terms/semantic_search/</link><pubDate>Sat, 18 Jul 2026 11:01:52 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/semantic_search/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>语义搜索解释查询背后的意图和上下文含义，超越了简单的关键词匹配。它使用嵌入将文本表示为高维空间中的向量，从而允许&amp;hellip;&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种理解查询词含义而非仅匹配关键词的搜索技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>嵌入&lt;/li>
&lt;li>向量相似度&lt;/li>
&lt;li>意图识别&lt;/li>
&lt;li>上下文理解&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>企业文档搜索&lt;/li>
&lt;li>电子商务产品发现&lt;/li>
&lt;li>知识库查询&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>null
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/retrieval-%E6%A3%80%E7%B4%A2/">retrieval (检索)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5/">embedding (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vector_database-%E5%90%91%E9%87%8F%E6%95%B0%E6%8D%AE%E5%BA%93/">vector_database (向量数据库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural_language_processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">natural_language_processing (自然语言处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>位置编码</title><link>https://terms-en.ai-term-hub.com/zh/terms/positional_encoding/</link><pubDate>Sat, 18 Jul 2026 11:01:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/positional_encoding/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>由于 Transformer 并行处理所有令牌，而不是像循环神经网络（RNN）那样按顺序处理，因此它缺乏对令牌顺序的固有认知。位置编码通过向输入嵌入添加特定的向量来保留这种顺序信息。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种将序列中令牌的相对或绝对位置信息注入 Transformer 模型的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>序列顺序&lt;/li>
&lt;li>自注意力机制&lt;/li>
&lt;li>正弦函数&lt;/li>
&lt;li>令牌嵌入&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>机器翻译&lt;/li>
&lt;li>文本摘要&lt;/li>
&lt;li>语言建模&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> math
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">get_positional_encoding&lt;/span>(seq_len, d_model):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pe &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>zeros(seq_len, d_model)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> position &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>arange(&lt;span style="color:#ae81ff">0&lt;/span>, seq_len, dtype&lt;span style="color:#f92672">=&lt;/span>torch&lt;span style="color:#f92672">.&lt;/span>float)&lt;span style="color:#f92672">.&lt;/span>unsqueeze(&lt;span style="color:#ae81ff">1&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> div_term &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>exp(torch&lt;span style="color:#f92672">.&lt;/span>arange(&lt;span style="color:#ae81ff">0&lt;/span>, d_model, &lt;span style="color:#ae81ff">2&lt;/span>)&lt;span style="color:#f92672">.&lt;/span>float() &lt;span style="color:#f92672">*&lt;/span> (&lt;span style="color:#f92672">-&lt;/span>math&lt;span style="color:#f92672">.&lt;/span>log(&lt;span style="color:#ae81ff">10000.0&lt;/span>) &lt;span style="color:#f92672">/&lt;/span> d_model))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pe[:, &lt;span style="color:#ae81ff">0&lt;/span>::&lt;span style="color:#ae81ff">2&lt;/span>] &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>sin(position &lt;span style="color:#f92672">*&lt;/span> div_term)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pe[:, &lt;span style="color:#ae81ff">1&lt;/span>::&lt;span style="color:#ae81ff">2&lt;/span>] &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>cos(position &lt;span style="color:#f92672">*&lt;/span> div_term)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">return&lt;/span> pe&lt;span style="color:#f92672">.&lt;/span>unsqueeze(&lt;span style="color:#ae81ff">0&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E6%9E%B6%E6%9E%84/">Transformer 架构&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%B5%8C%E5%85%A5-embedding/">嵌入 (Embedding)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6-attention-mechanism/">注意力机制 (Attention Mechanism)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%97%8B%E8%BD%AC%E4%BD%8D%E7%BD%AE%E7%BC%96%E7%A0%81-rope/">旋转位置编码 (RoPE)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>问答</title><link>https://terms-en.ai-term-hub.com/zh/terms/question_answering/</link><pubDate>Sat, 18 Jul 2026 11:01:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/question_answering/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>问答（QA）涉及从给定上下文或知识库中检索或生成对用户查询的准确响应。它包括依赖特定文档的封闭领域问答，以及基于通用知识的开放领域问答。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一项自然语言处理任务，系统自动提供用自然语言提出的问题的精确答案。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>信息检索&lt;/li>
&lt;li>语义理解&lt;/li>
&lt;li>上下文提取&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>虚拟助手&lt;/li>
&lt;li>搜索引擎&lt;/li>
&lt;li>客户支持自动化&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%A3%80%E7%B4%A2%E5%A2%9E%E5%BC%BA%E7%94%9F%E6%88%90-retrieval-augmented-generation/">检索增强生成 (Retrieval-Augmented Generation)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%91%BD%E5%90%8D%E5%AE%9E%E4%BD%93%E8%AF%86%E5%88%AB-named-entity-recognition/">命名实体识别 (Named Entity Recognition)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%96%87%E6%9C%AC%E6%91%98%E8%A6%81-text-summarization/">文本摘要 (Text Summarization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%84%8F%E5%9B%BE%E5%88%86%E7%B1%BB-intent-classification/">意图分类 (Intent Classification)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>命名实体识别</title><link>https://terms-en.ai-term-hub.com/zh/terms/named_entity_recognition/</link><pubDate>Sat, 18 Jul 2026 11:01:16 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/named_entity_recognition/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>命名实体识别（NER）是信息抽取的一个子任务，用于在文本中定位并将命名实体分类为预定义的类别，如人名、组织机构名、地名、医疗术语等。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一项自然语言处理任务，旨在将关键信息实体识别并分类到预定义的类别中。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>实体类型化&lt;/li>
&lt;li>词元分类&lt;/li>
&lt;li>序列标注&lt;/li>
&lt;li>BiLSTM-CRF 架构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>简历解析&lt;/li>
&lt;li>客户支持意图检测&lt;/li>
&lt;li>病历分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/information_extraction-%E4%BF%A1%E6%81%AF%E6%8A%BD%E5%8F%96/">information_extraction (信息抽取)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/part_of_speech_tagging-%E8%AF%8D%E6%80%A7%E6%A0%87%E6%B3%A8/">part_of_speech_tagging (词性标注)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/coreference_resolution-%E6%8C%87%E4%BB%A3%E6%B6%88%E8%A7%A3/">coreference_resolution (指代消解)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenization-%E5%88%86%E8%AF%8D/">tokenization (分词)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Embedding Model</title><link>https://terms-en.ai-term-hub.com/zh/terms/embedding_model/</link><pubDate>Sat, 18 Jul 2026 10:59:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/embedding_model/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>这些模型将高维数据映射到低维连续向量空间中，其中相似的项目彼此靠得更近。这种转换捕捉了语义关系，使得&amp;hellip;（原文截断）&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>嵌入模型将文本或图像等原始数据转换为表示语义意义的稠密数值向量。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>向量表示&lt;/li>
&lt;li>语义相似度&lt;/li>
&lt;li>降维&lt;/li>
&lt;li>特征提取&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>构建语义搜索引擎&lt;/li>
&lt;li>产品或内容推荐系统&lt;/li>
&lt;li>聚类相似的文档或图像&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> AutoTokenizer, AutoModel
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>model &lt;span style="color:#f92672">=&lt;/span> AutoModel&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;sentence-transformers/all-MiniLM-L6-v2&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokenizer &lt;span style="color:#f92672">=&lt;/span> AutoTokenizer&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;sentence-transformers/all-MiniLM-L6-v2&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>inputs &lt;span style="color:#f92672">=&lt;/span> tokenizer(&lt;span style="color:#e6db74">&amp;#39;Hello world&amp;#39;&lt;/span>, return_tensors&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;pt&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>embeddings &lt;span style="color:#f92672">=&lt;/span> model(&lt;span style="color:#f92672">**&lt;/span>inputs)&lt;span style="color:#f92672">.&lt;/span>last_hidden_state&lt;span style="color:#f92672">.&lt;/span>mean(dim&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">1&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/word2vec-%E8%AF%8D%E5%90%91%E9%87%8F%E6%A8%A1%E5%9E%8B/">Word2Vec (词向量模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E9%A2%84%E8%AE%AD%E7%BB%83%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">BERT (预训练语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vector-database-%E5%90%91%E9%87%8F%E6%95%B0%E6%8D%AE%E5%BA%93/">Vector Database (向量数据库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/similarity-search-%E7%9B%B8%E4%BC%BC%E5%BA%A6%E6%90%9C%E7%B4%A2/">Similarity Search (相似度搜索)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>解码器</title><link>https://terms-en.ai-term-hub.com/zh/terms/decoder/</link><pubDate>Sat, 18 Jul 2026 10:59:40 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/decoder/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在序列到序列（Seq2Seq）模型中，解码器接收由编码器生成的上下文向量，并逐步生成目标输出。它利用注意力机制来关注输入序列的相关部分，从而准确预测下一个输出元素。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>负责从编码后的潜在表示生成输出序列的神经网络组件。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>序列生成&lt;/li>
&lt;li>注意力机制&lt;/li>
&lt;li>潜在空间&lt;/li>
&lt;li>自回归预测&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>机器翻译（如英语到法语）&lt;/li>
&lt;li>文本摘要&lt;/li>
&lt;li>图像描述生成&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/encoder-%E7%BC%96%E7%A0%81%E5%99%A8/">Encoder (编码器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E5%8F%98%E6%8D%A2%E5%99%A8%E6%9E%B6%E6%9E%84/">Transformer (变换器架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rnn-%E5%BE%AA%E7%8E%AF%E7%A5%9E%E7%BB%8F%E7%BD%91%E7%BB%9C/">RNN (循环神经网络)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sequence-to-sequence-%E5%BA%8F%E5%88%97%E5%88%B0%E5%BA%8F%E5%88%97%E6%A8%A1%E5%9E%8B/">Sequence-to-Sequence (序列到序列模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>字节对编码 (BPE)</title><link>https://terms-en.ai-term-hub.com/zh/terms/bpe/</link><pubDate>Sat, 18 Jul 2026 10:59:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/bpe/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>字节对编码（BPE）是一种数据压缩技术，经过调整后应用于自然语言处理中，以处理未登录词（Out-of-Vocabulary）。它从单个字符的词汇表开始，并迭代地合并最频繁出现的字符对，直到达到预定的词汇表大小或收敛。这种方法允许模型将罕见词分解为更常见的子词单元，从而提高对未知词汇的处理能力。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>字节对编码是一种用于子词分词的算法，它通过迭代合并出现频率最高的字符对来构建词汇表。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>子词分词&lt;/li>
&lt;li>词汇表合并&lt;/li>
&lt;li>频率分析&lt;/li>
&lt;li>未登录词处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为大语言模型预处理文本&lt;/li>
&lt;li>处理形态丰富的语言&lt;/li>
&lt;li>减少神经网络中的词汇表规模&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> tiktoken
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>enc &lt;span style="color:#f92672">=&lt;/span> tiktoken&lt;span style="color:#f92672">.&lt;/span>get_encoding(&lt;span style="color:#e6db74">&amp;#34;cl100k_base&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokens &lt;span style="color:#f92672">=&lt;/span> enc&lt;span style="color:#f92672">.&lt;/span>encode(&lt;span style="color:#e6db74">&amp;#34;unhappiness&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(tokens)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wordpiece-wordpiece%E5%88%86%E8%AF%8D%E7%AE%97%E6%B3%95/">WordPiece (WordPiece分词算法)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentencepiece-sentencepiece%E5%88%86%E8%AF%8D%E5%B7%A5%E5%85%B7%E5%BA%93/">SentencePiece (SentencePiece分词工具库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenization-%E5%88%86%E8%AF%8D/">Tokenization (分词)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/subword-units-%E5%AD%90%E8%AF%8D%E5%8D%95%E5%85%83/">Subword Units (子词单元)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>自监督</title><link>https://terms-en.ai-term-hub.com/zh/terms/self_supervised/</link><pubDate>Sat, 18 Jul 2026 10:57:22 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/self_supervised/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>自监督学习是机器学习的一个子集，其监督信号自动从数据本身派生，消除了手动标注的需求。模型通常通过解决预设的代理任务来学习数据的内在结构和表示。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>自监督学习是一种技术，模型从输入数据中自动生成标签以学习表示，无需人工标注。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>pretext任务 (代理任务)&lt;/li>
&lt;li>掩码建模&lt;/li>
&lt;li>无标签数据&lt;/li>
&lt;li>表示学习&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>通过掩码语言建模训练BERT&lt;/li>
&lt;li>用于图像嵌入的对比学习&lt;/li>
&lt;li>预测大语言模型中的下一个词元&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/unsupervised-%E6%97%A0%E7%9B%91%E7%9D%A3/">unsupervised (无监督)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/contrastive_learning-%E5%AF%B9%E6%AF%94%E5%AD%A6%E4%B9%A0/">contrastive_learning (对比学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/masked_language_modeling-%E6%8E%A9%E7%A0%81%E8%AF%AD%E8%A8%80%E5%BB%BA%E6%A8%A1/">masked_language_modeling (掩码语言建模)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/representation_learning-%E8%A1%A8%E7%A4%BA%E5%AD%A6%E4%B9%A0/">representation_learning (表示学习)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>少样本</title><link>https://terms-en.ai-term-hub.com/zh/terms/few_shot/</link><pubDate>Sat, 18 Jul 2026 10:56:23 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/few_shot/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>少样本学习使机器学习模型能够从极其有限的数据中进行泛化，通常每个类别仅需一到十个示例。与需要数千个示例的传统监督学习不同，少样本学习利用预训练模型中提取的通用特征，使其能够在新任务上快速适应并表现良好。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种学习范式，模型在仅接触少量标注示例后便能正确执行任务。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>元学习&lt;/li>
&lt;li>泛化能力&lt;/li>
&lt;li>标签效率&lt;/li>
&lt;li>预训练&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>罕见病诊断&lt;/li>
&lt;li>聊天机器人中的自定义意图识别&lt;/li>
&lt;li>有限数据下的领域自适应&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/zero_shot-%E9%9B%B6%E6%A0%B7%E6%9C%AC/">zero_shot (零样本)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/one_shot-%E5%8D%95%E6%A0%B7%E6%9C%AC/">one_shot (单样本)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transfer_learning-%E8%BF%81%E7%A7%BB%E5%AD%A6%E4%B9%A0/">transfer_learning (迁移学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/meta_learning-%E5%85%83%E5%AD%A6%E4%B9%A0/">meta_learning (元学习)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Transformer</title><link>https://terms-en.ai-term-hub.com/zh/terms/transformer/</link><pubDate>Sat, 18 Jul 2026 10:55:40 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/transformer/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Transformer架构在《Attention Is All You Need》论文中被提出，彻底革新了自然语言处理及更多领域。它使用多头自注意力机制来权衡输入序列中不同部分的重要性。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种基于自注意力机制的深度学习架构，能够并行而非顺序地处理序列数据。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自注意力&lt;/li>
&lt;li>多头注意力&lt;/li>
&lt;li>位置编码&lt;/li>
&lt;li>编码器-解码器结构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>机器翻译&lt;/li>
&lt;li>文本生成&lt;/li>
&lt;li>图像识别（ViT）&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch.nn &lt;span style="color:#66d9ef">as&lt;/span> nn
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>attention_layer &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>MultiheadAttention(embed_dim&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">512&lt;/span>, num_heads&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">8&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/attention_mechanism-%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6/">attention_mechanism (注意力机制)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-bert%E6%A8%A1%E5%9E%8B/">bert (BERT模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpt-gpt%E6%A8%A1%E5%9E%8B/">gpt (GPT模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/self_attention-%E8%87%AA%E6%B3%A8%E6%84%8F%E5%8A%9B/">self_attention (自注意力)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>词元</title><link>https://terms-en.ai-term-hub.com/zh/terms/token/</link><pubDate>Sat, 18 Jul 2026 10:55:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/token/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>词元是文本或数据的离散单元，作为自然语言处理模型的基本输入元素。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>离散单元&lt;/li>
&lt;li>词汇表&lt;/li>
&lt;li>嵌入&lt;/li>
&lt;li>上下文窗口&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>LLM中的文本生成&lt;/li>
&lt;li>机器翻译预处理&lt;/li>
&lt;li>情感分析输入格式化&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenization-%E5%88%86%E8%AF%8D/">Tokenization (分词)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5/">Embedding (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vocabulary-%E8%AF%8D%E6%B1%87%E8%A1%A8/">Vocabulary (词汇表)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/subword-%E5%AD%90%E8%AF%8D/">Subword (子词)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>分词</title><link>https://terms-en.ai-term-hub.com/zh/terms/tokenization/</link><pubDate>Sat, 18 Jul 2026 10:55:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/tokenization/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>分词是自然语言处理（NLP）中的关键预处理步骤，它将非结构化文本转换为适合模型输入的结构化数据。该过程涉及将句子分解为更小的单元，如单词、子词或字符，以便模型能够有效理解语义。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>分词是将原始文本拆分为称为词元的较小单元的过程，这些单元可以被机器学习算法处理。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本拆分&lt;/li>
&lt;li>预处理&lt;/li>
&lt;li>WordPiece&lt;/li>
&lt;li>字节对编码&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为BERT训练准备数据集&lt;/li>
&lt;li>GPT模型的输入格式化&lt;/li>
&lt;li>情感分析的数据清洗&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> AutoTokenizer
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokenizer &lt;span style="color:#f92672">=&lt;/span> AutoTokenizer&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;bert-base-uncased&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokens &lt;span style="color:#f92672">=&lt;/span> tokenizer&lt;span style="color:#f92672">.&lt;/span>tokenize(&lt;span style="color:#e6db74">&amp;#39;Hello world!&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenizer-%E5%88%86%E8%AF%8D%E5%99%A8/">Tokenizer (分词器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vocabulary-%E8%AF%8D%E6%B1%87%E8%A1%A8/">Vocabulary (词汇表)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5/">Embedding (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/preprocessing-%E9%A2%84%E5%A4%84%E7%90%86/">Preprocessing (预处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>语义</title><link>https://terms-en.ai-term-hub.com/zh/terms/semantic/</link><pubDate>Sat, 18 Jul 2026 10:54:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/semantic/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>AI中的语义分析侧重于理解输入的底层含义，而不仅仅是其表面模式。这涉及将单词或符号映射到概念，捕捉关系等。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>与语言或数据中的意义相关，区别于句法结构或形式。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>意义表示&lt;/li>
&lt;li>向量嵌入&lt;/li>
&lt;li>上下文理解&lt;/li>
&lt;li>意图识别&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>语义搜索引擎&lt;/li>
&lt;li>情感分析&lt;/li>
&lt;li>知识图谱构建&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">NLP (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embeddings-%E5%B5%8C%E5%85%A5%E5%90%91%E9%87%8F/">Embeddings (嵌入向量)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ontology-%E6%9C%AC%E4%BD%93%E8%AE%BA/">Ontology (本体论)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/semantics-%E8%AF%AD%E4%B9%89%E5%AD%A6/">Semantics (语义学)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>提示词</title><link>https://terms-en.ai-term-hub.com/zh/terms/prompt/</link><pubDate>Sat, 18 Jul 2026 10:54:02 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/prompt/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>提示词是与大型语言模型及其他生成式AI系统进行交互的主要接口。它定义了模型输出的上下文、语气和约束。有效的提示……&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>提供给生成式AI模型的输入文本或指令，以引发特定的响应或行为。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>指令微调&lt;/li>
&lt;li>上下文窗口&lt;/li>
&lt;li>少样本学习&lt;/li>
&lt;li>分词&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>生成创意写作内容&lt;/li>
&lt;li>代码补全与调试&lt;/li>
&lt;li>客户服务聊天机器人交互&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/inference-%E6%8E%A8%E7%90%86/">Inference (推理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/llm-%E5%A4%A7%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">LLM (大语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/fine-tuning-%E5%BE%AE%E8%B0%83/">Fine-tuning (微调)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/context-%E4%B8%8A%E4%B8%8B%E6%96%87/">Context (上下文)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>帖子</title><link>https://terms-en.ai-term-hub.com/zh/terms/post/</link><pubDate>Sat, 18 Jul 2026 10:53:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/post/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在数字通信和AI数据语境中，“帖子”指在线分享的离散内容单元。它是训练自然语言处理模型、情感分析以及理解用户交互模式的主要数据来源。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>发布的内容片段，通常位于博客、社交媒体平台或论坛上，代表用户生成的信息或评论。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>用户生成内容&lt;/li>
&lt;li>NLP数据源&lt;/li>
&lt;li>社交媒体&lt;/li>
&lt;li>情感分析&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>基于对话数据训练聊天机器人&lt;/li>
&lt;li>通过情感分析分析公众舆论&lt;/li>
&lt;li>检测社交信息流中的垃圾邮件或假新闻&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/content-%E5%86%85%E5%AE%B9/">Content (内容)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/blog-%E5%8D%9A%E5%AE%A2/">Blog (博客)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/forum-%E8%AE%BA%E5%9D%9B/">Forum (论坛)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dataset-%E6%95%B0%E6%8D%AE%E9%9B%86/">Dataset (数据集)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>预训练</title><link>https://terms-en.ai-term-hub.com/zh/terms/pre_training/</link><pubDate>Sat, 18 Jul 2026 10:53:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/pre_training/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>预训练是深度学习中的一种基础技术，模型从海量数据中学习广泛的特征和模式，通常无需标签。这一过程使模型能够发展出通用的知识表示，从而在后续针对特定下游任务进行微调时，仅需少量数据即可达到高性能。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>在大型未标记数据集上训练机器学习模型的初始阶段，以便在针对特定任务进行微调之前学习通用表示。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>迁移学习&lt;/li>
&lt;li>特征提取&lt;/li>
&lt;li>大规模数据&lt;/li>
&lt;li>微调&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练BERT或GPT等语言模型&lt;/li>
&lt;li>使用ImageNet权重初始化卷积神经网络（CNN）&lt;/li>
&lt;li>构建多模态AI的基础模型&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> BertModel
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>model &lt;span style="color:#f92672">=&lt;/span> BertModel&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;bert-base-uncased&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Model is now pre-trained and ready for fine-tuning&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/fine-tuning-%E5%BE%AE%E8%B0%83/">Fine-tuning (微调)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/foundation-model-%E5%9F%BA%E7%A1%80%E6%A8%A1%E5%9E%8B/">Foundation Model (基础模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/unsupervised-learning-%E6%97%A0%E7%9B%91%E7%9D%A3%E5%AD%A6%E4%B9%A0/">Unsupervised Learning (无监督学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transfer-learning-%E8%BF%81%E7%A7%BB%E5%AD%A6%E4%B9%A0/">Transfer Learning (迁移学习)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>自然语言处理</title><link>https://terms-en.ai-term-hub.com/zh/terms/natural_language_processing/</link><pubDate>Sat, 18 Jul 2026 10:53:26 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/natural_language_processing/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>自然语言处理（NLP）是人工智能的一个子领域，它将计算语言学与统计、机器学习和深度学习模型相结合。它使机器能够&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>人工智能的一个分支，专注于使计算机能够理解、解释和生成人类语言。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>分词&lt;/li>
&lt;li>情感分析&lt;/li>
&lt;li>命名实体识别&lt;/li>
&lt;li>机器翻译&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>聊天机器人和虚拟助手&lt;/li>
&lt;li>自动化客户支持&lt;/li>
&lt;li>语言翻译服务&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> spacy
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>nlp &lt;span style="color:#f92672">=&lt;/span> spacy&lt;span style="color:#f92672">.&lt;/span>load(&lt;span style="color:#e6db74">&amp;#39;en_core_web_sm&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>doc &lt;span style="color:#f92672">=&lt;/span> nlp(&lt;span style="color:#e6db74">&amp;#39;Apple is looking at buying U.K. startup for $1 billion&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">for&lt;/span> ent &lt;span style="color:#f92672">in&lt;/span> doc&lt;span style="color:#f92672">.&lt;/span>ents:
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> print(ent&lt;span style="color:#f92672">.&lt;/span>text, ent&lt;span style="color:#f92672">.&lt;/span>label_)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/computational_linguistics-%E8%AE%A1%E7%AE%97%E8%AF%AD%E8%A8%80%E5%AD%A6/">computational_linguistics (计算语言学)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/machine_learning-%E6%9C%BA%E5%99%A8%E5%AD%A6%E4%B9%A0/">machine_learning (机器学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/deep_learning-%E6%B7%B1%E5%BA%A6%E5%AD%A6%E4%B9%A0/">deep_learning (深度学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/text_mining-%E6%96%87%E6%9C%AC%E6%8C%96%E6%8E%98/">text_mining (文本挖掘)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>多头注意力</title><link>https://terms-en.ai-term-hub.com/zh/terms/multi_head_attention/</link><pubDate>Sat, 18 Jul 2026 10:53:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/multi_head_attention/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>多头注意力通过并行运行多次标准注意力机制（使用不同的学习到的线性投影）来扩展标准注意力机制。这使得模型能够联合关注来自不同位置的不同表示子空间的信息。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Transformer模型中的一种机制，允许模型同时关注来自不同表示子空间的信息。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自注意力&lt;/li>
&lt;li>线性投影&lt;/li>
&lt;li>拼接&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自然语言处理 (NLP)&lt;/li>
&lt;li>机器翻译&lt;/li>
&lt;li>使用Vision Transformer进行图像分类&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch.nn &lt;span style="color:#66d9ef">as&lt;/span> nn
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">class&lt;/span> &lt;span style="color:#a6e22e">MultiHeadAttention&lt;/span>(nn&lt;span style="color:#f92672">.&lt;/span>Module):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">def&lt;/span> __init__(self, d_model, num_heads):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> super()&lt;span style="color:#f92672">.&lt;/span>__init__()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> self&lt;span style="color:#f92672">.&lt;/span>num_heads &lt;span style="color:#f92672">=&lt;/span> num_heads
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> self&lt;span style="color:#f92672">.&lt;/span>d_k &lt;span style="color:#f92672">=&lt;/span> d_model &lt;span style="color:#f92672">//&lt;/span> num_heads
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> self&lt;span style="color:#f92672">.&lt;/span>W_q &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>Linear(d_model, d_model)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> self&lt;span style="color:#f92672">.&lt;/span>W_k &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>Linear(d_model, d_model)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> self&lt;span style="color:#f92672">.&lt;/span>W_v &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>Linear(d_model, d_model)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> self&lt;span style="color:#f92672">.&lt;/span>W_o &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>Linear(d_model, d_model)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">forward&lt;/span>(self, x):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#75715e"># Simplified forward pass logic&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">pass&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/scaled-dot-product-attention-%E7%BC%A9%E6%94%BE%E7%82%B9%E7%A7%AF%E6%B3%A8%E6%84%8F%E5%8A%9B/">Scaled Dot-Product Attention (缩放点积注意力)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E5%8F%98%E6%8D%A2%E5%99%A8/">Transformer (变换器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5/">Embedding (嵌入)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>长上下文</title><link>https://terms-en.ai-term-hub.com/zh/terms/long/</link><pubDate>Sat, 18 Jul 2026 10:52:50 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/long/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能语境中，“长”通常描述处理大量输入的能力，如长文档或冗长的视频流。对于大语言模型而言，这涉及管理长上下文窗口，以保持对完整输入的理解和一致性。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>通常指扩展的数据序列，例如自然语言处理模型中的长上下文窗口。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>上下文窗口&lt;/li>
&lt;li>序列长度&lt;/li>
&lt;li>注意力机制&lt;/li>
&lt;li>内存管理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>文档摘要&lt;/li>
&lt;li>长篇幅内容生成&lt;/li>
&lt;li>代码库分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/context-window-%E4%B8%8A%E4%B8%8B%E6%96%87%E7%AA%97%E5%8F%A3/">Context Window (上下文窗口)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-transformer%E6%9E%B6%E6%9E%84/">Transformer (Transformer架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rnn-%E5%BE%AA%E7%8E%AF%E7%A5%9E%E7%BB%8F%E7%BD%91%E7%BB%9C/">RNN (循环神经网络)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/attention-%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6/">Attention (注意力机制)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>大语言模型</title><link>https://terms-en.ai-term-hub.com/zh/terms/llm/</link><pubDate>Sat, 18 Jul 2026 10:52:16 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/llm/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>大语言模型（LLM）是基于Transformer架构的高级人工智能系统，在包含大量文本和代码的数据集上进行训练。它们学习语言中的统计模式，&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种在海量文本语料库上训练的深度学习模型，用于理解和生成类人语言。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>Transformer架构&lt;/li>
&lt;li>词元预测&lt;/li>
&lt;li>预训练&lt;/li>
&lt;li>缩放定律&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>聊天机器人和虚拟助手&lt;/li>
&lt;li>内容生成&lt;/li>
&lt;li>代码补全&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gpt-%E7%94%9F%E6%88%90%E5%BC%8F%E9%A2%84%E8%AE%AD%E7%BB%83transformer/">GPT (生成式预训练Transformer)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%9F%BA%E4%BA%8Etransformer%E7%9A%84%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BA/">BERT (基于Transformer的双向编码器表示)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86-nlp/">自然语言处理 (NLP)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>代替 / 反而</title><link>https://terms-en.ai-term-hub.com/zh/terms/instead/</link><pubDate>Sat, 18 Jul 2026 10:52:05 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/instead/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>虽然 &amp;lsquo;instead&amp;rsquo; 不是一个技术性的 AI 算法术语，但在提示工程（Prompt Engineering）和自然语言理解中至关重要。它指示子句之间的对比或替代关系。在大型语言模型（LLM）训练（注：原文截断，此处补全语义）中，理解此类逻辑连接词有助于模型准确捕捉用户意图中的否定或转折约束。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Instead 是一个连接词或副词，表示替代、替换，或在另一项行动之外采取的替代行动。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>替代&lt;/li>
&lt;li>话语标记&lt;/li>
&lt;li>负向约束&lt;/li>
&lt;li>语言逻辑&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>用于生成替代输出的提示工程&lt;/li>
&lt;li>自然语言理解中的意图识别&lt;/li>
&lt;li>逻辑推理任务&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/alternative-%E6%9B%BF%E4%BB%A3%E6%96%B9%E6%A1%88/">alternative (替代方案)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/substitution-%E6%9B%BF%E6%8D%A2/">substitution (替换)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/constraint-%E7%BA%A6%E6%9D%9F/">constraint (约束)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompting-%E6%8F%90%E7%A4%BA%E5%B7%A5%E7%A8%8B/">prompting (提示工程)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>分层</title><link>https://terms-en.ai-term-hub.com/zh/terms/hierarchical/</link><pubDate>Sat, 18 Jul 2026 10:51:52 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/hierarchical/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>分层AI系统将信息或控制组织成嵌套层的树状结构。在强化学习中，分层强化学习（Hierarchical RL）将复杂任务分解为由高层管理的子目标&amp;hellip;&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>指组织为多个抽象级别的AI架构或学习策略，高级别控制低级别。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>抽象层级&lt;/li>
&lt;li>子目标设定&lt;/li>
&lt;li>特征提取&lt;/li>
&lt;li>模块化&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>复杂机器人任务分解&lt;/li>
&lt;li>深度神经网络特征学习&lt;/li>
&lt;li>带有语法树的自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/reinforcement-learning-%E5%BC%BA%E5%8C%96%E5%AD%A6%E4%B9%A0/">Reinforcement Learning (强化学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/abstraction-%E6%8A%BD%E8%B1%A1/">Abstraction (抽象)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multi-agent-systems-%E5%A4%9A%E6%99%BA%E8%83%BD%E4%BD%93%E7%B3%BB%E7%BB%9F/">Multi-agent Systems (多智能体系统)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tree-search-%E6%A0%91%E6%90%9C%E7%B4%A2/">Tree Search (树搜索)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>生成</title><link>https://terms-en.ai-term-hub.com/zh/terms/generation/</link><pubDate>Sat, 18 Jul 2026 10:51:41 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/generation/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能中，生成是指模型（特别是生成对抗网络 GAN 和基于 Transformer 的大语言模型）产生文本、图像等新颖内容的能力。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>生成式模型创建与训练分布相似的新数据实例的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>概率建模&lt;/li>
&lt;li>潜在空间&lt;/li>
&lt;li>词元预测&lt;/li>
&lt;li>合成数据&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自然语言生成&lt;/li>
&lt;li>图像合成&lt;/li>
&lt;li>代码自动补全&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/diffusion-models-%E6%89%A9%E6%95%A3%E6%A8%A1%E5%9E%8B/">Diffusion Models (扩散模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-transformer%E6%9E%B6%E6%9E%84/">Transformer (Transformer架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gan-%E7%94%9F%E6%88%90%E5%AF%B9%E6%8A%97%E7%BD%91%E7%BB%9C/">GAN (生成对抗网络)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompt-engineering-%E6%8F%90%E7%A4%BA%E5%B7%A5%E7%A8%8B/">Prompt Engineering (提示工程)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>上下文</title><link>https://terms-en.ai-term-hub.com/zh/terms/context/</link><pubDate>Sat, 18 Jul 2026 10:49:42 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/context/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在自然语言处理中，上下文对于消除歧义至关重要，例如根据前文理解代词或习语。现代架构（如Transformer）利用注意力机制来捕捉长距离依赖关系，从而更好地处理上下文信息。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>上下文是指帮助AI模型准确解释输入数据并生成相关响应的周围信息或环境。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>注意力机制&lt;/li>
&lt;li>序列建模&lt;/li>
&lt;li>歧义消解&lt;/li>
&lt;li>窗口大小&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>对话式聊天机器人&lt;/li>
&lt;li>文档摘要&lt;/li>
&lt;li>代码补全助手&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompt-engineering-%E6%8F%90%E7%A4%BA%E5%B7%A5%E7%A8%8B/">Prompt Engineering (提示工程)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/memory-%E8%AE%B0%E5%BF%86/">Memory (记忆)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/attention-%E6%B3%A8%E6%84%8F%E5%8A%9B/">Attention (注意力)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sequence-%E5%BA%8F%E5%88%97/">Sequence (序列)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>嵌入</title><link>https://terms-en.ai-term-hub.com/zh/terms/embedding/</link><pubDate>Sat, 18 Jul 2026 07:44:46 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/embedding/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>嵌入是数据的稠密向量表示，其中语义关系在几何空间中得以保留。通过将分类或高维输入转换为固定长度的向量，模型能够捕捉数据背后的深层含义。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种将单词或图像等离散对象映射到连续向量空间的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>向量空间&lt;/li>
&lt;li>语义相似性&lt;/li>
&lt;li>降维&lt;/li>
&lt;li>连续表示&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>情感分析等自然语言处理任务&lt;/li>
&lt;li>用于用户-物品匹配推荐系统&lt;/li>
&lt;li>基于视觉相似性的图像检索&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> numpy &lt;span style="color:#66d9ef">as&lt;/span> np
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Simulating a simple embedding lookup&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>embeddings &lt;span style="color:#f92672">=&lt;/span> np&lt;span style="color:#f92672">.&lt;/span>random&lt;span style="color:#f92672">.&lt;/span>rand(&lt;span style="color:#ae81ff">100&lt;/span>, &lt;span style="color:#ae81ff">128&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>word_index &lt;span style="color:#f92672">=&lt;/span> &lt;span style="color:#ae81ff">5&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>vector &lt;span style="color:#f92672">=&lt;/span> embeddings[word_index]
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/word2vec-%E8%AF%8D%E5%90%91%E9%87%8F%E6%A8%A1%E5%9E%8B/">Word2Vec (词向量模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E5%8F%98%E6%8D%A2%E5%99%A8%E6%9E%B6%E6%9E%84/">Transformer (变换器架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%BD%9C%E5%9C%A8%E7%A9%BA%E9%97%B4-latent-space/">潜在空间 (Latent Space)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%90%91%E9%87%8F%E6%95%B0%E6%8D%AE%E5%BA%93/">向量数据库&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>注意力机制</title><link>https://terms-en.ai-term-hub.com/zh/terms/attention_mechanism/</link><pubDate>Sat, 18 Jul 2026 07:44:10 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/attention_mechanism/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>注意力机制使模型能够动态地权衡输入序列中不同元素的重要性。与平等对待所有输入数据不同，它会根据上下文分配不同的权重，从而让模型聚焦于最相关的信息部分，显著提升了对长序列数据的处理能力。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种使神经网络在生成输出时能够专注于输入数据特定部分的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自注意力&lt;/li>
&lt;li>上下文向量&lt;/li>
&lt;li>加权求和&lt;/li>
&lt;li>Transformer 架构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>机器翻译模型&lt;/li>
&lt;li>图像描述生成&lt;/li>
&lt;li>文本摘要&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-transformer-%E6%A8%A1%E5%9E%8B/">Transformer (Transformer 模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multi-head-attention-%E5%A4%9A%E5%A4%B4%E6%B3%A8%E6%84%8F%E5%8A%9B/">Multi-Head Attention (多头注意力)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sequence-to-sequence-%E5%BA%8F%E5%88%97%E5%88%B0%E5%BA%8F%E5%88%97/">Sequence-to-Sequence (序列到序列)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>