<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Transformers on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/transformers/</link><description>Recent content in Transformers on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/transformers/index.xml" rel="self" type="application/rss+xml"/><item><title>XLM-RoBERTa</title><link>https://terms-en.ai-term-hub.com/zh/terms/xlm_roberta/</link><pubDate>Sat, 18 Jul 2026 11:38:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/xlm_roberta/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>XLM-RoBERTa（跨语言语言模型 RoBERTa）是由 Meta AI 开发的大规模多语言模型。它通过在一个涵盖 100 多种语言的多样化数据集上进行预训练，扩展了 RoBERTa 的架构，从而实现了强大的跨语言理解和生成能力，&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种基于 RoBERTa 的多语言 Transformer 模型，在超过 100 种语言的庞大文本数据集上进行预训练。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>多语言自然语言处理&lt;/li>
&lt;li>Transformer 架构&lt;/li>
&lt;li>跨语言迁移&lt;/li>
&lt;li>预训练&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>多语言情感分析&lt;/li>
&lt;li>跨语言问答&lt;/li>
&lt;li>机器翻译&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BAtransformer/">BERT (双向编码器表示Transformer)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/roberta-%E4%BC%98%E5%8C%96%E7%89%88%E7%9A%84bert/">RoBERTa (优化版的BERT)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/m-bert-%E5%A4%9A%E8%AF%AD%E8%A8%80bert/">M-BERT (多语言BERT)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentence-transformers-%E5%8F%A5%E5%AD%90transformer/">Sentence Transformers (句子Transformer)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>前缀微调</title><link>https://terms-en.ai-term-hub.com/zh/terms/prefix_tuning/</link><pubDate>Sat, 18 Jul 2026 11:30:18 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/prefix_tuning/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>前缀微调是一种用于预训练变换器的参数高效适应技术。与更新所有模型权重不同，它在输入序列前prepend（前置）一系列可训练的连续向量（即前缀），从而冻结主干网络并仅优化少量参数以适应新任务。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种参数高效的微调方法，通过在变换器层的输入端添加可训练的连续向量来实现适配。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>参数高效微调&lt;/li>
&lt;li>软提示&lt;/li>
&lt;li>变换器层&lt;/li>
&lt;li>冻结主干&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>少样本学习适配&lt;/li>
&lt;li>资源受限下的多任务学习&lt;/li>
&lt;li>为利基领域定制大型语言模型&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompt_tuning-%E6%8F%90%E7%A4%BA%E5%BE%AE%E8%B0%83/">prompt_tuning (提示微调)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/p_tuning-p-tuning/">p_tuning (P-Tuning)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/adapter_modules-%E9%80%82%E9%85%8D%E5%99%A8%E6%A8%A1%E5%9D%97/">adapter_modules (适配器模块)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/peft-%E5%8F%82%E6%95%B0%E9%AB%98%E6%95%88%E5%BE%AE%E8%B0%83%E6%8A%80%E6%9C%AF/">peft (参数高效微调技术)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>长上下文</title><link>https://terms-en.ai-term-hub.com/zh/terms/long_context/</link><pubDate>Sat, 18 Jul 2026 11:24:46 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/long_context/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>长上下文指的是基于 Transformer 的模型处理极长输入长度的能力，通常超过标准的 2k 或 4k 标记限制。这种能力使模型能够分析完整的文档、代码库或长文本，保持全局一致性和细节记忆。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>语言模型处理并保留包含数千或数百万个标记（token）的输入序列信息的能力。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>上下文窗口&lt;/li>
&lt;li>标记限制&lt;/li>
&lt;li>注意力机制&lt;/li>
&lt;li>位置编码&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>总结完整的法律合同&lt;/li>
&lt;li>分析完整的源代码仓库&lt;/li>
&lt;li>处理长篇音频转录文本&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/context-window-%E4%B8%8A%E4%B8%8B%E6%96%87%E7%AA%97%E5%8F%A3/">Context Window (上下文窗口)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-architecture-transformer-%E6%9E%B6%E6%9E%84/">Transformer Architecture (Transformer 架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rope-rotary-positional-embeddings-%E6%97%8B%E8%BD%AC%E4%BD%8D%E7%BD%AE%E7%BC%96%E7%A0%81/">RoPE (Rotary Positional Embeddings, 旋转位置编码)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/kv-cache-%E9%94%AE%E5%80%BC%E7%BC%93%E5%AD%98/">KV Cache (键值缓存)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>掩码填充</title><link>https://terms-en.ai-term-hub.com/zh/terms/fill_mask/</link><pubDate>Sat, 18 Jul 2026 11:17:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/fill_mask/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>掩码填充是 BERT 等基于 Transformer 模型中使用的一种基本预训练目标。该过程涉及掩盖文本序列中的随机标记，并训练模型预测被掩盖的原始标记，从而学习语言的深层语义表示。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种自然语言处理任务，模型根据上下文预测句子中缺失的标记。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>掩码语言建模&lt;/li>
&lt;li>上下文理解&lt;/li>
&lt;li>自监督学习&lt;/li>
&lt;li>标记预测&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>文本补全&lt;/li>
&lt;li>语义角色标注&lt;/li>
&lt;li>预训练基础&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-%E5%8F%8C%E5%90%91%E7%BC%96%E7%A0%81%E5%99%A8%E8%A1%A8%E7%A4%BA/">BERT (双向编码器表示)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/masked-language-model-%E6%8E%A9%E7%A0%81%E8%AF%AD%E8%A8%80%E6%A8%A1%E5%9E%8B/">Masked Language Model (掩码语言模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E8%BD%AC%E6%8D%A2%E5%99%A8%E6%9E%B6%E6%9E%84/">Transformer (转换器架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenization-%E5%88%86%E8%AF%8D/">Tokenization (分词)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>ExBERT</title><link>https://terms-en.ai-term-hub.com/zh/terms/exbert/</link><pubDate>Sat, 18 Jul 2026 11:16:37 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/exbert/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>ExBERT通过分析不同层中单个注意力头的重要性，为BERT Transformer模型提供可解释性。它使用基于梯度的归因或其他技术来量化每个组件对最终预测的贡献，从而帮助理解模型的决策过程。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种通过识别对特定输出贡献最大的注意力头和层来解释BERT预测的方法。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>注意力头&lt;/li>
&lt;li>模型可解释性&lt;/li>
&lt;li>梯度归因&lt;/li>
&lt;li>Transformer架构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>调试NLP模型&lt;/li>
&lt;li>理解句法与语义处理&lt;/li>
&lt;li>医疗文本分析中的可解释AI&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bert-bert%E6%A8%A1%E5%9E%8B/">bert (BERT模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/attention_mechanism-%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6/">attention_mechanism (注意力机制)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/explainable_ai-%E5%8F%AF%E8%A7%A3%E9%87%8A%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD/">explainable_ai (可解释人工智能)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer_models-transformer%E6%A8%A1%E5%9E%8B/">transformer_models (Transformer模型)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>位置编码</title><link>https://terms-en.ai-term-hub.com/zh/terms/positional_encoding/</link><pubDate>Sat, 18 Jul 2026 11:01:29 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/positional_encoding/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>由于 Transformer 并行处理所有令牌，而不是像循环神经网络（RNN）那样按顺序处理，因此它缺乏对令牌顺序的固有认知。位置编码通过向输入嵌入添加特定的向量来保留这种顺序信息。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种将序列中令牌的相对或绝对位置信息注入 Transformer 模型的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>序列顺序&lt;/li>
&lt;li>自注意力机制&lt;/li>
&lt;li>正弦函数&lt;/li>
&lt;li>令牌嵌入&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>机器翻译&lt;/li>
&lt;li>文本摘要&lt;/li>
&lt;li>语言建模&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> math
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">get_positional_encoding&lt;/span>(seq_len, d_model):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pe &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>zeros(seq_len, d_model)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> position &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>arange(&lt;span style="color:#ae81ff">0&lt;/span>, seq_len, dtype&lt;span style="color:#f92672">=&lt;/span>torch&lt;span style="color:#f92672">.&lt;/span>float)&lt;span style="color:#f92672">.&lt;/span>unsqueeze(&lt;span style="color:#ae81ff">1&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> div_term &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>exp(torch&lt;span style="color:#f92672">.&lt;/span>arange(&lt;span style="color:#ae81ff">0&lt;/span>, d_model, &lt;span style="color:#ae81ff">2&lt;/span>)&lt;span style="color:#f92672">.&lt;/span>float() &lt;span style="color:#f92672">*&lt;/span> (&lt;span style="color:#f92672">-&lt;/span>math&lt;span style="color:#f92672">.&lt;/span>log(&lt;span style="color:#ae81ff">10000.0&lt;/span>) &lt;span style="color:#f92672">/&lt;/span> d_model))
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pe[:, &lt;span style="color:#ae81ff">0&lt;/span>::&lt;span style="color:#ae81ff">2&lt;/span>] &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>sin(position &lt;span style="color:#f92672">*&lt;/span> div_term)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> pe[:, &lt;span style="color:#ae81ff">1&lt;/span>::&lt;span style="color:#ae81ff">2&lt;/span>] &lt;span style="color:#f92672">=&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>cos(position &lt;span style="color:#f92672">*&lt;/span> div_term)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">return&lt;/span> pe&lt;span style="color:#f92672">.&lt;/span>unsqueeze(&lt;span style="color:#ae81ff">0&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E6%9E%B6%E6%9E%84/">Transformer 架构&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%B5%8C%E5%85%A5-embedding/">嵌入 (Embedding)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%B3%A8%E6%84%8F%E5%8A%9B%E6%9C%BA%E5%88%B6-attention-mechanism/">注意力机制 (Attention Mechanism)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%97%8B%E8%BD%AC%E4%BD%8D%E7%BD%AE%E7%BC%96%E7%A0%81-rope/">旋转位置编码 (RoPE)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Encoder</title><link>https://terms-en.ai-term-hub.com/zh/terms/encoder/</link><pubDate>Sat, 18 Jul 2026 10:59:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/encoder/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>编码器处理原始输入序列或数据结构，并将它们转换为潜在空间表示，通常称为嵌入或代码。它们是 Transformer 和自编码器等架构的核心部分。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>编码器是神经网络的组件，它将输入数据转换为压缩的、有意义的表示形式。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>特征提取&lt;/li>
&lt;li>潜在空间&lt;/li>
&lt;li>序列处理&lt;/li>
&lt;li>压缩&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>在 Transformer 模型中处理输入文本&lt;/li>
&lt;li>在去噪自编码器中压缩图像&lt;/li>
&lt;li>从评论中提取情感特征&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch.nn &lt;span style="color:#66d9ef">as&lt;/span> nn
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#66d9ef">class&lt;/span> &lt;span style="color:#a6e22e">SimpleEncoder&lt;/span>(nn&lt;span style="color:#f92672">.&lt;/span>Module):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">def&lt;/span> __init__(self, input_dim, hidden_dim):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> super()&lt;span style="color:#f92672">.&lt;/span>__init__()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> self&lt;span style="color:#f92672">.&lt;/span>fc &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>Linear(input_dim, hidden_dim)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> 
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">def&lt;/span> &lt;span style="color:#a6e22e">forward&lt;/span>(self, x):
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span> &lt;span style="color:#66d9ef">return&lt;/span> torch&lt;span style="color:#f92672">.&lt;/span>relu(self&lt;span style="color:#f92672">.&lt;/span>fc(x))
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/decoder-%E8%A7%A3%E7%A0%81%E5%99%A8/">Decoder (解码器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E8%BD%AC%E6%8D%A2%E5%99%A8%E6%9E%B6%E6%9E%84/">Transformer (转换器架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/autoencoder-%E8%87%AA%E7%BC%96%E7%A0%81%E5%99%A8/">Autoencoder (自编码器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/latent-variable-%E6%BD%9C%E5%8F%98%E9%87%8F/">Latent Variable (潜变量)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>适配器</title><link>https://terms-en.ai-term-hub.com/zh/terms/adapter/</link><pubDate>Sat, 18 Jul 2026 10:59:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/adapter/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>适配器是一种参数高效的微调技术，主要用于大型语言模型和Transformer架构中。与计算成本高昂的更新所有模型权重不同，适配器仅训练少量新增参数，从而保留预训练知识并降低资源消耗。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>插入预训练模型中的轻量级模块，用于针对特定下游任务进行高效微调。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>参数高效微调&lt;/li>
&lt;li>迁移学习&lt;/li>
&lt;li>模块化架构&lt;/li>
&lt;li>灾难性遗忘&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为客服聊天机器人微调大型语言模型&lt;/li>
&lt;li>适配视觉模型以进行医学图像分析&lt;/li>
&lt;li>高效部署多个领域特定的模型&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/lora-%E4%BD%8E%E7%A7%A9%E9%80%82%E5%BA%94/">LoRA (低秩适应)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/prompt-tuning-%E6%8F%90%E7%A4%BA%E5%BE%AE%E8%B0%83/">Prompt Tuning (提示微调)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/fine-tuning-%E5%BE%AE%E8%B0%83/">Fine-Tuning (微调)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-transformer%E6%9E%B6%E6%9E%84/">Transformer (Transformer架构)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>注意力机制</title><link>https://terms-en.ai-term-hub.com/zh/terms/attention/</link><pubDate>Sat, 18 Jul 2026 10:59:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/attention/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>注意力机制使模型在处理输入（特别是文本等序列数据）时能够关注相关信息。通过计算注意力分数，模型确定哪些元素对当前任务最相关，从而捕捉长距离依赖关系并增强上下文理解。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种允许神经网络动态权衡输入序列不同部分重要性的机制。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自注意力&lt;/li>
&lt;li>上下文加权&lt;/li>
&lt;li>长距离依赖&lt;/li>
&lt;li>Transformer架构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>跨语言机器翻译&lt;/li>
&lt;li>长文档摘要&lt;/li>
&lt;li>图像描述生成和视觉问答&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-transformer%E6%9E%B6%E6%9E%84/">Transformer (Transformer架构)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/self-attention-%E8%87%AA%E6%B3%A8%E6%84%8F%E5%8A%9B/">Self-Attention (自注意力)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multi-head-attention-%E5%A4%9A%E5%A4%B4%E6%B3%A8%E6%84%8F%E5%8A%9B/">Multi-Head Attention (多头注意力)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sequence-modeling-%E5%BA%8F%E5%88%97%E5%BB%BA%E6%A8%A1/">Sequence Modeling (序列建模)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>自注意力机制</title><link>https://terms-en.ai-term-hub.com/zh/terms/self_attention/</link><pubDate>Sat, 18 Jul 2026 10:54:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/self_attention/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>自注意力使模型能够同时捕捉序列中所有位置之间的依赖关系，无论距离远近。通过计算每对标记之间的注意力分数，它使得……&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种允许神经网络根据彼此之间的相对重要性来权衡输入序列不同部分的机制。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>查询-键-值 (Query-Key-Value)&lt;/li>
&lt;li>注意力分数&lt;/li>
&lt;li>上下文加权&lt;/li>
&lt;li>并行处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>机器翻译&lt;/li>
&lt;li>文本摘要&lt;/li>
&lt;li>通过视觉Transformer进行图像分类&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> torch.nn &lt;span style="color:#66d9ef">as&lt;/span> nn
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>attn &lt;span style="color:#f92672">=&lt;/span> nn&lt;span style="color:#f92672">.&lt;/span>MultiheadAttention(embed_dim&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">512&lt;/span>, num_heads&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">8&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transformer-%E5%8F%98%E6%8D%A2%E5%99%A8/">Transformer (变换器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/multi-head-attention-%E5%A4%9A%E5%A4%B4%E6%B3%A8%E6%84%8F%E5%8A%9B/">Multi-Head Attention (多头注意力)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embeddings-%E5%B5%8C%E5%85%A5%E5%90%91%E9%87%8F/">Embeddings (嵌入向量)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sequence-modeling-%E5%BA%8F%E5%88%97%E5%BB%BA%E6%A8%A1/">Sequence Modeling (序列建模)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>