<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>BERT on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/bert/</link><description>Recent content in BERT on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/bert/index.xml" rel="self" type="application/rss+xml"/><item><title>WordPiece</title><link>https://terms-en.ai-term-hub.com/zh/terms/wordpiece/</link><pubDate>Sat, 18 Jul 2026 11:38:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/wordpiece/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>WordPiece 是一种广泛应用于 BERT 和 ALBERT 等自然语言处理模型的分词方法。它将单词分解为更小的子词单元，以应对形态学丰富性并减少词汇表大小，从而更好地处理未见过的单词。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种子词分词算法，通过递归合并最频繁出现的字符对来处理未登录词（OOV）。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>子词分词&lt;/li>
&lt;li>词汇扩展&lt;/li>
&lt;li>未登录词处理&lt;/li>
&lt;li>形态分析&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为 BERT 模型预处理文本&lt;/li>
&lt;li>处理低资源语言&lt;/li>
&lt;li>减小嵌入矩阵大小&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> BertTokenizer
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokenizer &lt;span style="color:#f92672">=&lt;/span> BertTokenizer&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;bert-base-uncased&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokens &lt;span style="color:#f92672">=&lt;/span> tokenizer&lt;span style="color:#f92672">.&lt;/span>tokenize(&lt;span style="color:#e6db74">&amp;#39;unhappiness&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(tokens)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AD%97%E8%8A%82%E5%AF%B9%E7%BC%96%E7%A0%81-byte-pair-encoding/">字节对编码 (Byte-Pair Encoding)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentencepiece/">SentencePiece&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%88%86%E8%AF%8D-tokenization/">分词 (Tokenization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E9%A2%84%E5%A4%84%E7%90%86-nlp-preprocessing/">NLP 预处理 (NLP preprocessing)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>