<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Tokenization on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/tokenization/</link><description>Recent content in Tokenization on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/tokenization/index.xml" rel="self" type="application/rss+xml"/><item><title>WordPiece</title><link>https://terms-en.ai-term-hub.com/zh/terms/wordpiece/</link><pubDate>Sat, 18 Jul 2026 11:38:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/wordpiece/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>WordPiece 是一种广泛应用于 BERT 和 ALBERT 等自然语言处理模型的分词方法。它将单词分解为更小的子词单元，以应对形态学丰富性并减少词汇表大小，从而更好地处理未见过的单词。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种子词分词算法，通过递归合并最频繁出现的字符对来处理未登录词（OOV）。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>子词分词&lt;/li>
&lt;li>词汇扩展&lt;/li>
&lt;li>未登录词处理&lt;/li>
&lt;li>形态分析&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为 BERT 模型预处理文本&lt;/li>
&lt;li>处理低资源语言&lt;/li>
&lt;li>减小嵌入矩阵大小&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> BertTokenizer
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokenizer &lt;span style="color:#f92672">=&lt;/span> BertTokenizer&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;bert-base-uncased&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokens &lt;span style="color:#f92672">=&lt;/span> tokenizer&lt;span style="color:#f92672">.&lt;/span>tokenize(&lt;span style="color:#e6db74">&amp;#39;unhappiness&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(tokens)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AD%97%E8%8A%82%E5%AF%B9%E7%BC%96%E7%A0%81-byte-pair-encoding/">字节对编码 (Byte-Pair Encoding)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentencepiece/">SentencePiece&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%88%86%E8%AF%8D-tokenization/">分词 (Tokenization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E9%A2%84%E5%A4%84%E7%90%86-nlp-preprocessing/">NLP 预处理 (NLP preprocessing)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>SentencePiece</title><link>https://terms-en.ai-term-hub.com/zh/terms/sentencepiece/</link><pubDate>Sat, 18 Jul 2026 11:33:07 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/sentencepiece/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>SentencePiece 是一个流行的开源文本归一化和分词库，广泛用于现代自然语言处理管道中。它执行联合词块和子词词汇表的无监督学习。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种无监督文本分词器和去分词器库，将原始文本视为子词序列用于自然语言处理预处理。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>子词分词&lt;/li>
&lt;li>词汇表学习&lt;/li>
&lt;li>去分词&lt;/li>
&lt;li>语言无关&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为变换器模型预处理数据&lt;/li>
&lt;li>处理多语言文本语料库&lt;/li>
&lt;li>减少语言模型的词汇量大小&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenizer-%E5%88%86%E8%AF%8D%E5%99%A8/">Tokenizer (分词器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bpe-%E5%AD%97%E8%8A%82%E5%AF%B9%E7%BC%96%E7%A0%81/">BPE (字节对编码)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/byte-pair-encoding-%E5%AD%97%E8%8A%82%E5%AF%B9%E7%BC%96%E7%A0%81/">Byte-Pair Encoding (字节对编码)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-preprocessing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86%E9%A2%84%E5%A4%84%E7%90%86/">NLP Preprocessing (自然语言处理预处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>字节对编码 (BPE)</title><link>https://terms-en.ai-term-hub.com/zh/terms/bpe/</link><pubDate>Sat, 18 Jul 2026 10:59:27 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/bpe/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>字节对编码（BPE）是一种数据压缩技术，经过调整后应用于自然语言处理中，以处理未登录词（Out-of-Vocabulary）。它从单个字符的词汇表开始，并迭代地合并最频繁出现的字符对，直到达到预定的词汇表大小或收敛。这种方法允许模型将罕见词分解为更常见的子词单元，从而提高对未知词汇的处理能力。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>字节对编码是一种用于子词分词的算法，它通过迭代合并出现频率最高的字符对来构建词汇表。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>子词分词&lt;/li>
&lt;li>词汇表合并&lt;/li>
&lt;li>频率分析&lt;/li>
&lt;li>未登录词处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为大语言模型预处理文本&lt;/li>
&lt;li>处理形态丰富的语言&lt;/li>
&lt;li>减少神经网络中的词汇表规模&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> tiktoken
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>enc &lt;span style="color:#f92672">=&lt;/span> tiktoken&lt;span style="color:#f92672">.&lt;/span>get_encoding(&lt;span style="color:#e6db74">&amp;#34;cl100k_base&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokens &lt;span style="color:#f92672">=&lt;/span> enc&lt;span style="color:#f92672">.&lt;/span>encode(&lt;span style="color:#e6db74">&amp;#34;unhappiness&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(tokens)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wordpiece-wordpiece%E5%88%86%E8%AF%8D%E7%AE%97%E6%B3%95/">WordPiece (WordPiece分词算法)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sentencepiece-sentencepiece%E5%88%86%E8%AF%8D%E5%B7%A5%E5%85%B7%E5%BA%93/">SentencePiece (SentencePiece分词工具库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenization-%E5%88%86%E8%AF%8D/">Tokenization (分词)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/subword-units-%E5%AD%90%E8%AF%8D%E5%8D%95%E5%85%83/">Subword Units (子词单元)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>