<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Preprocessing on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/preprocessing/</link><description>Recent content in Preprocessing on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/preprocessing/index.xml" rel="self" type="application/rss+xml"/><item><title>量化</title><link>https://terms-en.ai-term-hub.com/zh/terms/quantification/</link><pubDate>Sat, 18 Jul 2026 11:31:12 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/quantification/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能和数据科学的背景下，量化指的是将非数值数据（如文本、图像或主观意见）转换为可测量的数值的过程。这一过程使计算机能够处理和理解原本难以量化的信息。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>量化是将定性属性或抽象概念以数值形式表达以便进行分析的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>数值表示&lt;/li>
&lt;li>特征工程&lt;/li>
&lt;li>数据转换&lt;/li>
&lt;li>测量&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>自然语言处理 (NLP)&lt;/li>
&lt;li>情感评分&lt;/li>
&lt;li>图像像素值分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%BC%96%E7%A0%81-encoding/">编码 (Encoding)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%BD%92%E4%B8%80%E5%8C%96-normalization/">归一化 (Normalization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%90%91%E9%87%8F%E5%8C%96-vectorization/">向量化 (Vectorization)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>实例选择</title><link>https://terms-en.ai-term-hub.com/zh/terms/instance_selection/</link><pubDate>Sat, 18 Jul 2026 11:22:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/instance_selection/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>实例选择旨在通过移除冗余或噪声数据点来提高计算效率和模型性能。与特征选择不同，它作用于数据集的行。其目标。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种预处理技术，通过选择代表性实例的子集来减小数据集的大小。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>数据缩减&lt;/li>
&lt;li>去噪&lt;/li>
&lt;li>代表性子集&lt;/li>
&lt;li>计算效率&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>大规模数据集预处理&lt;/li>
&lt;li>加速最近邻搜索&lt;/li>
&lt;li>清理不平衡数据集&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%95%B0%E6%8D%AE%E9%87%87%E6%A0%B7-data-sampling/">数据采样 (Data sampling)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%AC%A0%E9%87%87%E6%A0%B7-under-sampling/">欠采样 (Under-sampling)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%8E%8B%E7%BC%A9%E6%9C%80%E8%BF%91%E9%82%BB-condensed-nearest-neighbor/">压缩最近邻 (Condensed Nearest Neighbor)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>特征哈希</title><link>https://terms-en.ai-term-hub.com/zh/terms/feature_hashing/</link><pubDate>Sat, 18 Jul 2026 11:17:13 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/feature_hashing/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>特征哈希，也称为哈希技巧（hashing trick），允许机器学习模型处理大型稀疏特征空间，而无需维护特征与索引之间的显式映射。通过应用哈希函数，模型可以直接将任意特征映射到固定的向量维度中，从而节省内存并简化特征工程流程。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种利用哈希函数将高维稀疏特征映射到固定大小向量的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>哈希函数&lt;/li>
&lt;li>稀疏向量&lt;/li>
&lt;li>降维&lt;/li>
&lt;li>内存效率&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>具有大型词汇表的文本分类&lt;/li>
&lt;li>拥有海量物品集的推荐系统&lt;/li>
&lt;li>实时流数据处理&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> sklearn.feature_extraction &lt;span style="color:#f92672">import&lt;/span> FeatureHasher
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> numpy &lt;span style="color:#66d9ef">as&lt;/span> np
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#75715e"># Example: Hashing text features&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>hasher &lt;span style="color:#f92672">=&lt;/span> FeatureHasher(n_features&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">10&lt;/span>, input_type&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#e6db74">&amp;#39;string&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>docs &lt;span style="color:#f92672">=&lt;/span> [&lt;span style="color:#e6db74">&amp;#39;hello world&amp;#39;&lt;/span>, &lt;span style="color:#e6db74">&amp;#39;hello python&amp;#39;&lt;/span>, &lt;span style="color:#e6db74">&amp;#39;world python&amp;#39;&lt;/span>]
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>hashed &lt;span style="color:#f92672">=&lt;/span> hasher&lt;span style="color:#f92672">.&lt;/span>transform(docs)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(hashed&lt;span style="color:#f92672">.&lt;/span>toarray())
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/one-hot-encoding-%E7%8B%AC%E7%83%AD%E7%BC%96%E7%A0%81/">One-hot encoding (独热编码)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bag-of-words-%E8%AF%8D%E8%A2%8B%E6%A8%A1%E5%9E%8B/">Bag of Words (词袋模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/dimensionality-reduction-%E9%99%8D%E7%BB%B4/">Dimensionality reduction (降维)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/sparse-matrix-%E7%A8%80%E7%96%8F%E7%9F%A9%E9%98%B5/">Sparse matrix (稀疏矩阵)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>特征缩放</title><link>https://terms-en.ai-term-hub.com/zh/terms/feature_scaling/</link><pubDate>Sat, 18 Jul 2026 11:17:13 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/feature_scaling/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>特征缩放通过标准化输入变量的范围，防止量级较大的特征主导学习过程。常见方法包括归一化（最小-最大缩放）和标准化（Z-score标准化）。这一预处理步骤对于基于距离的算法和梯度下降优化至关重要，能加速收敛并提高模型性能。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>将数据的独立变量或特征的范围进行归一化，以确保量级一致性的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>归一化&lt;/li>
&lt;li>标准化&lt;/li>
&lt;li>梯度下降&lt;/li>
&lt;li>数据预处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>神经网络训练&lt;/li>
&lt;li>K-means聚类&lt;/li>
&lt;li>支持向量机 (SVM)&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> sklearn.preprocessing &lt;span style="color:#f92672">import&lt;/span> StandardScaler
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> numpy &lt;span style="color:#66d9ef">as&lt;/span> np
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>X &lt;span style="color:#f92672">=&lt;/span> np&lt;span style="color:#f92672">.&lt;/span>array([[&lt;span style="color:#ae81ff">1&lt;/span>, &lt;span style="color:#ae81ff">2&lt;/span>], [&lt;span style="color:#ae81ff">3&lt;/span>, &lt;span style="color:#ae81ff">4&lt;/span>], [&lt;span style="color:#ae81ff">5&lt;/span>, &lt;span style="color:#ae81ff">6&lt;/span>]])
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>scaler &lt;span style="color:#f92672">=&lt;/span> StandardScaler()
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>X_scaled &lt;span style="color:#f92672">=&lt;/span> scaler&lt;span style="color:#f92672">.&lt;/span>fit_transform(X)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(X_scaled)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/min-max-scaling-%E6%9C%80%E5%B0%8F-%E6%9C%80%E5%A4%A7%E7%BC%A9%E6%94%BE/">Min-Max Scaling (最小-最大缩放)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/z-score-normalization-z-score%E6%A0%87%E5%87%86%E5%8C%96/">Z-score Normalization (Z-score标准化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/data-preprocessing-%E6%95%B0%E6%8D%AE%E9%A2%84%E5%A4%84%E7%90%86/">Data preprocessing (数据预处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/gradient-descent-%E6%A2%AF%E5%BA%A6%E4%B8%8B%E9%99%8D/">Gradient Descent (梯度下降)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>特征工程</title><link>https://terms-en.ai-term-hub.com/zh/terms/feature_engineering/</link><pubDate>Sat, 18 Jul 2026 11:17:01 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/feature_engineering/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>特征工程是利用领域专业知识，将原始数据转换为更能代表底层模式的特征的艺术。该过程包括创建新变量、转换现有数据以及选择最具预测力的特征，从而提升算法效果。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>利用领域知识创建新特征或修改现有特征，以增强机器学习模型性能的做法。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>领域知识&lt;/li>
&lt;li>数据转换&lt;/li>
&lt;li>模型性能&lt;/li>
&lt;li>变量创建&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>提高回归模型准确性&lt;/li>
&lt;li>增强分类边界&lt;/li>
&lt;li>优化时间序列预测&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>df[&lt;span style="color:#e6db74">&amp;#39;new_feature&amp;#39;&lt;/span>] &lt;span style="color:#f92672">=&lt;/span> df[&lt;span style="color:#e6db74">&amp;#39;feature_a&amp;#39;&lt;/span>] &lt;span style="color:#f92672">*&lt;/span> df[&lt;span style="color:#e6db74">&amp;#39;feature_b&amp;#39;&lt;/span>]
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%89%B9%E5%BE%81%E6%8F%90%E5%8F%96/">特征提取&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%95%B0%E6%8D%AE%E9%A2%84%E5%A4%84%E7%90%86/">数据预处理&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%A8%A1%E5%9E%8B%E8%B0%83%E4%BC%98/">模型调优&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E9%A2%86%E5%9F%9F%E4%B8%93%E9%95%BF/">领域专长&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>特征提取</title><link>https://terms-en.ai-term-hub.com/zh/terms/feature_extraction/</link><pubDate>Sat, 18 Jul 2026 11:17:01 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/feature_extraction/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>特征提取涉及将原始数据转换为一组更能代表潜在问题的特征，从而提高预测模型的准确性。该技术有助于减少数据维度并增强模型表现。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>从原始数据中推导有意义信息的过程，旨在降低维度并提高机器学习模型的性能。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>降维&lt;/li>
&lt;li>原始数据转换&lt;/li>
&lt;li>模式识别&lt;/li>
&lt;li>主成分&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>图像识别任务&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;li>音频信号处理&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> sklearn.decomposition &lt;span style="color:#f92672">import&lt;/span> PCA
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>pca &lt;span style="color:#f92672">=&lt;/span> PCA(n_components&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">2&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>reduced_data &lt;span style="color:#f92672">=&lt;/span> pca&lt;span style="color:#f92672">.&lt;/span>fit_transform(raw_data)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/pca/">PCA&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%B5%8C%E5%85%A5/">嵌入&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E7%89%B9%E5%BE%81%E9%80%89%E6%8B%A9/">特征选择&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%B7%B1%E5%BA%A6%E5%AD%A6%E4%B9%A0/">深度学习&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据探索</title><link>https://terms-en.ai-term-hub.com/zh/terms/data_exploration/</link><pubDate>Sat, 18 Jul 2026 11:12:36 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/data_exploration/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>数据探索，通常称为探索性数据分析（EDA），是机器学习工作流程中至关重要的初步步骤。它涉及总结数据的主要特征，经常使用可视化技术来揭示数据的内在结构和分布情况。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>在正式建模之前，对数据集进行初步分析，以发现模式、识别异常并验证假设。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>探索性数据分析&lt;/li>
&lt;li>可视化&lt;/li>
&lt;li>模式识别&lt;/li>
&lt;li>异常检测&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>在模型训练前识别特征之间的相关性&lt;/li>
&lt;li>检测和处理缺失值或异常值&lt;/li>
&lt;li>验证数据质量和分布假设&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> pandas &lt;span style="color:#66d9ef">as&lt;/span> pd
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> seaborn &lt;span style="color:#66d9ef">as&lt;/span> sns
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>df &lt;span style="color:#f92672">=&lt;/span> pd&lt;span style="color:#f92672">.&lt;/span>read_csv(&lt;span style="color:#e6db74">&amp;#39;data.csv&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>sns&lt;span style="color:#f92672">.&lt;/span>pairplot(df)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>plt&lt;span style="color:#f92672">.&lt;/span>show()
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/feature_engineering-%E7%89%B9%E5%BE%81%E5%B7%A5%E7%A8%8B/">feature_engineering (特征工程)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/data_cleaning-%E6%95%B0%E6%8D%AE%E6%B8%85%E6%B4%97/">data_cleaning (数据清洗)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/eda-%E6%8E%A2%E7%B4%A2%E6%80%A7%E6%95%B0%E6%8D%AE%E5%88%86%E6%9E%90/">EDA (探索性数据分析)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/statistical_analysis-%E7%BB%9F%E8%AE%A1%E5%88%86%E6%9E%90/">statistical_analysis (统计分析)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据标注</title><link>https://terms-en.ai-term-hub.com/zh/terms/data_annotation/</link><pubDate>Sat, 18 Jul 2026 11:12:24 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/data_annotation/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>这一关键步骤涉及为原始数据点附加有意义的元数据，以便算法能够学习输入与输出之间的关系。例如，在图像中物体周围绘制边界框……&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>数据标注是对原始数据（如图像或文本）进行标记的过程，使其适用于监督式机器学习训练。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>监督学习&lt;/li>
&lt;li>标记&lt;/li>
&lt;li>地面真值&lt;/li>
&lt;li>元数据&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练目标检测模型&lt;/li>
&lt;li>构建情感分析分类器&lt;/li>
&lt;li>创建语音识别数据集&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/ground-truth-%E5%9C%B0%E9%9D%A2%E7%9C%9F%E5%80%BC/">Ground Truth (地面真值)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/active-learning-%E4%B8%BB%E5%8A%A8%E5%AD%A6%E4%B9%A0/">Active Learning (主动学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/crowdsourcing-%E4%BC%97%E5%8C%85/">Crowdsourcing (众包)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/labeled-data-%E6%A0%87%E6%B3%A8%E6%95%B0%E6%8D%AE/">Labeled Data (标注数据)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>数据增强</title><link>https://terms-en.ai-term-hub.com/zh/terms/data_augmentation/</link><pubDate>Sat, 18 Jul 2026 11:12:24 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/data_augmentation/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>这种方法通过创建现有样本的修改版本来人工扩展训练数据集，例如旋转图像、在音频中添加噪声或在文本中进行同义词替换。它有助于防止……&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>数据增强是一种通过变换现有数据点来增加训练数据集多样性和规模的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>防止过拟合&lt;/li>
&lt;li>数据集扩展&lt;/li>
&lt;li>泛化能力&lt;/li>
&lt;li>变换&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>提高计算机视觉模型的鲁棒性&lt;/li>
&lt;li>在有限文本下增强自然语言处理模型的性能&lt;/li>
&lt;li>平衡不平衡的数据集&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> tensorflow.keras.preprocessing.image &lt;span style="color:#f92672">import&lt;/span> ImageDataGenerator
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>gen &lt;span style="color:#f92672">=&lt;/span> ImageDataGenerator(rotation_range&lt;span style="color:#f92672">=&lt;/span>&lt;span style="color:#ae81ff">20&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/regularization-%E6%AD%A3%E5%88%99%E5%8C%96/">Regularization (正则化)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/synthetic-data-%E5%90%88%E6%88%90%E6%95%B0%E6%8D%AE/">Synthetic Data (合成数据)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/transfer-learning-%E8%BF%81%E7%A7%BB%E5%AD%A6%E4%B9%A0/">Transfer Learning (迁移学习)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/overfitting-%E8%BF%87%E6%8B%9F%E5%90%88/">Overfitting (过拟合)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>分块</title><link>https://terms-en.ai-term-hub.com/zh/terms/chunking/</link><pubDate>Sat, 18 Jul 2026 11:09:54 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/chunking/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>分块是检索增强生成（RAG）和其他NLP管道中的关键预处理步骤。它涉及将文本划分为固定大小或语义单元（块），以适应上下文&amp;hellip;&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>上下文窗口&lt;/li>
&lt;li>RAG (检索增强生成)&lt;/li>
&lt;li>文本分割&lt;/li>
&lt;li>索引&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>构建RAG知识库&lt;/li>
&lt;li>处理长文档以进行摘要&lt;/li>
&lt;li>向量数据库数据摄入&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5%E5%90%91%E9%87%8F/">Embedding (嵌入向量)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vector-database-%E5%90%91%E9%87%8F%E6%95%B0%E6%8D%AE%E5%BA%93/">Vector Database (向量数据库)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenizer-%E5%88%86%E8%AF%8D%E5%99%A8/">Tokenizer (分词器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/rag-%E6%A3%80%E7%B4%A2%E5%A2%9E%E5%BC%BA%E7%94%9F%E6%88%90/">RAG (检索增强生成)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>分词</title><link>https://terms-en.ai-term-hub.com/zh/terms/tokenization/</link><pubDate>Sat, 18 Jul 2026 10:55:25 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/tokenization/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>分词是自然语言处理（NLP）中的关键预处理步骤，它将非结构化文本转换为适合模型输入的结构化数据。该过程涉及将句子分解为更小的单元，如单词、子词或字符，以便模型能够有效理解语义。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>分词是将原始文本拆分为称为词元的较小单元的过程，这些单元可以被机器学习算法处理。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本拆分&lt;/li>
&lt;li>预处理&lt;/li>
&lt;li>WordPiece&lt;/li>
&lt;li>字节对编码&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>为BERT训练准备数据集&lt;/li>
&lt;li>GPT模型的输入格式化&lt;/li>
&lt;li>情感分析的数据清洗&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">from&lt;/span> transformers &lt;span style="color:#f92672">import&lt;/span> AutoTokenizer
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokenizer &lt;span style="color:#f92672">=&lt;/span> AutoTokenizer&lt;span style="color:#f92672">.&lt;/span>from_pretrained(&lt;span style="color:#e6db74">&amp;#39;bert-base-uncased&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>tokens &lt;span style="color:#f92672">=&lt;/span> tokenizer&lt;span style="color:#f92672">.&lt;/span>tokenize(&lt;span style="color:#e6db74">&amp;#39;Hello world!&amp;#39;&lt;/span>)
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tokenizer-%E5%88%86%E8%AF%8D%E5%99%A8/">Tokenizer (分词器)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vocabulary-%E8%AF%8D%E6%B1%87%E8%A1%A8/">Vocabulary (词汇表)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/embedding-%E5%B5%8C%E5%85%A5/">Embedding (嵌入)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/preprocessing-%E9%A2%84%E5%A4%84%E7%90%86/">Preprocessing (预处理)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>