<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Code Generation on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/code-generation/</link><description>Recent content in Code Generation on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/code-generation/index.xml" rel="self" type="application/rss+xml"/><item><title>数据集：Bigcode/The Stack Dedup</title><link>https://terms-en.ai-term-hub.com/zh/terms/datasetbigcodethe_stack_dedup/</link><pubDate>Sat, 18 Jul 2026 11:12:42 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/datasetbigcodethe_stack_dedup/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>The Stack Dedup 是 The Stack（一个庞大的开源代码仓库）的一个专用子集。它应用严格的技术来消除冗余的代码片段，从而避免大型语言模型在训练时产生偏差。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>由 BigCode 整理的 The Stack 数据集的去重版本，旨在移除近乎重复的代码片段，以提供更清洁的训练数据。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>去重&lt;/li>
&lt;li>代码大语言模型&lt;/li>
&lt;li>数据质量&lt;/li>
&lt;li>The Stack&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>训练代码生成模型&lt;/li>
&lt;li>评估编码能力基准测试&lt;/li>
&lt;li>减少训练冗余&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/the-stack/">The Stack&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/bigcode/">BigCode&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E4%BB%A3%E7%A0%81%E6%90%9C%E7%B4%A2-code-search/">代码搜索 (Code Search)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%A4%A7%E6%A8%A1%E5%9E%8B%E8%AE%AD%E7%BB%83%E6%95%B0%E6%8D%AE-llm-training-data/">大模型训练数据 (LLM Training Data)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>