<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Speech Recognition on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/speech-recognition/</link><description>Recent content in Speech Recognition on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/speech-recognition/index.xml" rel="self" type="application/rss+xml"/><item><title>Whisper</title><link>https://terms-en.ai-term-hub.com/zh/terms/whisper/</link><pubDate>Sat, 18 Jul 2026 11:38:14 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/whisper/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Whisper 是一个通用语音识别模型，旨在处理多种语言、方言和口音。它是在数十万小时的多语言和 multitask 监督数据上训练而成的，具有强大的鲁棒性。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>由 OpenAI 开发的自动语音识别 (ASR) 系统，基于大量多样化音频数据集训练而成。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>自动语音识别&lt;/li>
&lt;li>多语言支持&lt;/li>
&lt;li>抗噪鲁棒性&lt;/li>
&lt;li>Transformer 架构&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>视频字幕生成&lt;/li>
&lt;li>会议或讲座转录&lt;/li>
&lt;li>语音命令处理&lt;/li>
&lt;/ul>
&lt;h2 id="code-example">Code Example&lt;/h2>
&lt;div class="highlight">&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-python" data-lang="python">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f92672">import&lt;/span> whisper
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>model &lt;span style="color:#f92672">=&lt;/span> whisper&lt;span style="color:#f92672">.&lt;/span>load_model(&lt;span style="color:#e6db74">&amp;#34;base&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>result &lt;span style="color:#f92672">=&lt;/span> model&lt;span style="color:#f92672">.&lt;/span>transcribe(&lt;span style="color:#e6db74">&amp;#34;audio.mp3&amp;#34;&lt;/span>)
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>print(result[&lt;span style="color:#e6db74">&amp;#34;text&amp;#34;&lt;/span>])
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%AD%E9%9F%B3%E8%BD%AC%E6%96%87%E6%9C%AC-speech-to-text/">语音转文本 (Speech-to-text)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86-natural-language-processing/">自然语言处理 (Natural Language Processing)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/openai/">OpenAI&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E9%9F%B3%E9%A2%91%E5%88%86%E7%B1%BB-audio-classification/">音频分类 (Audio classification)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>