<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Speech on 中文AI术语词典</title><link>https://terms-en.ai-term-hub.com/zh/tags/speech/</link><description>Recent content in Speech on 中文AI术语词典</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Sat, 18 Jul 2026 11:44:45 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/zh/tags/speech/index.xml" rel="self" type="application/rss+xml"/><item><title>语音</title><link>https://terms-en.ai-term-hub.com/zh/terms/voice/</link><pubDate>Sat, 18 Jul 2026 11:37:50 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/voice/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>在人工智能领域，语音涵盖了由人类声带产生并携带语言信息的声学信号。与一般音频不同，它特指与说话相关的信号。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>语音是指由人类发音产生的声音，它是语音识别和合成系统的主要输入模态。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>语音识别&lt;/li>
&lt;li>音频信号处理&lt;/li>
&lt;li>说话人识别&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>如 Siri 和 Alexa 等虚拟助手&lt;/li>
&lt;li>自动转录服务&lt;/li>
&lt;li>视障人士的辅助工具&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speech_recognition-%E8%AF%AD%E9%9F%B3%E8%AF%86%E5%88%AB/">speech_recognition (语音识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/text_to_speech-%E6%96%87%E6%9C%AC%E8%BD%AC%E8%AF%AD%E9%9F%B3/">text_to_speech (文本转语音)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/audio_processing-%E9%9F%B3%E9%A2%91%E5%A4%84%E7%90%86/">audio_processing (音频处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/nlp-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">nlp (自然语言处理)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>语音活动检测</title><link>https://terms-en.ai-term-hub.com/zh/terms/voice_activity_detection/</link><pubDate>Sat, 18 Jul 2026 11:37:50 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/voice_activity_detection/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>VAD 算法实时分析音频流，以区分活跃语音时段和非语音间隔（如背景噪声或停顿）。这对于优化带宽至关重要。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>语音活动检测（VAD）是一种信号处理技术，用于识别包含人类语音的音频片段与静音或噪声之间的区别。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>静音抑制&lt;/li>
&lt;li>信号分类&lt;/li>
&lt;li>实时处理&lt;/li>
&lt;li>抗噪性&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>高效的 VoIP 通信&lt;/li>
&lt;li>语音到文本引擎的预处理&lt;/li>
&lt;li>唤醒词检测系统&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speech_recognition-%E8%AF%AD%E9%9F%B3%E8%AF%86%E5%88%AB/">speech_recognition (语音识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/noise_cancellation-%E9%99%8D%E5%99%AA/">noise_cancellation (降噪)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/audio_segmentation-%E9%9F%B3%E9%A2%91%E5%88%86%E5%89%B2/">audio_segmentation (音频分割)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wake_word-%E5%94%A4%E9%86%92%E8%AF%8D/">wake_word (唤醒词)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Vibevoice</title><link>https://terms-en.ai-term-hub.com/zh/terms/vibevoice/</link><pubDate>Sat, 18 Jul 2026 11:37:39 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/vibevoice/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Vibevoice 是一种概念性或品牌化的文本转语音（TTS）技术方法，强调捕捉人类语音的‘氛围’或情感细微差别。与传统可能听起来单调的 TTS 不同，&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Vibevoice 指的是一种 AI 生成的语音合成风格，优先考虑自然、富有情感且具备语境意识的语音表达，而非僵硬的机械精度。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>情感 TTS&lt;/li>
&lt;li>韵律建模&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;li>语音克隆&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>AI 伴侣和聊天机器人&lt;/li>
&lt;li>游戏中的叙事讲述&lt;/li>
&lt;li>无障碍内容创作&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/text-to-speech-%E6%96%87%E6%9C%AC%E8%BD%AC%E8%AF%AD%E9%9F%B3/">Text-to-Speech (文本转语音)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/voice-synthesis-%E8%AF%AD%E9%9F%B3%E5%90%88%E6%88%90/">Voice Synthesis (语音合成)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/emotion-ai-%E6%83%85%E6%84%9F%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD/">Emotion AI (情感人工智能)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/audio-generation-%E9%9F%B3%E9%A2%91%E7%94%9F%E6%88%90/">Audio Generation (音频生成)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>母语识别</title><link>https://terms-en.ai-term-hub.com/zh/terms/native_language_identification/</link><pubDate>Sat, 18 Jul 2026 11:28:00 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/native_language_identification/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>母语识别（NLI）是自然语言处理的一个子领域，专注于识别说话者学习的第一语言。与一般的语言检测不同，NLI 分析说话者在发音、语法结构和词汇选择上无意识的母语特征。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>自动从说话者的语音或文本样本中确定其母语的过程。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>说话人画像&lt;/li>
&lt;li>语言特征&lt;/li>
&lt;li>口音识别&lt;/li>
&lt;li>自然语言处理&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>生物特征安全认证&lt;/li>
&lt;li>个性化客户服务交互&lt;/li>
&lt;li>社会语言学人口统计分析&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%AD%E8%A8%80%E6%A3%80%E6%B5%8B-language-detection/">语言检测 (Language Detection)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%B4%E8%AF%9D%E4%BA%BA%E5%88%86%E7%A6%BB-speaker-diarization/">说话人分离 (Speaker Diarization)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%8F%A3%E9%9F%B3%E8%AF%86%E5%88%AB-accent-identification/">口音识别 (Accent Identification)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E6%B3%95%E5%8C%BB%E8%AF%AD%E8%A8%80%E5%AD%A6-forensic-linguistics/">法医语言学 (Forensic Linguistics)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Moshi</title><link>https://terms-en.ai-term-hub.com/zh/terms/moshi/</link><pubDate>Sat, 18 Jul 2026 11:26:37 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/moshi/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Moshi 是由 Kyutai 创建的高级 AI 模型，它将语音和文本处理整合到一个统一的框架中。与传统系统在处理后转换语音为文本的系统不同，Moshi 学习联合表示。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>由 Kyutai 开发的语音语言模型，共同学习文本和音频表示，以实现无缝的多模态交互。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>多模态学习&lt;/li>
&lt;li>语音-文本联合建模&lt;/li>
&lt;li>韵律保留&lt;/li>
&lt;li>实时交互&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>构建自然语音助手&lt;/li>
&lt;li>增强带有情感语调的互动式故事讲述&lt;/li>
&lt;li>改善听障用户的辅助工具&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/kyutai-kyutai-%E5%85%AC%E5%8F%B8/">Kyutai (Kyutai 公司)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%A4%9A%E6%A8%A1%E6%80%81-ai-multimodal-ai/">多模态 AI (Multimodal AI)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%AD%E9%9F%B3%E8%AF%86%E5%88%AB-speech-recognition/">语音识别 (Speech Recognition)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E5%AF%B9%E8%AF%9D%E5%BC%8F-ai-conversational-ai/">对话式 AI (Conversational AI)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Hugging Face ASR 排行榜</title><link>https://terms-en.ai-term-hub.com/zh/terms/hf_asr_leaderboard/</link><pubDate>Sat, 18 Jul 2026 11:20:51 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/hf_asr_leaderboard/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>HF ASR 排行榜是由 Hugging Face 托管的一个社区驱动指标平台，追踪自动语音识别领域的最新性能表现。它允许研究人员和开发者&amp;hellip;&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Hugging Face 上的一个排名系统，用于评估和比较自动语音识别模型的性能。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>基准测试&lt;/li>
&lt;li>自动语音识别&lt;/li>
&lt;li>开源&lt;/li>
&lt;li>性能指标&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>部署时的模型选择&lt;/li>
&lt;li>研究进展跟踪&lt;/li>
&lt;li>社区贡献评估&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/hugging_face-hugging-face/">hugging_face (Hugging Face)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wer-%E8%AF%8D%E9%94%99%E8%AF%AF%E7%8E%87/">wer (词错误率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speech_model-%E8%AF%AD%E9%9F%B3%E6%A8%A1%E5%9E%8B/">speech_model (语音模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/benchmark-%E5%9F%BA%E5%87%86/">benchmark (基准)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Coqui</title><link>https://terms-en.ai-term-hub.com/zh/terms/coqui/</link><pubDate>Sat, 18 Jul 2026 11:11:39 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/coqui/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Coqui Technologies 是开源 AI 社区中的知名参与者，最著名的是其 TTS（文本转语音）引擎。该项目提供了预训练模型，能够生成自然 sounding 的语音&amp;hellip;&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>Coqui 是一家开源语音技术公司，以开发高质量的多语言文本转语音模型而闻名。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>文本转语音&lt;/li>
&lt;li>开源&lt;/li>
&lt;li>多语言合成&lt;/li>
&lt;li>声音克隆&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>视障人士的辅助工具&lt;/li>
&lt;li>视频自动配音&lt;/li>
&lt;li>交互式语音响应系统&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/tacotron-tacotron-%E6%A8%A1%E5%9E%8B/">Tacotron (Tacotron 模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/vits-vits-%E6%A8%A1%E5%9E%8B/">VITS (VITS 模型)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%AF%AD%E9%9F%B3%E8%AF%86%E5%88%AB-speech-recognition/">语音识别 (Speech Recognition)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86-nlp/">自然语言处理 (NLP)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>自动语音识别</title><link>https://terms-en.ai-term-hub.com/zh/terms/automatic_speech_recognition/</link><pubDate>Sat, 18 Jul 2026 11:08:15 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/automatic_speech_recognition/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>自动语音识别（ASR），也称为语音转文字，是语音处理的一个子领域，它利用人工智能将音频信号转录为书面文本。现代ASR系统能够处理各种口音、背景噪音和连续语音，广泛应用于人机交互领域。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>一种利用深度学习模型将口语转换为文本的技术。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>声学建模&lt;/li>
&lt;li>语言建模&lt;/li>
&lt;li>深度学习&lt;/li>
&lt;li>转录&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>语音助手（如Siri、Alexa）&lt;/li>
&lt;li>视频实时字幕生成&lt;/li>
&lt;li>专业人士的听写软件&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural_language_processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">natural_language_processing (自然语言处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speaker_identification-%E8%AF%B4%E8%AF%9D%E4%BA%BA%E8%AF%86%E5%88%AB/">speaker_identification (说话人识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/audio_processing-%E9%9F%B3%E9%A2%91%E5%A4%84%E7%90%86/">audio_processing (音频处理)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/deep_learning-%E6%B7%B1%E5%BA%A6%E5%AD%A6%E4%B9%A0/">deep_learning (深度学习)&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>ASR-complete</title><link>https://terms-en.ai-term-hub.com/zh/terms/asr_complete/</link><pubDate>Sat, 18 Jul 2026 11:04:01 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/zh/terms/asr_complete/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>术语“ASR-complete”表示自动语音识别系统在特定且定义明确的任务和数据集上，其性能已达到与人类转录员相当的水平。这是一个重要的里程碑，标志着系统在特定领域内的识别精度已满足实际应用的高标准要求。&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>ASR-complete 描述在标准化基准数据集上达到人类水平准确率的语音识别系统。&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>语音识别&lt;/li>
&lt;li>人类水平准确率&lt;/li>
&lt;li>错误率&lt;/li>
&lt;li>基准测试&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>评估 ASR 模型性能&lt;/li>
&lt;li>制定行业标准&lt;/li>
&lt;li>比较不同的声学模型&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/automatic-speech-recognition-%E8%87%AA%E5%8A%A8%E8%AF%AD%E9%9F%B3%E8%AF%86%E5%88%AB/">Automatic Speech Recognition (自动语音识别)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/wer-%E8%AF%8D%E9%94%99%E8%AF%AF%E7%8E%87/">WER (词错误率)&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/natural-language-processing-%E8%87%AA%E7%84%B6%E8%AF%AD%E8%A8%80%E5%A4%84%E7%90%86/">Natural Language Processing (自然语言处理)&lt;/a>&lt;/li>
&lt;/ul></description></item></channel></rss>