<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Transcription on English AI Terms Dictionary</title><link>https://terms-en.ai-term-hub.com/en/tags/transcription/</link><description>Recent content in Transcription on English AI Terms Dictionary</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 18 Jul 2026 11:44:44 +0000</lastBuildDate><atom:link href="https://terms-en.ai-term-hub.com/en/tags/transcription/index.xml" rel="self" type="application/rss+xml"/><item><title>Speaker Diarization</title><link>https://terms-en.ai-term-hub.com/en/terms/speaker_diarization/</link><pubDate>Sat, 18 Jul 2026 10:16:18 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/speaker_diarization/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Speaker Diarization is the task of partitioning an audio stream into homogeneous segments according to the identity of the speaker. It combines speaker change detection with speaker clustering to label segments with unique speaker IDs. This technology is essential for making multi-party conversations understandable in transcripts, often referred to as the &amp;lsquo;who said what&amp;rsquo; problem.&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>The process of determining &amp;lsquo;who spoke when&amp;rsquo; in an audio recording.&lt;/p>
&lt;h2 id="key-concepts">Key Concepts&lt;/h2>
&lt;ul>
&lt;li>Speaker clustering&lt;/li>
&lt;li>Identity labeling&lt;/li>
&lt;li>Who-said-what&lt;/li>
&lt;li>Audio segmentation&lt;/li>
&lt;/ul>
&lt;h2 id="use-cases">Use Cases&lt;/h2>
&lt;ul>
&lt;li>Automatic meeting minutes generation&lt;/li>
&lt;li>Interview transcription&lt;/li>
&lt;li>Broadcast media analysis&lt;/li>
&lt;/ul>
&lt;h2 id="related-terms">Related Terms&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speaker_change_detection/">speaker_change_detection&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/speech_to_text/">speech_to_text&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/voice_printing/">voice_printing&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://terms-en.ai-term-hub.com/en/terms/audio_analysis/">audio_analysis&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>Automatic Speech Recognition</title><link>https://terms-en.ai-term-hub.com/en/terms/automatic_speech_recognition/</link><pubDate>Sat, 18 Jul 2026 09:47:17 +0000</pubDate><guid>https://terms-en.ai-term-hub.com/en/terms/automatic_speech_recognition/</guid><description>&lt;h2 id="definition">Definition&lt;/h2>
&lt;p>Automatic Speech Recognition (ASR), also known as speech-to-text, is a subfield of speech processing that leverages artificial intelligence to transcribe audio signals into written text. Modern ASR systems typically employ deep neural networks, such as recurrent neural networks or transformers, to map acoustic features to linguistic units. This technology enables voice interfaces, transcription services, and accessibility tools for hearing-impaired users.&lt;/p>
&lt;h3 id="summary">Summary&lt;/h3>
&lt;p>A technology that converts spoken language into text using deep learning models.&lt;/p></description></item></channel></rss>