Listen to this lesson
Unlock audio and more
Audio streaming, downloadable PDFs and certificates come with Plus and Pro.
What it means
Automatic speech recognition became genuinely reliable ahead of most other AI capabilities, and it is now embedded almost everywhere: meeting notes, captions, voice input, call analytics, media search.
Accuracy is high on clear speech and degrades in predictable ways — background noise, crosstalk, strong accents underrepresented in training data, and specialist vocabulary. Speaker separation in multi-person recordings remains meaningfully harder than transcription itself.
A quiet risk with model-based transcription is that errors are fluent. Where older systems produced obvious garbage on unclear audio, a modern model may produce a plausible sentence that was never said — the same failure shape as hallucination, and harder to spot precisely because it reads naturally.
Why it matters
Transcription converts speech — a huge share of business communication — into something searchable and analyzable. The fluent-error property matters most where transcripts become records: medical, legal and compliance settings need verification against the audio rather than trust in the text.
In practice
Supply domain vocabulary where the tool allows it — names, product terms and acronyms are the common error source. For anything that becomes a record, spot-check against the audio rather than accepting the transcript.
Where this shows up
Tools and models in our catalog.
OpenAI TTS / WhisperOpenAI's text-to-speech (TTS) and speech-to-text (Whisper) APIs. Whisper is open-source and industry-leading for transcription accuracy across 100+ languages.
MAI-Transcribe-2Microsoft's in-house speech-to-text model, released September 3, 2026. Microsoft reports 60 languages at a 5.2% average word error rate on FLEURS, and 10x the speed of GPT-Transcribe, 7x ElevenLabs Scribe v2 and 5x Gemini 3.5 Transcribe. 10 cents/hour as a limited-time rate through end of 2026. A real-time Streaming version followed on October 1, 2026, returning first words in about 100 milliseconds at 54 cents/hour. Note the 5.2% is not comparable to v1.5's 2.4%, which averaged over 43 languages rather than 60.
Wispr FlowAI voice keyboard for Mac, Windows, iOS, and Android — press a hotkey, dictate into any app, and Flow transcribes with AI filler-word cleanup and context-aware formatting. Distinct from OpenAI Whisper despite the similar name.
DescriptAI-powered video and podcast editing platform. Edit video like a doc, remove filler words, clone your voice, and create AI overdub replacements.