Listen to this lesson
Unlock audio and more
Audio streaming, downloadable PDFs and certificates come with Plus and Pro.
What it means
Text-to-speech has moved from the flat, obviously synthetic voices of a decade ago to output that carries intonation, emphasis and emotional tone. Modern systems handle pacing and expressiveness well enough that most listeners do not identify them as synthetic in ordinary listening.
What still trips them is anything requiring interpretation rather than pronunciation: heteronyms whose reading depends on meaning, abbreviations that could be spoken or spelled, numbers and currency in ambiguous formats, and unusual proper nouns. These are the residual failures, and they are the reason narration scripts are frequently written differently from the text they narrate.
Realtime speech-to-speech models, which skip the text step entirely, are the current frontier and enable natural conversational latency.
Why it matters
TTS makes any written content consumable hands-free, which is a genuine accessibility gain and a genuine convenience one. The same capability underlies voice cloning, so the technology that narrates your documentation is the technology that impersonates a CEO on a phone call.
In practice
Expect to adjust source text for narration rather than feeding it verbatim — spelling out abbreviations and disambiguating heteronyms is normal authoring work, not a workaround. Always listen to output before publishing.
Where this shows up
Tools and models in our catalog.
ElevenLabsThe leading AI voice generation platform. Ultra-realistic text-to-speech in 32 languages, voice cloning, and a massive voice library. Used by 1M+ creators.
OpenAI TTS / WhisperOpenAI's text-to-speech (TTS) and speech-to-text (Whisper) APIs. Whisper is open-source and industry-leading for transcription accuracy across 100+ languages.
MAI-Voice-2.1Microsoft's in-house text-to-speech model, released October 1, 2026 as the successor to MAI-Voice-2. Speaks 23 languages across 26 locales, lets one voice keep a native accent in each, and clones a voice from a few seconds of audio behind consent checks. $22 per million characters, or $15 for the Flash version built for live voice agents. In Microsoft Foundry.
Murf AIProfessional AI voiceover studio with 120+ voices in 20+ languages. Includes timeline-based video sync, pitch/speed controls, and team collaboration.