Updated Aug 20, 2026

Text-to-Speech

TTS

Converting written text into spoken audio — now close enough to human that listeners often cannot tell.

Share

What it means

Text-to-speech has moved from the flat, obviously synthetic voices of a decade ago to output that carries intonation, emphasis and emotional tone. Modern systems handle pacing and expressiveness well enough that most listeners do not identify them as synthetic in ordinary listening.

What still trips them is anything requiring interpretation rather than pronunciation: heteronyms whose reading depends on meaning, abbreviations that could be spoken or spelled, numbers and currency in ambiguous formats, and unusual proper nouns. These are the residual failures, and they are the reason narration scripts are frequently written differently from the text they narrate.

Realtime speech-to-speech models, which skip the text step entirely, are the current frontier and enable natural conversational latency.

Why it matters

TTS makes any written content consumable hands-free, which is a genuine accessibility gain and a genuine convenience one. The same capability underlies voice cloning, so the technology that narrates your documentation is the technology that impersonates a CEO on a phone call.

In practice

Expect to adjust source text for narration rather than feeding it verbatim — spelling out abbreviations and disambiguating heteronyms is normal authoring work, not a workaround. Always listen to output before publishing.

Where this shows up

Tools and models in our catalog.

Related terms