Learning Objectives
- Understand what separates Gemini 3.5 Transcribe from conventional speech-to-text
- Explain why edited transcription is the wrong tool for some jobs and the right one for others
- Judge the speed gain against the accuracy gain before choosing a transcription engine
What Is Gemini 3.5 Transcribe?
Gemini 3.5 Transcribe is Google's speech-to-text model, announced on August 26, 2026. It replaces Chirp 3, the previous Google voice-to-text engine, and it arrives in the Gemini 3.5 branch ahead of the general-purpose Gemini 3.5 Pro model.
The interesting claim is not that it hears words more accurately. It is that the model decides what you meant to say. As you speak, it strips the "ums" and "uhs" that clutter natural speech, and when you correct yourself mid-sentence it applies the correction rather than transcribing both attempts. A custom vocabulary handles specialised jargon that a general model would mangle.
It works across 85 languages and handles up to three speakers in pre-recorded audio.
💡Key Concept
Transcription versus dictation. A conventional speech-to-text engine aims for a faithful record of the sounds produced — every false start, every repetition. A dictation-oriented model aims for the text you would have typed. Those are different products, and the gap between them is exactly where this model sits.
The Speed Gain Is the Real Gain
Google reports the model reaches finished text about 70 percent faster than Chirp 3. The accuracy improvement is far more modest: the live-speech error rate falls from 7.32 percent to 5.5 percent.
That difference matters when choosing between engines. A roughly two-percentage-point accuracy gain is worth having but will not change what a transcription workflow can do. Cutting the wait between speaking and seeing usable text changes whether voice input feels viable at all — the friction in dictation has always been the pause, and the manual cleanup afterwards.
Where It Runs Today
Availability is staged rather than universal. The model already powers the Rambler feature in Gboard, though that remains limited to Pixel 11 phones, with Google saying it will reach more devices later in the year. Voice input in the Gemini app on macOS uses it as of launch.
Developers get broader access. Antigravity carries the model with access to screen context and chat history where permitted, AI Studio uses it for voice-driven building, and it is callable directly from the Gemini API. Google says Chrome support is coming, which would make it available in any web text field.
What It Is Not Suited For
The editing behaviour that makes this model good at dictation makes it wrong for several jobs, and the distinction is easy to miss.
Any workflow needing a verbatim record — legal proceedings, research interviews, qualitative analysis, accessibility captioning where disfluency carries meaning — wants the false starts preserved. A model that silently removes them is producing a cleaner document and a less truthful one.
The same applies to forced alignment, where a transcript must map word-for-word onto the audio timeline for captioning or media synchronisation. If the text has been edited, it no longer matches the waveform, and the alignment breaks. For that work, a verbatim engine with word-level timestamps remains the correct choice.
⚠️Warning
Three speakers is a low ceiling for meetings. Multi-speaker support is capped at three in pre-recorded audio, which covers an interview or a small call and does not cover a typical meeting. Check that limit against your actual use before planning around it.
Pricing
- Rambler in Gboard on Pixel 11
- Voice input in the Gemini app on macOS
- Antigravity with screen context
- AI Studio voice-driven building
- Direct model access
- Custom vocabulary support
- Pricing per Gemini API rates
Consumer access arrives bundled with the surfaces that carry it rather than sold separately. API use is billed on standard Gemini API terms.
Strengths
- Removes the cleanup step — filler words and spoken self-corrections are handled during transcription rather than afterwards
- Substantially faster to usable text — about 70 percent quicker than Chirp 3, which is the difference that changes whether dictation is practical
- Broad language coverage — 85 languages in the initial release
- Custom vocabulary — domain jargon and proper nouns can be supplied rather than guessed at
- Reachable from several surfaces at once — the Gemini API, AI Studio and Antigravity from launch day, not a staged developer preview
Limitations and Considerations
- The accuracy gain is small — 7.32 percent to 5.5 percent on live speech, so do not switch engines expecting a step change in correctness
- Edited output is unsuitable for verbatim work — legal, research and accessibility uses generally need the disfluencies kept
- Breaks forced alignment — an edited transcript no longer maps onto the audio timeline, so captioning and media-sync workflows need a verbatim engine instead
- Three-speaker ceiling — enough for an interview, not for a meeting
- Consumer availability is narrow — Rambler is Pixel 11 only at launch, and Chrome support is announced rather than shipped
Related Tools
- OpenAI TTS / Whisper — the verbatim speech-to-text baseline, and the right choice when word-level alignment matters
- MAI-Transcribe-1.5 — Microsoft's in-house speech-to-text model, the closest cross-vendor comparison
- Wispr Flow — dictation-first voice keyboard pursuing the same cleaned-up-text goal at the application layer
- Gemini — the assistant whose macOS voice input this model now powers
Key Takeaways
- Gemini 3.5 Transcribe is Google's speech-to-text model announced August 26, 2026, replacing the Chirp 3 engine
- It edits while transcribing — stripping filler words and applying spoken self-corrections — across 85 languages and up to three speakers
- The speed gain is the substantive one at about 70 percent; live-speech accuracy moves only from 7.32 percent to 5.5 percent
- That editing makes it wrong for verbatim work and for forced alignment, where an unedited transcript with word-level timestamps is required
- It is live in the Gemini app on macOS, Antigravity, AI Studio and the Gemini API, with Gboard limited to Pixel 11 and Chrome still to come










