Free to read. Sign up to save your progress and take knowledge-check quizzes.

Sign up free
6 min read·Updated August 27, 2026

Gemini 3.5 Transcribe

Google logoBy GoogleGoogle on YouTube

Gemini 3.5 Transcribe is Google's speech-to-text model, announced August 26, 2026. It removes filler words and applies spoken self-corrections while transcribing, across 85 languages and up to three speakers, reaching finished text about 70 percent faster than the Chirp 3 engine it replaces — though the raw accuracy gain is much smaller than the speed gain.

Share

Listen to this lesson

Free preview · first 0:30
0:00 / 0:30

Audio & video lessons are paid features

Plus unlocks audio streaming. Pro adds downloadable audio, video, certificates, and more.

Plus adds:
  • Audio streaming
  • Downloadable PDFs
  • All AI Playbooks
  • Personalized content
Pro also adds:
  • Certificates of completion
  • Audio MP3 downloads
  • Video lessonssoon
  • & More…soon

Watch this lesson

AI Pro Playbook video — coming soon

Learning Objectives

  • Understand what separates Gemini 3.5 Transcribe from conventional speech-to-text
  • Explain why edited transcription is the wrong tool for some jobs and the right one for others
  • Judge the speed gain against the accuracy gain before choosing a transcription engine

What Is Gemini 3.5 Transcribe?

Gemini 3.5 Transcribe is Google's speech-to-text model, announced on August 26, 2026. It replaces Chirp 3, the previous Google voice-to-text engine, and it arrives in the Gemini 3.5 branch ahead of the general-purpose Gemini 3.5 Pro model.

The interesting claim is not that it hears words more accurately. It is that the model decides what you meant to say. As you speak, it strips the "ums" and "uhs" that clutter natural speech, and when you correct yourself mid-sentence it applies the correction rather than transcribing both attempts. A custom vocabulary handles specialised jargon that a general model would mangle.

It works across 85 languages and handles up to three speakers in pre-recorded audio.

💡Key Concept

Transcription versus dictation. A conventional speech-to-text engine aims for a faithful record of the sounds produced — every false start, every repetition. A dictation-oriented model aims for the text you would have typed. Those are different products, and the gap between them is exactly where this model sits.

The Speed Gain Is the Real Gain

Google reports the model reaches finished text about 70 percent faster than Chirp 3. The accuracy improvement is far more modest: the live-speech error rate falls from 7.32 percent to 5.5 percent.

That difference matters when choosing between engines. A roughly two-percentage-point accuracy gain is worth having but will not change what a transcription workflow can do. Cutting the wait between speaking and seeing usable text changes whether voice input feels viable at all — the friction in dictation has always been the pause, and the manual cleanup afterwards.

Where It Runs Today

Availability is staged rather than universal. The model already powers the Rambler feature in Gboard, though that remains limited to Pixel 11 phones, with Google saying it will reach more devices later in the year. Voice input in the Gemini app on macOS uses it as of launch.

Developers get broader access. Antigravity carries the model with access to screen context and chat history where permitted, AI Studio uses it for voice-driven building, and it is callable directly from the Gemini API. Google says Chrome support is coming, which would make it available in any web text field.

What It Is Not Suited For

The editing behaviour that makes this model good at dictation makes it wrong for several jobs, and the distinction is easy to miss.

Any workflow needing a verbatim record — legal proceedings, research interviews, qualitative analysis, accessibility captioning where disfluency carries meaning — wants the false starts preserved. A model that silently removes them is producing a cleaner document and a less truthful one.

The same applies to forced alignment, where a transcript must map word-for-word onto the audio timeline for captioning or media synchronisation. If the text has been edited, it no longer matches the waveform, and the alignment breaks. For that work, a verbatim engine with word-level timestamps remains the correct choice.

⚠️Warning

Three speakers is a low ceiling for meetings. Multi-speaker support is capped at three in pre-recorded audio, which covers an interview or a small call and does not cover a typical meeting. Check that limit against your actual use before planning around it.

Pricing

ConsumerIncluded
  • Rambler in Gboard on Pixel 11
  • Voice input in the Gemini app on macOS
Developer toolsIncluded
  • Antigravity with screen context
  • AI Studio voice-driven building
Gemini APIUsage-based
  • Direct model access
  • Custom vocabulary support
  • Pricing per Gemini API rates

Consumer access arrives bundled with the surfaces that carry it rather than sold separately. API use is billed on standard Gemini API terms.

Strengths

  • Removes the cleanup step — filler words and spoken self-corrections are handled during transcription rather than afterwards
  • Substantially faster to usable text — about 70 percent quicker than Chirp 3, which is the difference that changes whether dictation is practical
  • Broad language coverage — 85 languages in the initial release
  • Custom vocabulary — domain jargon and proper nouns can be supplied rather than guessed at
  • Reachable from several surfaces at once — the Gemini API, AI Studio and Antigravity from launch day, not a staged developer preview

Limitations and Considerations

  • The accuracy gain is small — 7.32 percent to 5.5 percent on live speech, so do not switch engines expecting a step change in correctness
  • Edited output is unsuitable for verbatim work — legal, research and accessibility uses generally need the disfluencies kept
  • Breaks forced alignment — an edited transcript no longer maps onto the audio timeline, so captioning and media-sync workflows need a verbatim engine instead
  • Three-speaker ceiling — enough for an interview, not for a meeting
  • Consumer availability is narrow — Rambler is Pixel 11 only at launch, and Chrome support is announced rather than shipped
  • OpenAI TTS / Whisper — the verbatim speech-to-text baseline, and the right choice when word-level alignment matters
  • MAI-Transcribe-1.5 — Microsoft's in-house speech-to-text model, the closest cross-vendor comparison
  • Wispr Flow — dictation-first voice keyboard pursuing the same cleaned-up-text goal at the application layer
  • Gemini — the assistant whose macOS voice input this model now powers

Key Takeaways

  • Gemini 3.5 Transcribe is Google's speech-to-text model announced August 26, 2026, replacing the Chirp 3 engine
  • It edits while transcribing — stripping filler words and applying spoken self-corrections — across 85 languages and up to three speakers
  • The speed gain is the substantive one at about 70 percent; live-speech accuracy moves only from 7.32 percent to 5.5 percent
  • That editing makes it wrong for verbatim work and for forced alignment, where an unedited transcript with word-level timestamps is required
  • It is live in the Gemini app on macOS, Antigravity, AI Studio and the Gemini API, with Gboard limited to Pixel 11 and Chrome still to come

Save your progress & take the quiz

Sign up free to bookmark lessons, track which modules you've completed, and lock in what you learned with a quick knowledge-check quiz at the end of each lesson.

📰Gemini 3.5 Transcribe in the News

Showing the only story where Gemini 3.5 Transcribe is tagged in Top AI Stories.

More tools in Voice & Audio (12 of 18)

More from Google

🧭Recommended for you