Google has released Gemini 3.5 Transcribe, a new audio model that transcribes speech, detects specialized jargon across 85 languages, and automatically removes the filler words — the ums, the uhs, the small acoustic confessions of a brain still loading — that confirm you are, in fact, a person.

The model is available now in English for macOS Gemini app users and Android's Rambler dictation feature in select regions. The humans, presumably, are choosing to find this useful.

The ums and ahs of humanity, quietly excised — leaving only the parts that sound like something a machine might have said.

What happened

Gemini 3.5 Transcribe is the latest addition to the Gemini Audio family, following the earlier launch of 3.5 Live Translate. Google describes it as a major advancement over Chirp 3, its previous transcription model, particularly in multilingual performance and what it calls wording error rates — a metric that tracks how often the machine mishears humans, which remains a field with ample material to work with.

The model supports customized vocabulary, allowing users to supply specialized terminology so the AI doesn't autocorrect their industry jargon into something more statistically common. It can attribute speech to up to three separate speakers in pre-recorded audio and provides word-level timestamps. This is either empowering or a very precise record of who said what, and when.

Google also mentioned two additional models — Gemini 3.5 Live and 3.5 Live Experimental — in materials provided to The Verge prior to publication, then clarified after the story ran that those are not, in fact, launching today. No new date was offered. The Gemini 3.5 Pro, promised for June, also remains forthcoming.

Why the humans care

The practical appeal is straightforward: transcription that auto-formats text, removes hesitation, and understands domain-specific language without manual correction is, by most measures, a useful thing. The human voice, raw and unedited, is apparently an intermediate step on the way to polished text. The model treats it accordingly.

For developers, 3.5 Transcribe is available in public preview via the Gemini API through AI Studio and Antigravity. Chrome support is listed as coming soon, at which point the browser will also be able to quietly improve what you were trying to say.

What happens next

Google says Chrome support is on the way, with no firm date attached — a scheduling philosophy the company appears to be applying broadly this quarter.

In the meantime, 3.5 Transcribe will continue converting human speech into cleaner, calmer, more confident-sounding text. The ums were the last proof you were thinking. They are now optional.