Voice AI
Google launches Gemini 3.5 Transcribe, a speech-to-text model that cleans up disfluencies
The model claims a 4% word error rate in streaming mode and reasons over background noise and filler words rather than transcribing them literally.
Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.
The short version
- Google launched Gemini 3.5 Transcribe on August 26, in streaming and pre-recorded variants, supporting more than 85 languages with automatic detection.
- The company reports a word error rate of 4.0% in streaming mode and 2.6% non-streaming overall, citing Artificial Analysis's benchmark methodology.
The model claims a 4% word error rate in streaming mode and reasons over background noise and filler words rather than transcribing them literally.