Google has introduced Gemini 3.5 Transcribe, a new AI model designed to convert spoken audio into accurate and easy-to-read text. The company said the model is its most precise speech-to-text model so far.
According to a Google Blog, Gemini 3.5 Transcribe can automatically detect more than 85 languages, including conversations where speakers switch between languages.
It is designed to handle background noise, technical terms and natural speech more effectively.
One of its key features is smart transcription. The model can remove filler words such as ‘um’ and ‘uh’, clean up repeated words and understand corrections made while speaking.
This means the final transcript can be more polished instead of simply copying every word from the recording.
For recorded audio, the model can also identify speakers and provide word-level timestamps. Google says it can support speaker attribution for up to three speakers, while support for more speakers is experimental.
The technology could be useful for meetings, interviews, call recordings, voice assistants, captions and other applications that need fast and accurate speech-to-text conversion.
Google has also made Gemini 3.5 Transcribe available to developers through its Gemini API, while a real-time version is available for live transcription applications.
Sudhir Chaudhary wins two big Honours at XIIᵗʰ BCS Ratna Award 2026
12th BCS Ratna Award: JioStar’s Aravamudhan is Lifetime Achievement honouree
XIIth BCS Ratna Award : JioStar CEO Entertainment Kevin Vaz honoured
XIIth BCS Ratna Award 2026: Media achievers honoured
12th BCS Ratna Award a roaring success; honours excellence in M&E
Twelfth BCS Ratna Award boasts stellar lineup; to be held on Aug 5
India b’band, telecom services subs bases rise in July: TRAI data
Anil Kapoor-hosted show to debut on JioStar network Sept 19
Google unveils new AI model Gemini 3.5 Transcribe
Zee to focus on sports, animation to build new opportunities
Danny Ramirez joins Universal Pictures’ ‘Miami Vice ’85’ cast 


