
Google has introduced Gemini 3.5 Transcribe, a new speech-to-text model.
The model can offer more accurate and context-aware transcriptions. It can handle background noise, filler words, self-corrections, specialized terms, and different accents. It can also automatically format transcribed text and detect more than 85 languages.
Gemini 3.5 Transcribe is available through two APIs. The Live API supports real-time streaming with sub-second latency, while the Interactions API can process recorded audio with speaker identification and word-level timestamps.
Google says the model achieved a 4.0% Word Error Rate (WER) for streaming and 2.6% for non-streaming use cases, based on measurements from Artificial Analysis. It also improves final transcription latency by up to 70% compared with the previous Chirp 3 model.

For developers, Gemini 3.5 Transcribe is currently in public preview through the Gemini API and Google Antigravity. It is also available in preview through the Gemini Enterprise Agent Platform.
Google says the feature is coming soon to Chrome, where users will be able to use their voice to type into web fields.

ManilaShaker is a tech media producing insightful and helpful content for our local and growing international audience. Our goal is to create a premier Philippine digital consumer electronics resource that provides the most objective reviews and comparisons globally.