Google has introduced Gemini 3.5 Transcribe, a new AI speech-to-text model designed to make voice typing faster, more accurate, and more natural. While users continue to wait for the rumored Gemini 3.5 Pro, the latest model is already powering the Pixel 11’s Gboard “Rambler” voice-input feature and is expected to expand across Google’s ecosystem.
According to Google, Gemini 3.5 Transcribe delivers significantly faster results than the company’s previous Chirp 3 speech-recognition engine. The new Google AI transcription model can reportedly process speech and produce final text approximately 70% faster. Google also measured a reduction in the live-audio error rate, from 7.32% with Chirp 3 to 5.5% with Gemini 3.5 Transcribe. Although the improvement is relatively modest, better accuracy could make correcting voice-typing mistakes much easier.
Gemini 3.5 Transcribe is designed to do more than simply convert speech into text. The model can identify the intended meaning of a sentence, remove filler words such as “um” and “uh,” and clean up verbal corrections as users speak. It also supports on-the-fly text editing and custom vocabulary, making it easier to accurately transcribe industry-specific terms, names, and other jargon.
Google says these voice-input improvements are available in 85 languages. The model can also process prerecorded audio featuring up to three speakers, making it useful for conversations, interviews, meetings, and other multi-speaker recordings.
However, the model’s ability to interpret and refine speech may not be suitable for every situation. Gemini 3.5 Transcribe appears particularly effective for short passages, where it can resolve hesitations, repeated words, and verbal stumbles. Yet because the AI may alter the original wording rather than transcribe it exactly, users may prefer traditional speech-to-text tools when accuracy and verbatim transcription are essential.
Source: arstechnica.com


