On August 26, 2026, Google announced "Gemini 3.5 Transcribe," a speech recognition model. The model is provided via the Gemini API in two versions: one for real-time applications and one for recorded audio. It can automatically detect over 85 languages and features capabilities such as filler word removal, correction of self-corrections, and automatic formatting. Google stated that the time required to complete the final transcription has improved by 70% compared to the previous generation, Chirp 3.
For Japanese, the model is scheduled to be integrated into Google AI Studio and "Rambler," the voice input feature of Gboard for the Pixel 11 series. Additionally, benchmarks by Artificial Analysis reported an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming.
Source: Gemini 3.5 Transcribe|4.0%は日本語を測った数字ではない - innovaTopia (Google News: Gemini)