On 2026-07-29, xAI announced Grok Voice Think Fast 2.0, the latest version of its voice dialogue AI model. This model employs a Speech-to-Speech approach, receiving user audio directly and generating responses in audio format without intermediate text conversion. Thanks to a proprietary architecture that allows inference to proceed in the background during conversation, the Time to First Audio has been significantly reduced from 1.25 seconds in the previous model to 0.70 seconds.

In terms of technical evolution, the amount of inference tokens used per response has been reduced to approximately 0.4 times that of the previous model, greatly improving inference efficiency. This enables the agent to complete tool calls, such as web searches and external API integrations, before it finishes speaking the first sentence. Additionally, xAI stated that transcription accuracy across 24 languages has improved, achieving 1.4 times the accuracy compared to the previous model.


Source: Grok Voice Think Fast 2.0とは?特徴や使い方を解説 - aismiley.co.jp (Google News: Grok)