English

Model ReleasesSpaceX AIGrok Voice Transcribe 2.0

SpaceXAI Releases Grok Voice Transcribe 2.0 with Improved Accuracy and Multilingual Support

SpaceXAI has released Grok Voice Transcribe 2.0, its latest speech-to-text model designed for high accuracy and cost-effectiveness in real-world environments. According to the company, the new model is twice as accurate as Grok Voice Transcribe 1.0 while maintaining the same pricing structure.

The model is built on the audio foundation model that powers Grok Voice, which is already used in customer support calls and Tesla vehicles. Grok Voice Transcribe 2.0 is specifically optimized for challenging audio conditions, such as noisy environments, multiple speakers, and varying accents. On the Artificial Analysis leaderboard, it currently ranks first for accuracy among 32 streaming models.

Grok Voice Transcribe 2.0 features significant improvements in multilingual accuracy, supporting dozens of languages and automatically detecting the language during transcription. It also supports mid-recording language switches in a single pass.

The model is available via REST API and WebSocket for real-time streaming. Pricing remains identical to the previous version, with batch transcription at $0.10 per hour and streaming at $0.20 per hour. SpaceXAI stated that Grok Voice Transcribe 2.0 will soon become the default in its Speech-to-Text API, and the 1.0 version will be deprecated in the coming weeks.

Sources

  1. Grok Voice Transcribe 2.0 (Hacker News Frontpage, 2026-09-18)
  2. Documentation