AI/TLDR

Grok Voice Think Fast 2.0

xAI's speech-to-speech voice-agent model, announced July 2026 — audio in, audio out, with reasoning and tool calls running in the background.

Grok Voice (speech-to-speech)API onlyGenerally available
Released
29 Jul 2026
Input
$0.004 / text input
License
Proprietary
Coverage
1 story

Overview

Grok Voice Think Fast 2.0 is xAI's speech-to-speech model for voice agents, announced on 29 July 2026 and served under the API model id `grok-voice-think-fast-2.0`. xAI's docs also expose the alias `grok-voice-latest`, which always resolves to the newest voice model; xAI's release notes record that the alias began routing to 2.0 on 5 August 2026. Version 1.0 is not marked deprecated and remains callable by its own id.

The model works on a continuous session rather than discrete request/response turns: audio goes in and spoken audio comes back, with the model reasoning in the background so that deliberation does not show up as a pause before the first word. xAI reports 0.70 seconds to first audio, down from 1.25 seconds for Think Fast 1.0, while spending roughly 0.4× the reasoning tokens per response of the earlier version.

It is built for agents that actually do work during the call. The Speech to Speech API supports five tool types — Collections search (`file_search`), `web_search`, `x_search`, MCP servers, and custom function tools with JSON schemas — and xAI executes the server-side tools (web search, X search, collections and MCP) itself. The docs list 20+ supported languages, including several Arabic and Spanish variants, Bengali, Chinese, Hindi, Indonesian, Japanese, Korean, Portuguese (Brazil and Portugal), Russian, Turkish and Vietnamese. Output voices come from the same roster as the Text to Speech API, and custom voices can be built through the Custom Voices API.

On transcription accuracy xAI reports a 1.5–2.0× word-error-rate improvement over Deepgram Nova 3 and ElevenLabs Scribe v2 across 24 languages, and a 1.4× improvement over Think Fast 1.0, with the gap widening in noisy conditions. Pricing is metered by audio time: $0.08 per minute of audio ($4.80 per hour) plus $0.004 for text input. Session history is dropped after 30 minutes of inactivity.

Released2026-07-29
LicenseProprietary
WeightsAPI only
ModalitiesAudio, Text
StatusGenerally available

Benchmarks

Grok Voice Think Fast 2.0 vs named peers, as published by xAI (figures credited to Artificial Analysis)

BenchmarkGrok Voice Think Fast 2.0Grok Voice Think Fast 1.0GPT-Realtime-2.1 (High)Gemini 3.1 Flash (High)
AA Speech-to-Speech Quality Index82.975.779.169.5
Speech Reasoning (Big Bench Audio)97.297.19696.6
Conversational Dynamics (Full Duplex Bench)95.177.895.774.3
Agentic Performance (τ-voice Bench)56.552.145.737.7
Time to First Audio0.7 s1.25 s2.98 s

Comparison source ↗

This model's scores

  1. AA Speech-to-Speech Quality Index82.9%
  2. Speech Reasoning (Big Bench Audio)97.2%
  3. Conversational Dynamics (Full Duplex Bench)95.1%
  4. Agentic Performance (τ-voice Bench)56.5%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$0.004 / text input

Audio is billed by time: $0.08 per minute ($4.80 per hour). Speech to Text is $0.10/hr (REST) or $0.20/hr (streaming); Text to Speech is $15.00 per 1M characters.

Pricing source ↗

Strengths

  • 0.70 seconds to first audio — xAI's published figure, against 1.25s for Think Fast 1.0 and 2.98s for Gemini 3.1 Flash
  • Reasoning happens in the background during the call, at roughly 0.4× the reasoning tokens per response of version 1.0
  • Five tool types in-session (collections search, web search, X search, MCP, custom functions), with server-side tools executed by xAI
  • Leads xAI's published agentic voice comparison at 56.5% on τ-voice Bench, ahead of GPT-Realtime-2.1 at 45.7%
  • Reported 1.5–2.0× lower word error rate than Deepgram Nova 3 and ElevenLabs Scribe v2 across 24 languages

Best for

  • Reach for it for phone-based support and sales agents that have to look things up or write to a system mid-conversation.
  • Reach for it when latency is the product — a noticeable pause before the first word reads as a broken line.
  • Reach for it for multilingual voice front ends across the 20+ languages the Speech to Speech API documents.
  • Reach for it when transcription accuracy under noise, accents and interruptions matters more than raw model size.

How to access

ProviderModel ID
xAI Speech to Speech API ↗grok-voice-think-fast-2.0

Grok Voice (speech-to-speech) — every version

The full lineage of the Grok Voice (speech-to-speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Grok Voice Think Fast 2.0current2026-07-29Proprietary
Grok Voice Think Fast 1.02026-04-23Proprietary

FAQ

What is the API model id for Grok Voice Think Fast 2.0?

`grok-voice-think-fast-2.0`. xAI's Speech to Speech docs also publish the alias `grok-voice-latest`, which always points at the newest voice model, so pinning the explicit id is what keeps a deployment on 2.0 specifically.

How much does Grok Voice Think Fast 2.0 cost?

xAI's pricing page lists $0.08 per minute of audio, which it also states as $4.80 per hour, plus $0.004 for text input. The separate Speech to Text API is $0.10/hr over REST and $0.20/hr streaming, and Text to Speech is $15.00 per million characters.

How does Grok Voice Think Fast 2.0 compare with GPT-Realtime and Gemini Live?

In the comparison xAI published at launch, crediting Artificial Analysis, it scores 82.9% on the AA Speech-to-Speech Quality Index against 79.1% for GPT-Realtime-2.1 (High) and 69.5% for Gemini 3.1 Flash (High), and 56.5% on τ-voice Bench against 45.7% and 37.7%. GPT-Realtime-2.1 edges it on Full Duplex Bench, 95.7% to 95.1%.

How fast does it start speaking?

xAI reports 0.70 seconds to first audio, against 1.25 seconds for Grok Voice Think Fast 1.0 and 2.98 seconds for Gemini 3.1 Flash (High). It reaches that while using about 0.4× the reasoning tokens per response that version 1.0 used.

Can it call tools during a live call?

Yes. The Speech to Speech API supports Collections search (`file_search`), `web_search`, `x_search`, MCP servers, and custom function tools defined with JSON schemas. xAI executes the server-side tools — web search, X search, collections and MCP — on its own infrastructure.

Which languages does it support?

xAI's Speech to Speech documentation lists 20+ languages, including multiple Arabic variants, Bengali, Simplified Chinese, English, French, German, Hindi, Indonesian, Italian, Japanese, Korean, Brazilian and European Portuguese, Russian, Mexican and European Spanish, Turkish and Vietnamese. The transcription-accuracy comparison in the launch post was measured across 24 languages.