█

AI/TLDR

Mercury Voice

Inception's diffusion language model tuned for voice agents, generally available to enterprise customers from September 29, 2026 — 320 ms median time to first answer token on real customer-service prompts.

Mercury VoiceAPI onlyGenerally available to enterprise customers through the OpenAI-compatible Inception API; access is enabled per organization by Inception's sales team. Previewed alongside Mercury 2.5 in September 2026.
Released
29 Sep 2026
Context
128K
Input
$0.40 / 1M tokens
License
Proprietary

Overview

Mercury Voice is Inception's diffusion language model tuned to power voice agents, made generally available to enterprise customers on 29 September 2026 after a preview announced alongside Mercury 2.5. It is a text model that sits in the LLM slot of a voice pipeline: it reasons, calls tools, returns structured outputs and follows long system prompts, served through an OpenAI-compatible chat-completions endpoint so it plugs into LiveKit, Pipecat, Vapi, Retell or a custom voice stack.

Inception's metric for the release is time to first answer token (TTFAT) — how long until the model starts producing the words the caller hears, after all reasoning is done — measured against a conversational budget of about 500 ms. On OpenCall customer-service prompts, Inception reports 320 ms median (p50) and 750 ms p95 TTFAT at low reasoning effort, which it describes as 5.9× faster than GPT-6 Luna (no reasoning) and the only model in its test with a median under the 500 ms budget.

On quality, Inception publishes an average across τ³-bench Telecom, Retail and Airline, IFBench and BFCL v4 multi-turn: 70.5 for Mercury Voice at low effort, against 59.0 for Gemma 4 31B (non-reasoning) and 51.4 for GPT-4.1, the non-reasoning models it names as the usual defaults for voice agents. These are Inception's own figures from the launch post; no independent evaluation is cited.

Pricing is $0.40 per million input tokens and $1.50 per million output tokens, with a 50% launch discount to $0.20 / $0.75. Inception estimates that this works out to about $0.009 per minute of conversation for a typical voice-agent profile, compared with $0.045 per minute for GPT-4.1. Weights are not published.

Released2026-09-29
LicenseProprietary
WeightsAPI only
Context128K
Max output50K
ArchitectureDiffusion language model (dLLM) tuned for voice agents, with low, medium and high reasoning-effort settings
ModalitiesText
StatusGenerally available to enterprise customers through the OpenAI-compatible Inception API; access is enabled per organization by Inception's sales team. Previewed alongside Mercury 2.5 in September 2026.

Benchmarks

Bar chart of time to first answer token on OpenCall customer-service prompts, p50 and p95, for Mercury Voice at low, medium and high effort against Cerebras GPT OSS 120B, GPT-4.1, Gemini 3.5 Flash Lite, GLM-5.3 Flash, Gemma 4 31B, GPT-6 Luna, GPT-5.6 Luna, Qwen 3.5 397B and Claude Haiku 4.5, with a ~500 ms conversational budget line
Time to first answer token on real voice-agent prompts (OpenCall), p50 and p95 — Inception
Bar chart of average quality scores: Mercury Voice low 70.5, high 69.9 and medium 69 against GLM-5.3 Flash, Gemma 4 31B, Qwen 3.5 397B, GPT-6 Luna, Cerebras GPT OSS 120B, Gemini 3.5 Flash Lite, GPT-5.6 Luna, Claude Haiku 4.5 and GPT-4.1
Average quality across τ³-bench Telecom, Retail, Airline, IFBench and BFCL v4 — Inception
Scatter plot of average τ³-bench quality against average p50 time to first answer token on a log scale, with Mercury Voice (low) the only model inside the ~500 ms conversational budget
τ³-bench quality vs. latency (mean of Telecom, Retail and Airline) — Inception

Mercury Voice (low effort) vs the non-reasoning models Inception names as voice-agent defaults

BenchmarkMercury Voice (Low)Gemma 4 31B (non-reasoning)GPT-4.1
OpenCall TTFAT p50 / p95320 ms / 750 ms1.24 s / 3.51 s870 ms / 1.45 s
Average quality70.55951.4
τ³-bench Telecom77.25048.2
τ³-bench Airline786654
τ³-bench Retail69.369.364
IFBench67.454.145.1
BFCL V4 multi-turn60.555.645.5
Price ($/1M in · out)$0.40 · $1.50Varies by provider$2.00 · $8.00

Comparison source ↗

This model's scores

  1. Average quality (composite, low effort)70.5
  2. τ³-bench Telecom (low effort)77.2%
  3. τ³-bench Airline (low effort)78%
  4. τ³-bench Retail (low effort)69.3%
  5. IFBench (low effort)67.4%
  6. BFCL V4 multi-turn (low effort)60.5%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$0.40 / 1M tokens
Cached input$0.04 / 1M tokens
Output$1.50 / 1M tokens

Inception launched Mercury Voice at 50% off: $0.20 input, $0.02 cached input and $0.75 output per million tokens.

Pricing source ↗

Strengths

  • 320 ms median and 750 ms p95 time to first answer token on OpenCall customer-service prompts at low reasoning effort, per Inception
  • Three reasoning-effort settings — low, medium and high — so a voice agent can trade latency for depth per deployment
  • Tool calling and structured outputs, with long system prompts, for agents that act during a call
  • 128K-token context window with up to 50,000 output tokens
  • OpenAI-compatible endpoint that drops into LiveKit, Pipecat, Vapi, Retell or a custom voice stack
  • Published price of $0.40 input / $1.50 output per million tokens, which Inception puts at about $0.009 per minute of conversation

Best for

  • Customer-service phone agents, where Inception's TTFAT measurements were taken on real OpenCall prompts
VideoHow OpenCall powers its voice agents with MercuryInception ↗
  • Automated drive-thru ordering — Audivi AI uses it for real-time orders, order modifications and upsells
  • Financial-services calls such as negotiating payment plans and handling objections, as Altur does
  • Appointment scheduling and other voice workflows that need tool calls without an audible pause

How to access

VideoA Mercury Voice agent booking a dental appointment in a live conversationInception ↗

FAQ

Does Mercury Voice take audio input?

No. Inception's docs list it as a text model on the v1/chat/completions endpoint. It fills the LLM slot of a voice pipeline — between speech-to-text and text-to-speech — in platforms such as LiveKit, Pipecat, Vapi and Retell.

How fast is Mercury Voice?

Inception reports a 320 ms median (p50) and 750 ms p95 time to first answer token on OpenCall customer-service prompts at low reasoning effort; at medium effort the figures are 550 ms and 1.96 s, and at high effort 1.11 s and 4.84 s. OpenCall reported median model response latency close to 170 ms on its own production workload.

What does Mercury Voice cost?

$0.40 per million input tokens, $0.04 per million cached input tokens and $1.50 per million output tokens. At launch Inception discounted it 50%, to $0.20, $0.02 and $0.75. Inception estimates about $0.009 per minute of conversation for a typical voice-agent profile.

How do I get access to Mercury Voice?

It is available to enterprise customers through the OpenAI-compatible Inception API. Inception's docs say to contact sales@inceptionlabs.ai to enable it for an organization.

What benchmarks did Inception publish for Mercury Voice?

At low effort: 77.2 on τ³-bench Telecom, 78.0 on τ³-bench Airline, 69.3 on τ³-bench Retail, 67.4 on IFBench and 60.5 on BFCL V4 multi-turn, for an average quality score of 70.5. Inception compares these with Gemma 4 31B (non-reasoning) at 59.0 and GPT-4.1 at 51.4 on the same average.