Overview
Mercury Voice is Inception's diffusion language model tuned to power voice agents, made generally available to enterprise customers on 29 September 2026 after a preview announced alongside Mercury 2.5. It is a text model that sits in the LLM slot of a voice pipeline: it reasons, calls tools, returns structured outputs and follows long system prompts, served through an OpenAI-compatible chat-completions endpoint so it plugs into LiveKit, Pipecat, Vapi, Retell or a custom voice stack.
Inception's metric for the release is time to first answer token (TTFAT) — how long until the model starts producing the words the caller hears, after all reasoning is done — measured against a conversational budget of about 500 ms. On OpenCall customer-service prompts, Inception reports 320 ms median (p50) and 750 ms p95 TTFAT at low reasoning effort, which it describes as 5.9× faster than GPT-6 Luna (no reasoning) and the only model in its test with a median under the 500 ms budget.
On quality, Inception publishes an average across τ³-bench Telecom, Retail and Airline, IFBench and BFCL v4 multi-turn: 70.5 for Mercury Voice at low effort, against 59.0 for Gemma 4 31B (non-reasoning) and 51.4 for GPT-4.1, the non-reasoning models it names as the usual defaults for voice agents. These are Inception's own figures from the launch post; no independent evaluation is cited.
Pricing is $0.40 per million input tokens and $1.50 per million output tokens, with a 50% launch discount to $0.20 / $0.75. Inception estimates that this works out to about $0.009 per minute of conversation for a typical voice-agent profile, compared with $0.045 per minute for GPT-4.1. Weights are not published.
| Released | 2026-09-29 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Context | 128K |
| Max output | 50K |
| Architecture | Diffusion language model (dLLM) tuned for voice agents, with low, medium and high reasoning-effort settings |
| Modalities | Text |
| Status | Generally available to enterprise customers through the OpenAI-compatible Inception API; access is enabled per organization by Inception's sales team. Previewed alongside Mercury 2.5 in September 2026. |
Benchmarks



Mercury Voice (low effort) vs the non-reasoning models Inception names as voice-agent defaults
| Benchmark | Mercury Voice (Low) | Gemma 4 31B (non-reasoning) | GPT-4.1 |
|---|---|---|---|
| OpenCall TTFAT p50 / p95 | 320 ms / 750 ms | 1.24 s / 3.51 s | 870 ms / 1.45 s |
| Average quality | 70.5 | 59 | 51.4 |
| τ³-bench Telecom | 77.2 | 50 | 48.2 |
| τ³-bench Airline | 78 | 66 | 54 |
| τ³-bench Retail | 69.3 | 69.3 | 64 |
| IFBench | 67.4 | 54.1 | 45.1 |
| BFCL V4 multi-turn | 60.5 | 55.6 | 45.5 |
| Price ($/1M in · out) | $0.40 · $1.50 | Varies by provider | $2.00 · $8.00 |
This model's scores
- Average quality (composite, low effort)70.5
- τ³-bench Telecom (low effort)77.2%
- τ³-bench Airline (low effort)78%
- τ³-bench Retail (low effort)69.3%
- IFBench (low effort)67.4%
- BFCL V4 multi-turn (low effort)60.5%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.40 / 1M tokens |
|---|---|
| Cached input | $0.04 / 1M tokens |
| Output | $1.50 / 1M tokens |
Inception launched Mercury Voice at 50% off: $0.20 input, $0.02 cached input and $0.75 output per million tokens.
Strengths
- 320 ms median and 750 ms p95 time to first answer token on OpenCall customer-service prompts at low reasoning effort, per Inception
- Three reasoning-effort settings — low, medium and high — so a voice agent can trade latency for depth per deployment
- Tool calling and structured outputs, with long system prompts, for agents that act during a call
- 128K-token context window with up to 50,000 output tokens
- OpenAI-compatible endpoint that drops into LiveKit, Pipecat, Vapi, Retell or a custom voice stack
- Published price of $0.40 input / $1.50 output per million tokens, which Inception puts at about $0.009 per minute of conversation
Best for
- Customer-service phone agents, where Inception's TTFAT measurements were taken on real OpenCall prompts

- Automated drive-thru ordering — Audivi AI uses it for real-time orders, order modifications and upsells
- Financial-services calls such as negotiating payment plans and handling objections, as Altur does
- Appointment scheduling and other voice workflows that need tool calls without an audible pause
How to access

| Provider | Model ID |
|---|---|
| Inception API (enterprise) ↗ | — |
FAQ
Does Mercury Voice take audio input?
No. Inception's docs list it as a text model on the v1/chat/completions endpoint. It fills the LLM slot of a voice pipeline — between speech-to-text and text-to-speech — in platforms such as LiveKit, Pipecat, Vapi and Retell.
How fast is Mercury Voice?
Inception reports a 320 ms median (p50) and 750 ms p95 time to first answer token on OpenCall customer-service prompts at low reasoning effort; at medium effort the figures are 550 ms and 1.96 s, and at high effort 1.11 s and 4.84 s. OpenCall reported median model response latency close to 170 ms on its own production workload.
What does Mercury Voice cost?
$0.40 per million input tokens, $0.04 per million cached input tokens and $1.50 per million output tokens. At launch Inception discounted it 50%, to $0.20, $0.02 and $0.75. Inception estimates about $0.009 per minute of conversation for a typical voice-agent profile.
How do I get access to Mercury Voice?
It is available to enterprise customers through the OpenAI-compatible Inception API. Inception's docs say to contact sales@inceptionlabs.ai to enable it for an organization.
What benchmarks did Inception publish for Mercury Voice?
At low effort: 77.2 on τ³-bench Telecom, 78.0 on τ³-bench Airline, 69.3 on τ³-bench Retail, 67.4 on IFBench and 60.5 on BFCL V4 multi-turn, for an average quality score of 70.5. Inception compares these with Gemma 4 31B (non-reasoning) at 59.0 and GPT-4.1 at 51.4 on the same average.