█

AI/TLDR

Eleven v4

ElevenLabs' expressive text-to-speech model, released 28 September 2026 — a new architecture, 90+ languages, inline audio tags, and voice cloning from 10 seconds of audio.

Eleven (text-to-speech)API onlyGenerally available
Released
28 Sep 2026
Parameters
Not disclosed
Input
$0.08 / 1K characters
License
Proprietary
Coverage
1 story

Overview

Eleven v4 is the text-to-speech model ElevenLabs launched on 28 September 2026, alongside the low-latency Eleven v4 Turbo. It is served in the ElevenLabs API as `eleven_v4` and is also available in ElevenAgents and ElevenCreative. ElevenLabs built it on an entirely new architecture and aims it at content where delivery matters: audiobooks, character voiceovers, dubbing and narration.

Delivery is directed in the script. Inline audio tags such as [laughs], [whispers], [said angrily in French accent] or [light rain] set emotion, accent and sound effects, and a line's delivery can also be described in natural language; ElevenLabs says v4 follows tag sequences more reliably than v3. SSML tags such as <break> are disabled in v4 in favour of tags like [pause]. IPA phoneme support was improved for custom pronunciations, and a single request takes up to 10,000 characters (about 10 minutes of audio), against 5,000 for Eleven v3.

The model family covers 90+ languages, up from 70+ for Eleven v3, and a voice recorded in one language can speak the others with a native accent while keeping its identity. Instant Voice Clones work from 10 seconds of audio, and Professional Voice Clones, which v3 did not support, are back. ElevenLabs says speaker identity stays stable across regenerations and multi-speaker scenes, and that request stitching for long-form content is more reliable. Every clone requires verified consent from the voice's owner.

In ElevenLabs' blind head-to-head preference tests (September 2026), listeners preferred Eleven v4 in 81% of pairs against Cartesia Sonic 3.6, 81% against Inworld TTS-2, 72% against Google Gemini 3.8 Flash-Lite TTS and 65% against Google Gemini 3.8 Flash TTS. ElevenLabs also reports a #1 rank on Artificial Analysis' Provider Voice Arena leaderboard in September 2026.

On the API pricing page Eleven v4 lists at $0.08 per 1,000 characters, with a launch discount to $0.022 per 1,000 characters until 12 October 2026. It uses the same credit pricing as ElevenLabs' other text-to-speech models, so it is available on every plan, including the free tier's 10,000 monthly credits.

Released2026-09-28
LicenseProprietary
WeightsAPI only
ParametersNot disclosed
Max output10,000 characters per request (~10 minutes of audio)
ArchitectureNot disclosed (ElevenLabs describes an entirely new architecture relative to Eleven v3)
ModalitiesText, Audio
StatusGenerally available

Benchmarks

ElevenLabs' bar chart of the share of blind head-to-head pairs won by Eleven v4: 81% against Cartesia Sonic 3.6, 81% against Inworld TTS-2, 72% against Google Gemini 3.8 Flash-Lite TTS and 65% against Google Gemini 3.8 Flash TTS.
Share of blind head-to-head preference pairs won by Eleven v4 (ties counted as half), as published by ElevenLabs (September 2026). — ElevenLabs

This model's scores

  1. Blind preference win rate vs Cartesia Sonic 3.681%
  2. Blind preference win rate vs Inworld TTS-281%
  3. Blind preference win rate vs Gemini 3.8 Flash-Lite TTS72%
  4. Blind preference win rate vs Gemini 3.8 Flash TTS65%

Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.

Pricing

Input$0.08 / 1K characters

ElevenLabs API list price for Eleven v4 text-to-speech; discounted 72% to $0.022 per 1K characters until 12 October 2026. Plans range from Free (10,000 characters included) to Business ($990/month).

Pricing source ↗

Strengths

  • Preferred in 65–81% of blind head-to-head pairs against Cartesia Sonic 3.6, Inworld TTS-2 and Google's Gemini 3.8 Flash and Flash-Lite TTS in ElevenLabs' September 2026 tests
  • Inline audio tags for emotion, accent, pauses and sound effects, followed more reliably than in Eleven v3
  • 90+ languages, with a cloned voice speaking each one in a native accent
  • Instant Voice Clones from 10 seconds of audio, plus Professional Voice Clones
  • 10,000 characters per request (about 10 minutes of audio), double Eleven v3's 5,000
  • Every voice in ElevenLabs' 17,500+ voice library works with it

Best for

  • Reach for it for audiobooks and long-form narration, where stable speaker identity and reliable request stitching keep a production consistent
  • Reach for it for game, animation and character voiceovers that need a wide emotional range and directed delivery
  • Reach for it for dubbing and localisation that keeps one brand or actor voice across 90+ languages
  • Reach for Eleven v4 Turbo instead for voice agents and other real-time use, where latency matters more than studio-grade output

How to access

VideoElevenLabs' launch video for Eleven v4 and Eleven v4 Turbo.ElevenLabs ↗
ProviderModel ID
ElevenLabs API ↗eleven_v4

Eleven (text-to-speech) — every version

The full lineage of the Eleven (text-to-speech) line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.

VersionReleasedContextLicense
Eleven v4current2026-09-28—Proprietary
Eleven v4 Turbo2026-09-28—Proprietary

FAQ

What is Eleven v4?

Eleven v4 is a text-to-speech model ElevenLabs launched on 28 September 2026. It is served in the ElevenLabs API as eleven_v4 and in ElevenAgents and ElevenCreative, and is built for expressive content such as audiobooks, character voiceovers and dubbing.

What changed from Eleven v3?

ElevenLabs lists a new architecture with higher audio quality and a wider emotional range, more reliable audio-tag following, stable speaker identity across regenerations, the return of Professional Voice Clones, 90+ languages instead of 70+, and a 10,000-character request limit instead of 5,000.

How much audio do I need to clone a voice?

Instant Voice Clones work from 10 seconds of audio. For the highest fidelity, Eleven v4 also supports Professional Voice Clones. Clones made before v4 launched need retraining with v4 to work effectively, and every clone requires verified consent from the voice's owner.

What does Eleven v4 cost?

The ElevenLabs API pricing page lists $0.08 per 1,000 characters, discounted to $0.022 until 12 October 2026. It uses the same credits as other ElevenLabs text-to-speech models, so it works on every plan including the free tier.

How does it differ from Eleven v4 Turbo?

Eleven v4 is tuned for produced content where quality matters most. Eleven v4 Turbo is the low-latency variant, with about 100 ms median inference latency and about 150 ms median time to first speech, for voice agents and real-time use. Both share the same expressive range and support Professional Voice Clones.