ElevenLabs · 2026-09-28 · major
Eleven v4 — ElevenLabs' speech models cover 90+ languages and clone from 10s
Eleven v4 and Eleven v4 Turbo are ElevenLabs' new text-to-speech models. They support 90+ languages (up from 70), clone a voice from 10 seconds of audio and take stackable inline tags for emotion and tone. Live in the API now.

ElevenLabs' fourth-generation voices speak more languages, clone faster and take stage directions inline.
Key specs
| Turbo median inference | ~100 ms |
|---|
Quick facts
| Maker | ElevenLabs |
|---|---|
| Models | eleven_v4, eleven_v4_turbo |
| Languages | 90+ (up from 70) |
| Instant voice clone | 10 seconds of audio |
| Limit per request | 10,000 characters (about 10 min of audio) |
| Availability | ElevenAgents, ElevenCreative, ElevenAPI |
What is it?
Eleven v4 is ElevenLabs' most expressive text-to-speech model, released on September 28, 2026, together with a low-latency twin, Eleven v4 Turbo. The full model targets content such as audiobooks and dubbing, while Turbo is tuned for real-time voice agents. ElevenLabs says the biggest quality gains are in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
How does it work?
Delivery is steered with inline tags written into the script, such as [laughs] or [said angrily in French accent], and several tags can be stacked in order. The model keeps the surrounding text in context to adjust expression over longer passages, and TechCrunch reports it can start speaking as soon as an upstream language model begins producing text. Instant voice clones now need 10 seconds of audio, and Professional Voice Clones are supported.
Why does it matter?
For teams building voice agents, a Turbo model with about 100 ms median inference and wider language coverage means fewer pauses and fewer markets left out. ElevenLabs claims v4 ranked first on Artificial Analysis in September 2026 and won 65–81% of blind head-to-head preference tests; TechCrunch notes that 55% of the company's revenue now comes from enterprise customers, who are the main buyers of these agents.
Who is it for?
voice-agent builders, dubbing and audiobook teams
Frequently asked questions
- What is the difference between Eleven v4 and Eleven v4 Turbo?
- Eleven v4 is ElevenLabs' highest-quality text-to-speech model for content creation and long-form projects. Eleven v4 Turbo is the real-time variant for conversational voice agents, with a median inference latency of about 100 ms according to ElevenLabs' model docs. Both support 90+ languages and a 10,000-character limit per request.
- How much audio does Eleven v4 need to clone a voice?
- Eleven v4 creates an Instant Voice Clone from 10 seconds of audio, according to ElevenLabs. The model also supports Professional Voice Clones, and ElevenLabs says v4 gives noticeably better speaker similarity to the source voice and more consistent output across repeated generations and multi-character dialogue.
- Where can I use Eleven v4?
- Eleven v4 and Eleven v4 Turbo are live in ElevenAgents, ElevenCreative and the ElevenAPI, using the model IDs eleven_v4 and eleven_v4_turbo. ElevenLabs is also giving Creator plans and above three times the usual credits until October 12, 2026 to try the new models.
Try it
Set model_id to eleven_v4 or eleven_v4_turbo in the ElevenLabs text-to-speech API