Previous text-to-speech models were limited in expressiveness, not sound quality. Specifically:
Eleven v3 was built from the ground up to produce speech that:
Sighs, whispers, laughs, and reacts — feeling genuinely alive
v3 is not just an upgrade in clarity — it's a fundamental shift toward emotional realism
Audio tags are written instructions embedded directly in your script that tell the model how to deliver specific words or phrases.
| Tag | Effect |
|---|---|
[whispers] | Delivers text in a whisper |
[sighs] | Adds a sigh sound |
[excited] | Raises energy and enthusiasm |
[whispers] Something's coming… [sighs] I can feel it.
You can also combine multiple tags for layered emotional control throughout a script.
Audio tags give you director-level control over tone without recording a human actor
A brand new capability allowing you to generate conversations between multiple speakers in a single audio output.
Instead of stitching together separate audio clips, v3 generates one cohesive, natural-sounding conversation
Eleven v3 supports 70+ languages, covering high-demand global languages comprehensively.
Combined with expressiveness features, this means:
| Use Case | Why v3 Works |
|---|---|
| Audiobooks | Long-form, expressive narration |
| Film/Video production | Cinematic emotional range |
| Game development | Character dialogue with personality |
| Immersive storytelling | Multi-speaker, reactive speech |
| Limitation | Impact |
|---|---|
| Higher latency | Not suitable for real-time use |
| Requires prompt engineering | Less reliable without careful scripting |
| PVCs not fully optimized | Lower clone quality currently |
For real-time and conversational use cases → use v2.5 Turbo or Flash
v3 trades speed and simplicity for depth and expressiveness — choose based on your use case
| User Type | Discount |
|---|---|
| Self-serve UI | 80% off (~5× cheaper) |
| Enterprise UI | 80% off business plan pricing |
Pricing returns to standard Multilingual V2 rates.
The alpha period is the best time to experiment with v3 at minimal cost
Step 1 → Log in to ElevenLabs UI
Step 2 → Select "Eleven v3 (alpha)" in the model dropdown
Step 3 → Paste your script with [audio tags] or dialogue JSON
Step 4 → Generate and refine
Eleven v3 (alpha)
│
├── Audio Tags ──────── Emotional inline control
├── Dialogue Mode ───── Multi-speaker conversations
├── 70+ Languages ───── Global expressiveness
└── Use Case Fit
├── YES → Audiobooks, Film, Games
└── NO → Real-time, Conversational AI
The central lesson is that v3 prioritizes expressiveness over speed, making it a powerful tool for creative production — but one that rewards careful scripting and prompt engineering.