Eleven v4: What Has Changed in ElevenLabs Text-to-Speech and How to Control Emotion
AI News

Eleven v4: What Has Changed in ElevenLabs Text-to-Speech and How to Control Emotion

01 Oct 2026
Explore Eleven v4 and v4 Turbo: Audio Tags, language support, latency, voice cloning, and how to choose a model for voiceovers and AI agents.

On September 28, 2026, ElevenLabs introduced Eleven v4 and Eleven v4 Turbo. Both models turn text into speech, but they serve different purposes: v4 is designed for expressive voiceovers, while Turbo targets interactive applications and voice agents. The update’s main promise is more precise emotional delivery and a consistent voice across long scenes and dialogues. These are the developer’s claims; you should evaluate quality on your own scripts separately.

What Is New in Eleven v4

In its announcement, ElevenLabs highlights improvements in handling context, pacing, and the character of spoken lines. The model is intended to preserve a speaker’s recognizable voice more reliably across repeated generations and transitions between segments. Support for Professional Voice Clones has been added, and the company says joining separate generations into a long recording has become more reliable. These changes are particularly relevant to creators of audiobooks, videos, and game dialogue: transitions between lines are often where inconsistencies in synthetic speech become noticeable.

The documentation lists support for more than 90 languages across the v4 family. Standard v4 has a stated limit of 10,000 characters and supports dialogue with multiple speakers. These are technical capabilities of the service, rather than a promise of identical pronunciation and expressiveness in every language. Before producing a large batch of videos, it is worth checking names, word stress, numbers, and specialist terms in the language you plan to publish in.

Audio Tags: How to Set the Emotion in a Voiceover

Audio Tags are short instructions in square brackets within a script. For example, [whispers] requests a whisper, while [excited] helps indicate an enthusiastic delivery. Place a tag before the phrase you want delivered that way. According to ElevenLabs’ guide, the direction continues through the line until you specify another one. You can combine characteristics, such as [whispering, playful].

There is no fixed, exhaustive list of commands: the service supports delivery descriptions in ordinary language. This does not mean that every wording will always produce the same result. Start with one clear instruction, listen to the output, and only then add refinements. Asking for several conflicting emotions at once makes it harder to identify which part of the instruction affected the voice.

A short line works well for an initial test: “[excited] The new collection is here. [whispers] And we saved this version for those paying close attention.” This is a sample script for an experiment, not a finished recording: how it sounds depends on the selected voice and the generation. Compare versions with and without tags using exactly the same text.

ElevenLabs also explains that a tag is not necessary on every line: v4 can infer emotional delivery from context. Instructions are more useful at points where the intonation needs to change noticeably. Punctuation and sentence length also affect rhythm. It therefore makes sense to get the script into shape first, then direct its performance.

How Eleven v4 Turbo Differs from v4

Turbo is aimed at use cases where a spoken response needs to arrive quickly. The company reports median inference latency of around 100 ms. The same announcement includes another metric: around 150 ms to the first audible speech in a comparative test. These figures measure different stages and do not mean that a complete voice assistant always responds in a tenth of a second.

The documentation clarifies that model latency excludes application and network overhead. In a real assistant, speech recognition, response preparation, and calls to external tools add time. Choosing Turbo on the basis of a single number in an announcement is therefore insufficient. For a video, the quality of the final audio track matters more; for a conversation, it is the entire pause between a person’s question and the start of the response.

Developers should check the integration interface: the September 28 changelog lists the Text to Dialogue API for eleven_v4 and Text to Dialogue WebSocket for eleven_v4_turbo. The arrival of a new model does not mean its identifier can be inserted into any existing request without other changes.

How to Test the Model Before a Real Project

Put together a small set of excerpts: a neutral explanation, an emotional line, a dialogue between two characters, and a paragraph with difficult names. Save the text, voice, and settings for each one. Then assess not only how pleasant the output sounds, but also missing words, word stress, unwanted sounds, and whether a successful delivery can be reproduced.

If your task involves cloning, use your own voice or a recording you have permission to use. For a commercial project, separately check the terms of your chosen plan, rights to the source material, and the cost of repeated generations. Free registration alone does not explain every restriction on use.

According to ElevenLabs, both models are available in ElevenCreative, ElevenAgents, and through ElevenAPI. The update is worth testing if you need controllable emotion or natural dialogue. A decision to switch is best made after a comparison using your own script: promotional demos help illustrate the possibilities, but do not replace checking the finished voiceover.