Contents
The same text sounds different depending on the voice, tags and punctuation. If the delivery is only described as "make it emotional", the model guesses; if the text is marked up, it follows the markup.
In Eleven v4 the delivery comes from three things: the voice, the text with tags, and two settings. The voice matters most: it sets the timbre, pace and manner, and tags and settings work on top of it.
Eleven v4 is the default model in new projects, so this guide is written for it. Tags and punctuation work the same way in Eleven v3; the differences are in a separate section.
What you need
- The line or script you want to voice.
- The Voiceover mode in Generation with Eleven v4 selected.
Step by step
- 1
Pick a voice for the manner
Preview several voices on one line.
- 2
Prepare the text
Numbers and abbreviations as words, short phrases, punctuation in place.
- 3
Place the tags
A tag before the phrase where the emotion changes; a detailed tag at the start of a paragraph for pace.
- 4
Draft in Eleven v4 Turbo
Check the pace and tags: Turbo costs half as much as v4.
- 5
Voice the final take in Eleven v4
Use the same tags; if a take does not work, generate another one.
Start with the voice
A voice repeats the manner of the recording it was made from. An upbeat voice sounds upbeat more easily, a calm one sounds calm. Eleven v4 can perform delivery the recording never had, such as a whisper from a newsreader voice, but tags work most reliably when the delivery is close to the voice.
Compare voices on the same text
Preview two or three voices on the same line with tags and pick the one closest to the manner you need.
Your own voice from a clean recording
Eleven v4 reproduces the recording a voice was made from, including noise, clicks and volume jumps. For a New voice, use 1–2 minutes of clean speech in one manner.
Another language, a native accent
In another language, a voice in Eleven v4 speaks with that language's native accent. To add an accent, try a tag like [strong French accent].
Tags: how to set emotion and delivery
A tag is a short instruction in square brackets where the delivery should change. Type “[” in the voiceover field and Givon suggests 12 tags. The list is open: any description of a sound of up to 40 characters in brackets goes to the model.
ElevenLabs writes all its tag examples in English, so we recommend English tags too. The voiceover text itself can be in any language.
Emotions
[excited], [sad], [angry], [curious], [sarcastic], [nervous], [calm]
Manner
[whispers], [shouting], [softly], [flatly], [deadpan], [cheerfully]
Reactions
[laughs], [sighs], [exhales], [clears throat], [gasps], [crying]
Pauses and pace
[pause], [long pause], [rushed], [slows down], [hesitates]
Sounds
[applause], [door slams], [phone buzzing], [light rain]
Experimental
[sings], [strong French accent]: not every voice handles them
Where to put tags
Before or right after the phrase
[annoyed] This is hard. — annoyance on the whole phrase. This is hard. [sighs] — a sigh after it.
Stacked, to build an emotion
[hesitant] [nervous] I... I'm not sure. Stack tags or join them with a comma in one pair of brackets: [excited, happy].
A change of emotion inside the text
[hesitant] I didn't mean to say that. [regretful] It just came out. Each new tag is a new turn in tone.
A direction tag at the start of a paragraph
A detailed description sets the pace and tone for the whole paragraph: [Quiet, measured narration], [Gradually building energy], [Low, steady voice, restrained urgency].
I'm fine. [flatly] I'm fine. [quietly, after a pause] I'm fine. [angrily, fed up] I'm FINE.
The same words, four different answers.
Punctuation and structure
The text shapes delivery as much as tags do. Eleven v4 reads the whole text and takes tone from context: exclamations, questions, remarks like "she said, her voice trembling".
- Ellipsis…
- A pause and weight, sometimes with a hint of hesitation
- CAPITALS
- Stress on a word: "I need it NOW"
- A dash at the end of a line
- A cut-off phrase
- A new line
- A longer pause than after a period
- Commas
- Breathing and rhythm inside a phrase
Numbers, abbreviations and pronunciation
Numbers as words
ElevenLabs advises against digits and symbols, especially in multilingual models: "twenty percent" is more reliable than "20%". The same goes for dates, times, phone numbers and amounts.
Abbreviations in full
"kg" becomes "kilograms", "St." becomes "Street", and a website address is spelled out. Remove emojis and rare symbols.
Difficult words in the Glossary
Add a name, brand or term the voice gets wrong to Library → Voices → Glossary: the rule applies to every voiceover generated from text.
Transcription in Eleven v4
Eleven v4 understands IPA between slashes, for example /ˌbaɪoʊˈkemɪstri/, with stress marks. Check the result on your voice.
Expressiveness and Similarity in Eleven v4
The settings set a range, not an exact result: every generation is a new take. The defaults of 50% and 75% are usually enough, and delivery is easier to change with tags.
- Higher Expressiveness
- Livelier and more expressive, but takes differ more: make a few and pick the best
- Lower Expressiveness
- Steady and predictable, good for instructions and long text
- Higher Similarity
- Closer to the source voice, but flaws of the recording repeat too
Templates by task
[excited] The new collection is live! [whispers] And the first hundred buyers get twenty percent off. [laughs] Grab yours before it's gone.
[curious] Why does coffee after lunch wake you up less than in the morning? [pause] It comes down to cortisol. [low, steady voice] In the morning your level is already high...
[sighs] Here we go again... [sarcastic] Sure, that was a VERY good idea. [exhales] Fine. [curious] So what is it this time?
[professional] Thank you for reaching out. [sympathetic] I understand how frustrating this is. [reassuring] We will refund you within three days.
[Quiet, reflective narration] The city was still asleep... [Building tension, measured pace] The footsteps behind grew closer. [Gentle pause, then quiet realization] It was him.
What is different in Eleven v3
Tags, punctuation and writing advice are the same in Eleven v3. The settings and the limit differ.
Delivery instead of settings
v3 has four presets: Natural for balanced narration, Calm for steadier, slower explanations, Expressive for more emotion with controlled diction, and Ad for energetic short ad copy.
Tags and presets
The steadier the preset, the less the voice responds to tags. For tag-heavy text, use Expressive or Natural.
The voice matters more than tags
v3 depends more on the voice: a voice recorded loudly whispers poorly on a tag. Pick a voice with the manner you need.
Up to 4,000 characters
v3 takes twice as much text per request as v4, which is handy for long narration in one piece.
What to avoid in tags
Actions you cannot hear
[standing], [grinning] and [pacing] do nothing: a tag must describe a sound, not a pose or a facial expression.
A tag that sounds like a sound effect
Eleven v4 does both speech and sounds, so [crying] sometimes comes out as the sound of crying. Describe the voice: [trembling, tearful voice], [low, gravelly voice].
Delivery against the voice's character
A meditation voice shouts unconvincingly, and a strict business voice laughs poorly on [giggles]. Pick a voice for the task.
Repeating what the text already says
If the text says "he laughed out loud", do not replace it with a tag: the model reads the phrase, and you can add [chuckles] next to it.
FAQ
Can I write tags in another language?
The model gets the tag as is, but ElevenLabs gives all its examples in English. English tags work more reliably, while the voiceover text can be in any language.
How many tags can I use?
There is no limit, but tags count toward the character limit: 2,000 in Eleven v4 and 4,000 in Eleven v3. One tag per change of emotion is usually enough.
Why was a tag read aloud or played as a sound?
Most often the delivery does not fit the voice, or the tag looks like a sound effect. Describe the voice in words, for example [trembling, tearful voice], or pick a voice with a wider range.
How do I add a pause?
With [pause] or [long pause], an ellipsis or a new line. Eleven v4 and v3 do not support SSML tags such as <break>.
Why do two takes sound different?
Every generation is a new performance, like an actor's take. The higher the Expressiveness, the bigger the difference. Lower it for a steadier result.


