How to write text for voiceover: tags and delivery in Eleven v4 and v3

How to pick a voice, place emotion tags and pauses, set Expressiveness and Similarity, and write text so Eleven v4 reads it with the delivery you want. Plus what works differently in Eleven v3.
Contents
  1. Step by step
  2. Start with the voice
  3. Tags: how to set emotion and delivery
  4. Where to put tags
  5. Punctuation and structure
  6. Numbers, abbreviations and pronunciation
  7. Expressiveness and Similarity in Eleven v4
  8. Templates by task
  9. What is different in Eleven v3
  10. What to avoid in tags
  11. FAQ

The same text sounds different depending on the voice, tags and punctuation. If the delivery is only described as "make it emotional", the model guesses; if the text is marked up, it follows the markup.

In Eleven v4 the delivery comes from three things: the voice, the text with tags, and two settings. The voice matters most: it sets the timbre, pace and manner, and tags and settings work on top of it.

Eleven v4 is the default model in new projects, so this guide is written for it. Tags and punctuation work the same way in Eleven v3; the differences are in a separate section.

What you need

  • The line or script you want to voice.
  • The Voiceover mode in Generation with Eleven v4 selected.

Step by step

  1. 1

    Pick a voice for the manner

    Preview several voices on one line.

  2. 2

    Prepare the text

    Numbers and abbreviations as words, short phrases, punctuation in place.

  3. 3

    Place the tags

    A tag before the phrase where the emotion changes; a detailed tag at the start of a paragraph for pace.

  4. 4

    Draft in Eleven v4 Turbo

    Check the pace and tags: Turbo costs half as much as v4.

  5. 5

    Voice the final take in Eleven v4

    Use the same tags; if a take does not work, generate another one.

Start with the voice

A voice repeats the manner of the recording it was made from. An upbeat voice sounds upbeat more easily, a calm one sounds calm. Eleven v4 can perform delivery the recording never had, such as a whisper from a newsreader voice, but tags work most reliably when the delivery is close to the voice.

  • Compare voices on the same text

    Preview two or three voices on the same line with tags and pick the one closest to the manner you need.

  • Your own voice from a clean recording

    Eleven v4 reproduces the recording a voice was made from, including noise, clicks and volume jumps. For a New voice, use 1–2 minutes of clean speech in one manner.

  • Another language, a native accent

    In another language, a voice in Eleven v4 speaks with that language's native accent. To add an accent, try a tag like [strong French accent].

Tags: how to set emotion and delivery

A tag is a short instruction in square brackets where the delivery should change. Type “[” in the voiceover field and Givon suggests 12 tags. The list is open: any description of a sound of up to 40 characters in brackets goes to the model.

ElevenLabs writes all its tag examples in English, so we recommend English tags too. The voiceover text itself can be in any language.

  • Emotions

    [excited], [sad], [angry], [curious], [sarcastic], [nervous], [calm]

  • Manner

    [whispers], [shouting], [softly], [flatly], [deadpan], [cheerfully]

  • Reactions

    [laughs], [sighs], [exhales], [clears throat], [gasps], [crying]

  • Pauses and pace

    [pause], [long pause], [rushed], [slows down], [hesitates]

  • Sounds

    [applause], [door slams], [phone buzzing], [light rain]

  • Experimental

    [sings], [strong French accent]: not every voice handles them

Where to put tags

  • Before or right after the phrase

    [annoyed] This is hard. — annoyance on the whole phrase. This is hard. [sighs] — a sigh after it.

  • Stacked, to build an emotion

    [hesitant] [nervous] I... I'm not sure. Stack tags or join them with a comma in one pair of brackets: [excited, happy].

  • A change of emotion inside the text

    [hesitant] I didn't mean to say that. [regretful] It just came out. Each new tag is a new turn in tone.

  • A direction tag at the start of a paragraph

    A detailed description sets the pace and tone for the whole paragraph: [Quiet, measured narration], [Gradually building energy], [Low, steady voice, restrained urgency].

One phrase, different meanings

I'm fine. [flatly] I'm fine. [quietly, after a pause] I'm fine. [angrily, fed up] I'm FINE.

The same words, four different answers.

Punctuation and structure

The text shapes delivery as much as tags do. Eleven v4 reads the whole text and takes tone from context: exclamations, questions, remarks like "she said, her voice trembling".

Ellipsis…
A pause and weight, sometimes with a hint of hesitation
CAPITALS
Stress on a word: "I need it NOW"
A dash at the end of a line
A cut-off phrase
A new line
A longer pause than after a period
Commas
Breathing and rhythm inside a phrase

Numbers, abbreviations and pronunciation

  • Numbers as words

    ElevenLabs advises against digits and symbols, especially in multilingual models: "twenty percent" is more reliable than "20%". The same goes for dates, times, phone numbers and amounts.

  • Abbreviations in full

    "kg" becomes "kilograms", "St." becomes "Street", and a website address is spelled out. Remove emojis and rare symbols.

  • Difficult words in the Glossary

    Add a name, brand or term the voice gets wrong to Library → Voices → Glossary: the rule applies to every voiceover generated from text.

  • Transcription in Eleven v4

    Eleven v4 understands IPA between slashes, for example /ˌbaɪoʊˈkemɪstri/, with stress marks. Check the result on your voice.

Expressiveness and Similarity in Eleven v4

The settings set a range, not an exact result: every generation is a new take. The defaults of 50% and 75% are usually enough, and delivery is easier to change with tags.

Higher Expressiveness
Livelier and more expressive, but takes differ more: make a few and pick the best
Lower Expressiveness
Steady and predictable, good for instructions and long text
Higher Similarity
Closer to the source voice, but flaws of the recording repeat too

Templates by task

Ad

[excited] The new collection is live! [whispers] And the first hundred buyers get twenty percent off. [laughs] Grab yours before it's gone.

Explainer

[curious] Why does coffee after lunch wake you up less than in the morning? [pause] It comes down to cortisol. [low, steady voice] In the morning your level is already high...

Character

[sighs] Here we go again... [sarcastic] Sure, that was a VERY good idea. [exhales] Fine. [curious] So what is it this time?

Business voice

[professional] Thank you for reaching out. [sympathetic] I understand how frustrating this is. [reassuring] We will refund you within three days.

Long story

[Quiet, reflective narration] The city was still asleep... [Building tension, measured pace] The footsteps behind grew closer. [Gentle pause, then quiet realization] It was him.

What is different in Eleven v3

Tags, punctuation and writing advice are the same in Eleven v3. The settings and the limit differ.

  • Delivery instead of settings

    v3 has four presets: Natural for balanced narration, Calm for steadier, slower explanations, Expressive for more emotion with controlled diction, and Ad for energetic short ad copy.

  • Tags and presets

    The steadier the preset, the less the voice responds to tags. For tag-heavy text, use Expressive or Natural.

  • The voice matters more than tags

    v3 depends more on the voice: a voice recorded loudly whispers poorly on a tag. Pick a voice with the manner you need.

  • Up to 4,000 characters

    v3 takes twice as much text per request as v4, which is handy for long narration in one piece.

What to avoid in tags

  • Actions you cannot hear

    [standing], [grinning] and [pacing] do nothing: a tag must describe a sound, not a pose or a facial expression.

  • A tag that sounds like a sound effect

    Eleven v4 does both speech and sounds, so [crying] sometimes comes out as the sound of crying. Describe the voice: [trembling, tearful voice], [low, gravelly voice].

  • Delivery against the voice's character

    A meditation voice shouts unconvincingly, and a strict business voice laughs poorly on [giggles]. Pick a voice for the task.

  • Repeating what the text already says

    If the text says "he laughed out loud", do not replace it with a tag: the model reads the phrase, and you can add [chuckles] next to it.

FAQ

Can I write tags in another language?

The model gets the tag as is, but ElevenLabs gives all its examples in English. English tags work more reliably, while the voiceover text can be in any language.

How many tags can I use?

There is no limit, but tags count toward the character limit: 2,000 in Eleven v4 and 4,000 in Eleven v3. One tag per change of emotion is usually enough.

Why was a tag read aloud or played as a sound?

Most often the delivery does not fit the voice, or the tag looks like a sound effect. Describe the voice in words, for example [trembling, tearful voice], or pick a voice with a wider range.

How do I add a pause?

With [pause] or [long pause], an ellipsis or a new line. Eleven v4 and v3 do not support SSML tags such as <break>.

Why do two takes sound different?

Every generation is a new performance, like an actor's take. The higher the Expressiveness, the bigger the difference. Lower it for a steadier result.

Read also