Eleven v4 and v4 Turbo: the new ElevenLabs voices in Givon AI

ElevenLabs released Eleven v4 and v4 Turbo, and we added both models to Voiceover. Here is how v4 differs from v3, how to set delivery with tags, what a voiceover costs and when v3 comes in handy.
Contents
  1. What's new
  2. Leaderboard
  3. v4 and v3
  4. v4 or Turbo
  5. Tags
  6. Settings
  7. Where it works
  8. When v3
  9. Coming soon
  10. Tips
  11. FAQ

On September 28 ElevenLabs released Eleven v4 and Eleven v4 Turbo, a new generation of text-to-speech models. According to the company, v4 reads text in the context of the whole scene, follows emotion tags more closely and speaks 90+ languages, up from 70+ in Eleven v3. We added both models to Givon AI, and they work in the Voiceover mode.

What changed in Eleven v4

According to ElevenLabs, v4 runs on a new architecture. The company highlights four changes:

  • Scene context

    The model reads the whole text, so a line sounds like a reply to what came before it rather than a standalone phrase.

  • Closer tag following

    Square-bracket tags work more reliably than in v3, and you can stack them: the model plays them in order.

  • More languages

    The ElevenLabs list has 91 languages, up from 74 in v3. New ones include Uzbek, Tajik, Mongolian, Cantonese, Amharic, Burmese, Zulu and 12 more.

  • The same voice in every take

    When you regenerate a line, the voice stays the same and only the performance changes.

First place on the Artificial Analysis leaderboard

On the Artificial Analysis leaderboard, listeners blindly pick the better of two recordings of the same text. As of September 28, Eleven v4 is first and 150 points ahead of Eleven v3:

ModelElo score
Eleven v41319
Cartesia Sonic 3.61276
Gemini 3.8 Flash TTS1267
Qwen-Audio-3.0-TTS-Plus1258
Inworld Realtime TTS-21246
Eleven v31169
Eleven Multilingual v21094
Provider Voice leaderboard as of September 28, 2026. The v4 score comes from 1,674 comparisons on release day and may still change.

ElevenLabs also cites its own test: in a blind comparison with Cartesia Sonic 3.6, Inworld TTS-2 and two Gemini TTS models, about 75% of listeners chose v4.

Eleven v4 and Eleven v3 in Givon AI

ParameterEleven v4Eleven v3
Text per requestUp to 2,000 charactersUp to 4,000 characters
How to set deliveryTags in the text, Expressiveness and SimilarityTags in the text and Delivery: Natural, Calm, Expressive, Ad
Languages according to ElevenLabs90+70+
Price in tokens8, 14, 24 or 448, 14, 24 or 44
Artificial Analysis score13191169

v4 costs the same as v3, and the price depends on text length: 8 tokens up to 250 characters, 14 up to 500, 24 up to 1,000 and 44 for longer text. Square-bracket tags count toward the characters. Plan discounts do not apply to voiceover.

In one request v4 voices up to 2,000 characters and returns word timings along with the audio, so Givon builds captions right away. Split a long script into several voiceovers.

Listen to the same text in v3 and v4: the Marina voice, the first take with no cherry-picking. The recordings are in Russian; the English version of the text is the explainer example below.

v4 or v4 Turbo

ParameterEleven v4Eleven v4 Turbo
Built forVideo voiceover, characters and audiobooksVoice assistants and agents where latency matters
Price in tokens8, 14, 24 or 444, 7, 12 or 22
Text per requestUp to 2,000 charactersUp to 2,000 characters
SettingsTags, Expressiveness, SimilarityTags, Expressiveness, Similarity

ElevenLabs built Turbo for real-time conversation: the model responds in about 100 ms. In Givon Turbo has another advantage: it costs half as much as v4.

Turbo fits drafts. Use it to check the text, pace and tag placement. v4 is for the final take: once the script is approved, regenerate it in v4 with the same tags.

The ad example below on both models, in Russian, with the Marina voice and default settings:

Delivery through tags

v4 has no Delivery presets: tone is set with square-bracket tags right in the text. Type “[” in the voiceover field and Givon suggests 12 tags: [whispers], [excited], [laughs], [sighs], [sad], [angry], [curious], [sarcastic], [crying], [shouting], [pause] and [exhales].

You can also write your own tag: any text of up to 40 characters in square brackets goes to the model as is, for example [nervous laugh], [low, steady voice] or [under his breath].

  • Tag before the phrase

    A tag changes the part of the text that follows it. When the emotion changes, add a new tag.

  • Stacked tags

    [sighs] [sad] means a sigh first, then a sad line. v4 plays tags in order.

  • Pauses and emphasis

    An ellipsis adds a pause and weight to a phrase, and a word in CAPITALS is stressed.

  • Sounds in the text

    v4 also understands sound tags such as [applause] or [door slams], but they sound different on different voices.

Example for an ad

[excited] The new collection is live! [whispers] And the first hundred buyers get twenty percent off. [laughs] Grab yours before it's gone.

Example for an explainer

[curious] Why does coffee after lunch wake you up less than in the morning? [pause] It comes down to cortisol. [low, steady voice] In the morning your level is already high, and caffeine adds almost nothing... But two or three hours after you get up, a cup works at full strength.

A character in Eleven v4

[sighs] Here you go again... [sarcastic] Sure, that was a VERY good idea. [exhales] Fine. [curious] So, what is it this time?

Recorded in Russian (Marina voice, Expressiveness 75%, Similarity 75%); the text is its English translation.

ElevenLabs marks [sings] and [strong French accent] as experimental: they work on some voices and not on others. How to pick a voice, place tags and pauses and write text for v4 is in the voiceover script guide.

Expressiveness and Similarity

Instead of Delivery, v4 has two percentage settings. Besides the preset values, each has a Custom value option.

  • Expressiveness, 50% by default

    Higher makes speech more expressive but increases the difference between takes. Lower makes the voice steadier and more predictable.

  • Similarity, 75% by default

    How closely the voiceover matches the selected voice. A high value reproduces the timbre more precisely, along with any flaws of the recording the voice was made from.

The settings are available in Generation on desktop and mobile, in the Agent and in the API.

Where Eleven v4 works

  • Generation

    In the Voiceover mode, pick Eleven v4 or Eleven v4 Turbo in the model list, write the text and choose a Voice.

  • Editor

    The finished voiceover goes on the Voiceover track, and its word timings go straight into captions.

  • Studio

    Avatar speech is voiced with Eleven v4, with text up to 2,000 characters.

  • Agent

    The Agent proposes a voiceover in v4 or v4 Turbo and sets Expressiveness and Similarity itself.

  • API and MCP

    The elevenlabs-tts-v4 and elevenlabs-tts-v4-turbo models with the stability and similarity parameters.

  • Glossary

    Pronunciation rules from Library → Voices → Glossary also apply in v4.

New projects use v4 for voiceover by default, and voice previews play in v4 too. If you picked Eleven v3 before, the prompt bar remembers that choice, so switch the model yourself.

When Eleven v3 comes in handy

Eleven v3 stays in the model list and fits two cases:

  • Long text in one piece

    v3 takes up to 4,000 characters per request, v4 up to 2,000.

  • Ready-made presets

    Natural, Calm, Expressive and Ad are available only in v3.

Coming soon: dialogue with several voices

v4 can voice a scene with lines from different voices in one generation, with speakers reacting to what was just said. We are already bringing dialogue to Givon and will announce it separately.

How to get the delivery you want

  • A tag came out as a sound

    Describe the delivery in words: instead of [crying], write [trembling, tearful voice].

  • Match the voice to the delivery

    A whisper works best with a soft voice, a shout with an energetic one. Preview a few voices on the same text.

  • Pace

    Ellipses and [pause] slow speech down; short phrases and [excited] speed it up.

  • Another language

    In another language the voice speaks with that language's native accent.

FAQ

What is ElevenLabs v4?

Eleven v4 is an ElevenLabs text-to-speech model released on September 28, 2026. It came out together with Eleven v4 Turbo for voice assistants. Both are available in Givon AI in the Voiceover mode.

Which languages does Eleven v4 speak?

The ElevenLabs list has 91 languages, including English, Spanish, Russian and Uzbek. The model detects the language of the text itself.

How is Eleven v4 different from v3?

v4 follows tags better, takes the context of the whole text into account and knows more languages: 90+ versus 70+. In Givon v4 has a 2,000-character limit versus 4,000 in v3, and instead of Delivery presets it has tags plus the Expressiveness and Similarity settings. The price is the same.

How much does an Eleven v4 voiceover cost?

8 tokens for text up to 250 characters, 14 up to 500, 24 up to 1,000 and 44 up to 2,000. Eleven v4 Turbo costs half as much: 4, 7, 12 and 22 tokens.

Do I need an ElevenLabs account?

No. Eleven v4 runs inside Givon AI and is paid with Givon tokens, so you do not need a separate ElevenLabs subscription.

Can I clone a voice from 10 seconds of audio?

ElevenLabs says so in its announcement, and its v4 documentation recommends one to two minutes of audio. In Givon you create your own voice in Library → Voices → New voice from 1–2 minutes of clean speech.

Should I pick v4 or v4 Turbo?

Turbo is for drafts and trying variants: it costs half as much. v4 is for the final voiceover that goes into the video.

Models in this article

Read also