Pronunciation dictionary guide: Fix any word in TTS on Eleven v4
- Written by
- Jack Limebear
- Published
ListenListen to this article
Eleven v4 sets a new bar for Text to Speech model pronunciation performance, handling abbreviations, technical terms, names, and mixed-language text more accurately than any model before it. Still, every brand has its own vocabulary, from unique product names to industry-specific jargon, meaning you’ll want full control over how those words sound.
Eleven v4 offers a range of ways to fix pronunciation, giving you granular control over how the model speaks. You can build a thorough pronunciation dictionary that applies across every script you write in the TTS app. Or, write IPA directly into your text for a one-off fix.
Here’s a quick sample of Eleven v4 handling a script packed with tricky terms:
This guide covers how to create and use a TTS pronunciation dictionary and all of the surrounding strategies you can use to make your audio sound as close to what you imagined as possible.
Summary
- A pronunciation dictionary is a set of rules that you write to tell a Text to Speech model how to say specific words or phrases.
- TTS pronunciation rules are either alias-based, which swap in different text, or phoneme-based, which use phonetic spelling.
- Eleven v4 supports phoneme rules and inline IPA (International Phonetic Alphabet).
- SSML phoneme and break tags don’t work on Eleven v4, but you can use a pronunciation dictionary or Audio Tags for a natural-language alternative.
- You can apply up to three pronunciation dictionaries per request through ElevenAPI.
What is a pronunciation dictionary?
A pronunciation dictionary is a tool that allows you to configure exactly how a text to speech model says certain words. By adding certain words or phrases to a pronunciation dictionary, you define how they’ll sound whenever they come up in a TTS script.
Pronunciation dictionaries are typically used for:
- Brand names: Brand names are a common entry into pronunciation dictionaries, especially for those with unique or unusual spellings.
- Acronyms: Certain acronyms have a specific way that you should say them. For example, AEO should sound like /ˌeɪ.iːˈoʊ/ ("A-E-O"), not /eɪˈjoʊ/ ("ay-yo"). A hard dictionary rule makes sure the model gets it right every time you write it.
- Place names and other proper nouns: Certain place names may sound quite different from how you would imagine they’re pronounced when looking at them. Worcestershire is a famous example here that a pronunciation guide could help avoid any issues with.
- Industry terms: Industry jargon or specialized terms from the medical or legal fields, for example, can benefit from a pronunciation dictionary if the model hasn’t explicitly trained on data from those sectors.
Outside of text to speech, pronunciation dictionaries are typically crowd-sourced repositories of audio clips from native speakers, which learners use to hear how certain words are spoken.

How to control TTS pronunciation on Eleven v4
Eleven v4 is ranked the highest quality Text to Speech model by Artificial Analysis' Provider Voice Arena.1 Part of what makes Eleven v4 a market leader is the model’s ability to handle complex text while still making the output sound natural.
To give users the best possible experience with Eleven v4, the model offers a range of ways to customize pronunciation, making sure every output is up to your accuracy standards.
Here is how you can control pronunciation with Eleven v4:
- Inline IPA: You can write IPA into your scripts between forward slashes to direct pronunciation without using the pronunciation dictionary. We’ll touch on the specifics of how to do this shortly.
- Phoneme rules in dictionaries: Write phoneme-based rules into your pronunciation dictionaries to control the phonetic sound of recurring words.
- Fluent pronunciation across languages: Eleven v4 can generate natural-sounding speech in 90+ languages. Keep native pronunciation in each language while using the same voice for a consistent sonic branding experience.
Here’s how different ElevenLabs TTS models stack up for pronunciation control:
Model | Alias Rules (Dictionary) | Phoneme Rules (Dictionary) | Inline IPA (/.../) | SSML Phoneme Tags (<phoneme>) |
Eleven v4 | Yes | Yes | Yes | No |
Eleven v4 Turbo | Yes | Yes | Yes | No |
Eleven v3 | Yes | Yes | Yes | No |
Eleven Flash v2 | Yes | Yes | No | Yes (English only) |
Eleven Turbo v2 | Yes | No | No | Yes (English only) |
Eleven Multilingual v2 | Yes | No | No | No |

What’s the difference between alias and phoneme rules in a pronunciation dictionary?
Alias pronunciation rules replace one piece of text with another before the model speaks the word. This happens behind the scenes whenever a model generates audio, replacing the sound of a word/phrase with what you’ve included in your pronunciation dictionary.
Here are some examples of alias pronunciation rules:
- “SaaS” becomes “sass”
- “UN” becomes “United Nations”
- “OAuth” becomes “oh auth”
On the other hand, phoneme rules give the model an exact phonetic spelling to use. These are more challenging to create, as you’ll have to write or generate the IPA transcription yourself and enter it into your dictionary. For example, "Kubernetes" becomes /ˌkuː.bərˈnɛt.iːz/.
How to fix brand name pronunciation in a TTS model
Brand names are one of the most common entries into pronunciation dictionaries, as they’re not always real words. If your company has invented a word or a unique spelling for your brand, TTS models may struggle to pronounce it correctly, as there is very little accessible training data that includes that specific pronunciation.
Here’s how to fix brand name pronunciation in Eleven v4:
- Test the default sound: Open the Text to Speech app and write your brand name. Generate a clip and listen to how it sounds. If there is a pronunciation error, you’ll need to proceed to the next step.
- Try an alias rule: The easiest way to fix a brand pronunciation error is to create an alias rule. They’re quick to write and intuitive.
- Switch to a phoneme rule: If your alias rule doesn’t quite work, resort to IPA. Transcribe your brand’s name using IPA and then create a rule within the Text to Speech pronunciation dictionary to fix its pronunciation.
- Cover every capitalization: Be sure to include the IPA transcription for both lowercase and uppercase versions of your brand’s name, if you use both.
- Use inline IPA for one-off changes: If you only want a quick fix for a one-off generation, you can simply add the IPA right into your script instead of configuring a dictionary entry.
After you’ve established a brand name rule, it’ll exist in every shared project going forward, whether in your AI agent or accessed through an API.
How to create a pronunciation dictionary with ElevenLabs
You can create a pronunciation dictionary by adding rules in the Text to Speech app, by uploading a .pls file, or through ElevenAPI. In this guide, we’ll cover the first option, which only takes a few steps.

Here are the steps you need to follow to create a pronunciation dictionary:
- Open the pronunciation dictionaries panel: From the Text to Speech page, click the search bar at the top, type "Pronunciation dictionary," and select Pronunciation dictionaries under Actions.
- Create or choose a dictionary: Click New to start a fresh dictionary, or select an existing one from the list on the left.
- Add a rule: Click Add rule to create a new row. Enter the word you want to fix. Type the word or phrase exactly as it appears in your scripts into the Input field. Rules are case-sensitive, so match the capitalization you use.
- Choose the rule type: Select Alias to swap in different text, or Phoneme to enter an IPA or CMU spelling.
- Enter the replacement: Type the alias or phonetic spelling into the Output field. Use the voice selector at the top of the panel to hear the rule with the voice you're working with.
- Save your changes: Click Save changes at the bottom of the panel.
To attach a pronunciation guide to an agent or apply one through the API, follow our pronunciation dictionary docs.
How to use IPA in text to speech on Eleven v4
For small changes or uncommon words, you don’t necessarily have to build out a new entry in your TTS pronunciation dictionary to fix a mistake. Instead, you can use inline IPA to detail exactly how the model should pronounce a word.
In Eleven v4, you can activate inline IPA by using forward slashes. Like Audio Tags, these go directly into your script to make things as easy as possible. Here are some examples of IPA in Eleven v4:
- Technical terms: The term "/ˌbaɪoʊˈkemɪstri/" refers to the study of chemical processes.
- Medical vocabulary: "/ɡluːˈkoʊs/" and "/ˈɪnsjəlɪn/" are commonly used to manage conditions like "/ˌdaɪəˈbiːtiːz/".
- Brand and product names: Our servers run on "/ˌɛndʒɪnˈɛks/" to handle traffic spikes.
As a side note, and for those without a degree in linguistics, using an IPA converter will help you get the right IPA to paste into your script from a natural-language input.
Why SSML phoneme tags don’t work on Eleven v4
Eleven v4 doesn’t support SSML, instead opting for a range of more flexible features, like Audio Tags and pronunciation dictionaries, that give you more control over the final audio production.
Every feature that SSML offers has a direct replacement in v4:
SSML tag | What it did | Eleven v4 replacement |
<phoneme> | Set a phonetic pronunciation | Inline IPA (/…/) or a dictionary phoneme rule |
<sub alias> | Swapped text before speaking | A dictionary alias rule |
<break> | Added a pause | Audio Tags like [pause], ellipses, or line breaks |
<prosody rate> | Changed speaking speed | Pacing Audio Tags like [slowly] or [rushed] |
<emphasis> | Stressed a word | Capital letters or a delivery Audio Tag |
Get started with natural-sounding text to speech pronunciation on Eleven v4
With a range of tools to help refine the accuracy of pronunciations on Eleven v4, you can write detailed scripts and have the model perform them as you imagined.
Build pronunciation dictionaries in the ElevenLabs Text to Speech app and apply them at scale with ElevenAPI.
Sign up to get started or learn more about Eleven v4 today.
FAQ
- Artificial Analysis, Provider Voice Arena Preference Elo, September 30, 2026


