使用发音词典

本指南介绍如何通过编程方式管理发音词典。

操作指南 · 假设你已完成 ElevenAPI 快速入门。

概览

发音词典可让你自定义 AI 智能体对特定单词或短语的发音。这在以下场景尤其有用:

  • 修正名称、地点或技术术语的发音
  • 确保不同对话中的发音保持一致
  • 自定义地区发音差异

ElevenLabs 同时支持 IPA 和 CMU 字母表。

发音词典音素标签仅适用于 eleven_v4、eleven_flash_v2 和 eleven_v3 模型。

其他模型会跳过词典音素标签并使用默认发音。对于其他模型,请改用 别名标签来替换拼写或短语,以获得所需发音。

如果想在英语以外的语言中使用 IPA 和 CMU 发音,必须 切换到 eleven_v4 模型。

快速入门

本指南假设你已设置 API 密钥和 SDK。如果尚未完成,请先完成 快速入门。

1

创建发音词典文件

本示例将为单词 tomato 创建一个发音词典文件。

此规则将使用“IPA”字母表,并为 tomato 和 Tomato 设置不同的发音。PLS 文件区分大小写,因此需同时包含带大写 “T” 和不带大写 “T” 的写法。

可以使用 Claude 或 ChatGPT 等 AI 工具,帮助为特定单词生成 IPA 或 CMU 标注。

dictionary.pls
<?xml version="1.0" encoding="UTF-8"?>
<lexicon version="1.0"
xmlns="http://www.w3.org/2005/01/pronunciation-lexicon"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://www.w3.org/2005/01/pronunciation-lexicon
http://www.w3.org/TR/2007/CR-pronunciation-lexicon-20071212/pls.xsd"
alphabet="ipa" xml:lang="en-US">
<lexeme>
<grapheme>tomato</grapheme>
<phoneme>/tə'meɪtoʊ/</phoneme>
</lexeme>
<lexeme>
<grapheme>Tomato</grapheme>
<phoneme>/tə'meɪtoʊ/</phoneme>
</lexeme>
</lexicon>
2

通过 SDK 从文件创建发音词典

根据所用语言,新建名为 example.py 或 example.mts 的文件,并添加以下代码:

from elevenlabs import ElevenLabs, PronunciationDictionaryVersionLocator
from elevenlabs.play import play
elevenlabs = ElevenLabs()
with open("dictionary.pls", "rb") as f:
# this dictionary changes how tomato is pronounced
pronunciation_dictionary = elevenlabs.pronunciation_dictionaries.create_from_file(
file=f.read(), name="example"
)
audio_1 = elevenlabs.text_to_speech.convert(
text="Without the dictionary: tomato",
voice_id="aMSt68OGf4xUZAnLpTU8",
model_id="eleven_flash_v2",
)
audio_2 = elevenlabs.text_to_speech.convert(
text="With the dictionary: tomato",
voice_id="aMSt68OGf4xUZAnLpTU8",
model_id="eleven_flash_v2",
pronunciation_dictionary_locators=[
PronunciationDictionaryVersionLocator(
pronunciation_dictionary_id=pronunciation_dictionary.id,
version_id=pronunciation_dictionary.version_id,
)
],
)
# play the audio
play(audio_1)
play(audio_2)
3

执行代码

python example.py

应能通过扬声器听到两版音频:一版使用发音词典,另一版未使用。

后续步骤