专业语音克隆快速入门

本指南介绍如何使用 PVC API 创建专业语音克隆。

本指南将介绍如何使用 PVC API 创建专业语音克隆(PVC)。要通过控制台创建 PVC,请参阅专业语音克隆产品指南。

创建 PVC 需要使用 Creator 版或更高版本。

如需深入了解 IVC 和 PVC 的工作原理及其适用场景,请参阅语音克隆:工作原理。

如果不确定法律允许的范围,请查阅服务 条款和我们的 AI 安全 信息,了解更多详情。

通过 API 创建 PVC 的步骤比创建即时语音克隆多得多。这是因为 PVC 更复杂,需要更多数据和微调才能创建高质量克隆声音。

使用专业语音克隆 API

本指南假定你已设置 API 密钥和 SDK。如果尚未设置,请先完成 快速入门。

1

创建 PVC 音色

根据所选语言,新建名为 example.py 或 example.mts 的文件,然后添加以下代码以创建 PVC 音色:

# example.py
import os
import time
import base64
from contextlib import ExitStack
from io import BytesIO
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs
load_dotenv()
elevenlabs = ElevenLabs(
api_key=os.getenv("ELEVENLABS_API_KEY"),
)
voice = elevenlabs.voices.pvc.create(
name="My Professional Voice Clone",
language="en",
description="A professional voice clone of my voice"
)
print(voice)
2

上传音频文件

接下来,上传用于训练 PVC 的音频样本文件。有关如何从音频文件获得最佳效果的更多信息,请查看 PVC 产品指南中的提示和建议部分。

# Define the list of file paths explicitly
# Replace with the paths to your audio and/or video files.
# The more files you add, the better the clone will be.
sample_file_paths = [
"/path/to/your/first_sample.mp3",
"/path/to/your/second_sample.wav",
"relative/path/to/another_sample.mp4"
]
samples = None
files_to_upload = []
# Use ExitStack to manage multiple open files
with ExitStack() as stack:
for filepath in sample_file_paths:
# Open each file and add it to the stack
audio_file = stack.enter_context(open(filepath, "rb"))
filename = os.path.basename(filepath)
# Create a File object for the SDK
files_to_upload.append(
BytesIO(audio_file.read())
)
samples = elevenlabs.voices.pvc.samples.create(
voice_id=voice.voice_id,
files=files_to_upload # Pass the list of File objects
)
3

开始说话人分离

此步骤将尝试把音频文件分离为不同说话人。如果上传的音频包含多个说话人,则必须执行此步骤。

sample_ids_to_check = []
for sample in samples:
if sample.sample_id:
print(f"Starting separation for sample: {sample.sample_id}")
elevenlabs.voices.pvc.samples.speakers.separate(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
sample_ids_to_check.append(sample.sample_id)
while sample_ids_to_check:
# Create a copy of the list to iterate over, so we can remove items from the original
ids_in_batch = list(sample_ids_to_check)
for sample_id in ids_in_batch:
status_response = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample_id
)
status = status_response.status
print(f"Sample {sample_id} status: {status}")
if status == "completed" or status == "failed":
sample_ids_to_check.remove(sample_id)
if sample_ids_to_check:
# Wait before the next poll cycle
time.sleep(5) # Wait for 5 seconds
print("All samples have been processed or removed from polling.")
4

获取说话人音频

由于上一步需要一些时间才能完成,请在上一步完成后,通过单独的进程运行以下步骤。

说话人分离完成后,每个样本都会有一个说话人列表。对于包含多个说话人的样本,你需要选择用于 PVC 的说话人。若要识别说话人,可以获取并试听每个说话人的音频。

# Get the list of samples from the voice created in Step 3
voice = elevenlabs.voices.get(voice_id=voice_id)
samples = voice.samples
# Loop over each sample and save the audio for each speaker to a file
speaker_audio_output_dir = "path/to/speakers/"
if not os.path.exists(speaker_audio_output_dir):
os.makedirs(speaker_audio_output_dir)
for sample in samples:
speaker_info = elevenlabs.voices.pvc.samples.speakers.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id
)
# Proceed only if separation is actually complete
if getattr(speaker_info, 'status', 'unknown') != "completed":
continue
if hasattr(speaker_info, 'speakers') and speaker_info.speakers:
speaker_list = speaker_info.speakers
if isinstance(speaker_info.speakers, dict):
speaker_list = speaker_info.speakers.values()
for speaker in speaker_list:
audio_response = elevenlabs.voices.pvc.samples.speakers.audio.get(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
speaker_id=speaker.speaker_id
)
audio_base64 = audio_response.audio_base_64
audio_data = base64.b64decode(audio_base64)
output_filename = os.path.join(speaker_audio_output_dir, f"sample_{sample.sample_id}_speaker_{speaker.speaker_id}.mp3")
with open(output_filename, "wb") as f:
f.write(audio_data)
5

使用说话人 ID 更新样本

说话人分离完成后,可以更新样本,选择要用于 PVC 的说话人。

elevenlabs.voices.pvc.samples.update(
voice_id=voice.voice_id,
sample_id=sample.sample_id,
selected_speaker_ids=[speaker.speaker_id]
)
6

验证 PVC

开始训练前,需要进行验证,以确保你有权使用该声音。首先请求验证 CAPTCHA。

captcha_response = elevenlabs.voices.pvc.verification.captcha.get(voice.voice_id)
# Save captcha image to file
captcha_buffer = base64.b64decode(captcha_response)
with open('captcha.png', 'wb') as f:
f.write(captcha_buffer)

图片包含多行文本,声音所有者需要朗读并录音。完成后,提交录音以验证声音所有者的身份。

elevenlabs.voices.pvc.verification.captcha.verify(
voice_id=voice.voice_id,
recording=open('path/to/recording.mp3', 'rb')
)
7

(可选)请求人工验证

如果无法完成 CAPTCHA 验证,可以请求人工验证。请注意,处理时间会更长。

仅应在之前的验证步骤失败或无法完成时使用此方式,例如声音所有者有视力障碍时。

如需了解人工验证所需文件的列表,请联系支持团队,因为每种情况可能不同。

elevenlabs.voices.pvc.verification.request(
voice_id=voice.voice_id,
files=[open('path/to/verification/files.txt', 'rb')],
)
8

训练 PVC

接下来,开始训练流程。所需时间取决于提供的样本时长和数量。

elevenlabs.voices.pvc.train(
voice_id=voice.voice_id,
# Specify the model the PVC should be trained on
model_id="eleven_multilingual_v2"
)
# Poll the fine tuning status until it is complete or fails
# This example specifically checks for the eleven_multilingual_v2 model
while True:
voice_details = elevenlabs.voices.get(voice_id=voice.voice_id)
fine_tuning_state = None
if voice_details.fine_tuning and voice_details.fine_tuning.state:
fine_tuning_state = voice_details.fine_tuning.state.get("eleven_multilingual_v2")
if fine_tuning_state:
progress = None
if voice_details.fine_tuning.progress and voice_details.fine_tuning.progress.get("eleven_multilingual_v2"):
progress = voice_details.fine_tuning.progress.get("eleven_multilingual_v2")
print(f"Fine tuning progress: {progress}")
if fine_tuning_state == "fine_tuned" or fine_tuning_state == "failed":
print("Fine tuning completed or failed")
break
# Wait for 5 seconds before polling again
time.sleep(5)
9

使用新创建的音色

PVC 验证完成后,可像使用其他任何音色一样使用它。有关如何使用音色的更多信息,请参阅文本转语音快速入门。

后续步骤