多声道文本转语音

本指南介绍如何通过文本转语音 API 使用多声道转录模式。

操作指南 · 假设你已完成文本转语音 快速入门。

概述

多声道文本转语音功能可转录每个声道包含不同说话人的音频文件。这尤其适用于将说话人隔离在独立音频声道中的录音,无需说话人分离即可获得更清晰的转录结果。

每个声道都会独立处理,并根据声道编号自动分配说话人 ID(声道 0 → speaker_0、声道 1 → speaker_1 等)。系统会从输入音频文件中提取各个声道,并并行转录。默认情况下,API 为每个声道返回一份转录文本;设置 multichannel_output_style=combined 后,则会返回一份合并所有声道的转录文本,以开始时间排序的单个列表形式呈现,并为每个词标记其 channel_index。

常见使用场景

  • 立体声访谈录音 - 采访者在左声道,受访者在右声道
  • 多轨播客录音 - 每位参与者录制在独立音轨中
  • 呼叫中心录音 - 客服与客户位于不同声道
  • 会议录音 - 各参与者隔离在独立声道中
  • 庭审记录 - 多方录制在不同声道中

要求

  • 拥有API 密钥的 ElevenLabs 账户
  • 多声道音频文件(WAV、MP3 或其他支持的格式)
  • 每个音频文件最多 5 个声道
  • 每个声道应只包含 1 位说话人

工作原理

1

准备多声道音频

确保音频文件中的说话人被隔离在不同声道中。多声道功能最多支持 5 个声道,每个声道对应一位特定说话人:

  • 声道 0 → speaker_0
  • 声道 1 → speaker_1
  • 声道 2 → speaker_2
  • 声道 3 → speaker_3
  • 声道 4 → speaker_4
2

配置 API 参数

发起文本转语音请求时,必须设置:

  • use_multi_channel:true
  • diarize:false(多声道模式通过声道区分说话人)

还可通过以下参数控制响应格式:

  • multichannel_output_style:separate(默认)为每个声道返回一份转录文本。combined 会将所有声道合并为一份转录文本,其中词语按开始时间排序,每个词都带有 channel_index,与标准单声道响应格式一致。combined 需要时间戳(timestamps_granularity 不得为 none),且不支持 webhook 传送或实体检测/脱敏。

多声道模式不能使用 num_speakers 参数,因为说话人数量由声道数自动确定。多声道模式假定每个声道恰好有 1 位说话人。如果一个声道有多位说话人,系统会为该声道中的所有说话人分配相同的 speaker id。

3

处理响应

默认情况下(multichannel_output_style=separate),多声道音频返回的响应格式与单声道不同:

如果设置 use_multi_channel: true,但提供单声道(单声道)音频文件, 将收到标准单声道响应,而非多声道格式。仅当音频文件实际包含多个声道时, 才会返回多声道响应格式。

{
"language_code": "en",
"language_probability": 0.98,
"text": "Hello world",
"words": [...]
}

使用 multichannel_output_style=combined 时,响应采用与单声道转录相同的扁平格式(顶层为 text 和 words,没有 transcripts 数组),所有声道合并为按开始时间排序的单个列表。每个词都包含标识其声道的 channel_index(以及 speaker_id)。

实现

基础多声道转录

以下是转录包含 2 位说话人的立体声音频文件的完整示例:

from elevenlabs import ElevenLabs
elevenlabs = ElevenLabs(api_key="YOUR_API_KEY")
def transcribe_multichannel(audio_file_path):
with open(audio_file_path, 'rb') as audio_file:
result = elevenlabs.speech_to_text.convert(
file=audio_file,
model_id='scribe_v2',
use_multi_channel=True,
diarize=False,
timestamps_granularity='word'
)
return result
# Process the response
result = transcribe_multichannel('stereo_interview.wav')
if hasattr(result, 'transcripts'): # Multichannel response
for transcript in result.transcripts:
channel = transcript.channel_index
text = transcript.text
print(f"Channel {channel} (speaker_{channel}): {text}")
else: # Single channel response (fallback)
print(f"Text: {result.text}")

创建对话转录文本

获取按时间排序、对话式转录文本的最简单方式是请求 multichannel_output_style=combined。API 会返回一个按开始时间排序的 words 列表,每个词都带有 channel_index 和 speaker_id:

合并输出(推荐)
with open("stereo_interview.wav", "rb") as audio_file:
result = elevenlabs.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
use_multi_channel=True,
multichannel_output_style="combined",
diarize=False,
timestamps_granularity="word",
)
for word in result.words:
if word.type == "word":
print(f"speaker_{word.channel_index}: {word.text}")

如果使用默认的 separate 输出,也可以在客户端合并各声道的转录文本:

def create_conversation_transcript(multichannel_result):
"""Create a conversation-style transcript with speaker labels"""
all_words = []
if hasattr(multichannel_result, 'transcripts'):
# Collect all words from all channels
for transcript in multichannel_result.transcripts:
for word in transcript.words or []:
if word.type == 'word':
all_words.append({
'text': word.text,
'start': word.start,
'speaker_id': word.speaker_id,
'channel': transcript.channel_index
})
# Sort by timestamp
all_words.sort(key=lambda w: w['start'])
# Group consecutive words by speaker
conversation = []
current_speaker = None
current_text = []
for word in all_words:
if word['speaker_id'] != current_speaker:
if current_text:
conversation.append({
'speaker': current_speaker,
'text': ' '.join(current_text)
})
current_speaker = word['speaker_id']
current_text = [word['text']]
else:
current_text.append(word['text'])
# Add the last segment
if current_text:
conversation.append({
'speaker': current_speaker,
'text': ' '.join(current_text)
})
return conversation
# Format the output
conversation = create_conversation_transcript(result)
for turn in conversation:
print(f"{turn['speaker']}: {turn['text']}")

在多声道模式中使用 webhook

多声道转录支持webhook 传送,可进行异步处理:

webhook 返回 separate(按声道)格式。webhook 传送目前不支持 multichannel_output_style=combined, 请使用同步请求,或在客户端合并各声道的 webhook 负载。

from elevenlabs import ElevenLabs
elevenlabs = ElevenLabs(api_key="YOUR_API_KEY")
async def transcribe_multichannel_with_webhook(audio_file_path):
with open(audio_file_path, 'rb') as audio_file:
result = await elevenlabs.speech_to_text.convert_async(
file=audio_file,
model_id='scribe_v2',
use_multi_channel=True,
diarize=False,
webhook=True # Enable webhook delivery
)
print(f"Transcription started with task ID: {result.task_id}")
return result.task_id

错误处理

常见验证错误

错误:多声道模式不支持说话人分离,并会根据说话人所在声道分配说话人。

解决方案:使用多声道模式时,始终设置 diarize=false。

错误:启用 use_multi_channel 时,无法指定 num_speakers。说话人数量会根据声道数 自动确定。解决方案:从请求中移除 num_speakers 参数。

错误:多声道模式最多支持 5 个声道,但该音频文件包含 X 个声道。

解决方案:仅处理前 5 个声道,或预处理音频以减少 声道数。

错误:multichannel_output_style=‘combined’ 需要时间戳;请将 timestamps_granularity 设置为 ‘word’ 或 ‘character’。

解决方案:合并输出按时间排序词语,因此请将 timestamps_granularity 设置为 word (默认值)或 character。

错误:webhook 传送暂不支持 multichannel_output_style=‘combined’。

解决方案:对 combined 使用同步请求;或者使用 webhook 时保留默认 separate 输出, 并在客户端合并。

最佳实践

音频准备

为获得最佳结果:- 使用 16kHz 采样率以提升性能 - 处理前移除静音或未使用的 声道 - 确保每个声道只包含 1 位说话人 - 尽可能使用无损格式 (WAV)以获得最佳质量

性能优化

并发成本会随声道数量线性增长。一个 60 秒的 3 声道文件,其并发成本是单声道文件的 3 倍。

可使用以下公式估算多声道音频的处理时间:

Processing Time=(D⋅0.3)+2+(N⋅0.5)Processing\ Time = (D \cdot 0.3) + 2 + (N \cdot 0.5)

其中:

  • DD = 文件时长(秒)
  • NN = 声道数
  • 0.30.3 = 处理速度系数(约为实时速度的 30%)
  • 22 = 固定开销(秒)
  • 0.50.5 = 每声道开销(秒)

示例:对于一个 60 秒的立体声文件(2 个声道):

Processing Time=(60⋅0.3)+2+(2⋅0.5)=18+2+1=21 secondsProcessing\ Time = (60 \cdot 0.3) + 2 + (2 \cdot 0.5) = 18 + 2 + 1 = 21\ seconds

内存注意事项

对于大型多声道文件,请考虑流式处理或分块:

def process_large_multichannel_file(file_path, chunk_duration=300):
"""Process large files in chunks (5-minute segments)"""
from pydub import AudioSegment
from elevenlabs import ElevenLabs
import os
elevenlabs = ElevenLabs(api_key="YOUR_API_KEY")
audio = AudioSegment.from_file(file_path)
duration_ms = len(audio)
chunk_size_ms = chunk_duration * 1000
all_transcripts = []
for start_ms in range(0, duration_ms, chunk_size_ms):
end_ms = min(start_ms + chunk_size_ms, duration_ms)
# Extract chunk
chunk = audio[start_ms:end_ms]
chunk_file = f"temp_chunk_{start_ms}.wav"
chunk.export(chunk_file, format="wav")
# Transcribe chunk using SDK
with open(chunk_file, 'rb') as audio_file:
result = elevenlabs.speech_to_text.convert(
file=audio_file,
model_id='scribe_v2',
use_multi_channel=True,
diarize=False,
timestamps_granularity='word'
)
# Adjust timestamps
if hasattr(result, 'transcripts'):
for transcript in result.transcripts:
for word in transcript.words or []:
word.start += start_ms / 1000
word.end += start_ms / 1000
all_transcripts.extend(result.transcripts)
# Clean up
os.remove(chunk_file)
return all_transcripts

常见问题

API 会返回错误。你需要选择要发送到 API 的 5 个声道,或在发送前将部分声道混合为更少的声道。

可以,但没有必要。如果发送单声道音频并设置 use_multi_channel=true,将收到标准单声道响应,而非多声道格式。

可以。设置 multichannel_output_style=combined,即可获得一份合并所有声道并按开始时间排序的转录文本,每个词都带有 channel_index。这与标准单声道响应格式一致。该模式需要时间戳,且不适用于 webhook 传送。

说话人 ID 根据声道编号确定:声道 0 为 speaker_0,声道 1 为 speaker_1,依此类推。

可以,每个声道都独立处理,可以检测不同语言。语言检测会按声道进行。使用 multichannel_output_style=combined 时,顶层 language_code 反映置信度最高的声道,而每个词仍带有其 channel_index。

后续步骤