Python SDK

ElevenAgents SDK:カスタマイズ可能なインタラクティブ音声エージェントを数分で導入。

ElevenAgentsの概要もご覧ください。

インストール

プロジェクトにelevenlabs Pythonパッケージをインストールします。

pip install elevenlabs
# or
poetry add elevenlabs

デフォルトのオーディオ入力/出力実装を使用する場合は、pyaudioエクストラも必要です。

pip install "elevenlabs[pyaudio]"
# or
poetry add "elevenlabs[pyaudio]"

pyaudioパッケージのインストールには、追加のシステム依存関係が必要になる場合があります。

詳しくは、PyAudioパッケージのREADMEをご覧ください。

Debianベースのシステムでは、以下のコマンドで依存関係をインストールできます。

sudo apt-get update
sudo apt-get install libportaudio2 libportaudiocpp0 portaudio19-dev libasound-dev libsndfile1-dev -y

使用方法

この例では、ElevenLabs Agentsエージェントとの会話を実行するシンプルなスクリプトを作成します。

まず、必要な依存関係をインポートします。

import os
import signal
from elevenlabs.client import ElevenLabs
from elevenlabs.conversational_ai.conversation import Conversation
from elevenlabs.conversational_ai.default_audio_interface import DefaultAudioInterface

次に、環境変数からエージェントIDとAPIキーを読み込みます。

agent_id = os.getenv("AGENT_ID")
api_key = os.getenv("ELEVENLABS_API_KEY")

APIキーが必要なのは、認証が有効な非パブリックエージェントのみです。 パブリックエージェントでは設定する必要はなく、設定しなくてもコードは問題なく動作します。

次に、ElevenLabsクライアントインスタンスを作成します。

elevenlabs = ElevenLabs(api_key=api_key)

次に、Conversationインスタンスを初期化します。

conversation = Conversation(
# API client and agent ID.
elevenlabs,
agent_id,
# Assume auth is required when API_KEY is set.
requires_auth=bool(api_key),
# Use the default audio interface.
audio_interface=DefaultAudioInterface(),
# Simple callbacks that print the conversation to the console.
callback_agent_response=lambda response: print(f"Agent: {response}"),
callback_agent_response_correction=lambda original, corrected: print(f"Agent: {original} -> {corrected}"),
callback_user_transcript=lambda transcript: print(f"User: {transcript}"),
# Uncomment if you want to see latency measurements.
# callback_latency_measurement=lambda latency: print(f"Latency: {latency}ms"),
# Uncomment if you want to receive audio alignment data with character-level timing.
# callback_audio_alignment=lambda alignment: print(f"Alignment: {alignment.chars}"),
)

ここでは、会話にシステムのデフォルトオーディオ入力/出力デバイスを使用するDefaultAudioInterfaceを使用しています。 elevenlabs.conversational_ai.conversation.AudioInterfaceをサブクラス化して、独自のオーディオインターフェースを実装することもできます。

これで会話を開始できます。必要に応じて、会話をユーザーに紐付けるために独自のエンドユーザーIDを渡すことをおすすめします。

conversation.start_session(
user_id=user_id # optional field
)

ユーザーがCtrl+Cを押したときに正常に終了できるよう、end_session()を呼び出すシグナルハンドラーを追加します。

signal.signal(signal.SIGINT, lambda sig, frame: conversation.end_session())

最後に、会話が終了するのを待ってから会話IDを出力します(会話履歴の確認やデバッグに使用できます)。

conversation_id = conversation.wait_for_session_end()
print(f"Conversation ID: {conversation_id}")

あとはスクリプトを実行して、エージェントとの会話を開始するだけです。

# For public agents:
AGENT_ID=youragentid python demo.py
# For private agents:
AGENT_ID=youragentid ELEVENLABS_API_KEY=yourapikey python demo.py