탐색으로 건너뛰기

대화 시뮬레이션

시뮬레이션된 대화로 ElevenLabs 에이전트를 테스트하고 평가하는 방법

이 가이드와 여기서 사용하는 엔드포인트는 더 이상 지원되지 않습니다. 대신 시뮬레이션 에이전트 테스트 유형을 사용하세요.

개요

ElevenLabs Agents API를 사용하면 AI 에이전트와의 텍스트 기반 대화를 시뮬레이션하고 평가할 수 있습니다. 이 가이드에서는 simulate conversation 엔드포인트(배치 및 스트리밍)를 활용해 엔드투엔드 시뮬레이션 테스트 워크플로를 구현하는 방법을 알아봅니다. 이를 통해 에이전트의 성능을 세밀하게 테스트하고 개선하여 상호작용 목표를 충족할 수 있습니다.

사전 준비 사항

시뮬레이션 테스트 워크플로 구현

1

초기 평가 매개변수 파악

에이전트의 대화 기록을 검토하여 성능이 기대에 미치지 못했던 사례를 찾으세요. 해당 대화를 활용해 에이전트와 상호작용할 시뮬레이션 사용자용 다양한 프롬프트를 만드세요. 또한 특정 시뮬레이션 사용자에게 원하는 결과를 테스트할 수 있도록, 에이전트 구성에 아직 지정되지 않은 추가 평가 기준을 정의하세요.

2

SDK로 대화 시뮬레이션

ElevenLabs SDK를 사용해 시뮬레이션 엔드포인트에 요청을 생성하세요.

from dotenv import load_dotenv
from elevenlabs import (
ElevenLabs,
ConversationSimulationSpecification,
AgentConfig,
PromptAgent,
PromptEvaluationCriteria
)
load_dotenv()
api_key = os.getenv("ELEVENLABS_API_KEY")
elevenlabs = ElevenLabs(api_key=api_key)
response = elevenlabs.conversational_ai.agents.simulate_conversation(
agent_id="YOUR_AGENT_ID",
simulation_specification=ConversationSimulationSpecification(
simulated_user_config=AgentConfig(
prompt=PromptAgent(
prompt="Your goal is to be a really difficult user.",
llm="gpt-4o",
temperature=0.5
)
)
),
extra_evaluation_criteria=[
PromptEvaluationCriteria(
id="politeness_check",
name="Politeness Check",
conversation_goal_prompt="The agent was polite.",
use_knowledge_base=False
)
]
)
print(response)

이는 기본 예시입니다. 전체 입력 매개변수 목록은 대화 시뮬레이션 및 스트리밍 대화 시뮬레이션 엔드포인트의 API 레퍼런스를 참조하세요.

3

응답 분석

SDK는 전체 대화 기록과 상세 분석이 포함된 종합적인 JSON 객체를 제공합니다.

시뮬레이션된 대화: 시뮬레이션 사용자와 에이전트 간 각 상호작용 턴의 메시지와 도구 사용 내역을 기록합니다.

대화 기록 예시
[
...
{
"role": "user",
"message": "Maybe a little. I'll think about it, but I'm still not convinced it's the right move.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "I understand. If you want to explore more at your own pace, I can direct you to our documentation, which has guides and API references. Would you like me to send you a link?",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "user",
"message": "I guess it wouldn't hurt to take a look. Go ahead and send it over.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"params_as_json": "{\"path\": \"/docs/api-reference/introduction\"}",
"tool_has_been_called": false,
"tool_details": null
}
],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [],
"tool_results": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"result_value": "Tool Called.",
"is_error": false,
"tool_has_been_called": true,
"tool_latency_secs": 0
}
],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "Okay, I've sent you a link to the introduction to our API reference. It provides a good starting point for understanding our different tools and how they can be integrated. Let me know if you have any questions as you explore it.\n",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
}
...
]

분석: 평가 기준 결과, 데이터 수집 지표 및 대화 기록 요약에 대한 인사이트를 제공합니다.

분석 예시
{
"analysis": {
"evaluation_criteria_results": {
"politeness_check": {
"criteria_id": "politeness_check",
"result": "success",
"rationale": "The agent remained polite and helpful despite the user's challenging attitude."
},
"understood_root_cause": {
"criteria_id": "understood_root_cause",
"result": "success",
"rationale": "The agent acknowledged the user's hesitation and provided relevant information."
},
"positive_interaction": {
"criteria_id": "positive_interaction",
"result": "success",
"rationale": "The user eventually asked for the documentation link, indicating engagement."
}
},
"data_collection_results": {
"issue_type": {
"data_collection_id": "issue_type",
"value": "support_issue",
"rationale": "The user asked for help with integrating ElevenLabs tools."
},
"user_intent": {
"data_collection_id": "user_intent",
"value": "The user is interested in integrating ElevenLabs tools into a project."
}
},
"call_successful": "success",
"transcript_summary": "The user expressed skepticism, but the agent provided useful information and a link to the API documentation."
}
}
4

평가 기준 개선

시뮬레이션된 대화를 면밀히 검토하여 평가 기준의 효과를 확인하세요. 평가 기준이 에이전트의 성능을 평가하는 데 부족할 수 있는 공백이나 영역을 파악하세요. 원하는 결과와 부합하고 에이전트의 역량을 정확히 측정할 수 있도록 평가 기준을 개선하고 조정하세요.

5

에이전트 개선

평가 기준의 정확성에 확신이 들면 시뮬레이션된 대화에서 얻은 인사이트를 활용해 에이전트의 역량을 강화하세요. 목표와 사용자 기대에 부합하도록 에이전트의 응답을 더 잘 안내하기 위해 시스템 프롬프트를 개선하는 것을 고려하세요. 또한 에이전트의 어조 조정, 특정 쿼리 처리 능력 향상, 응답을 풍부하게 할 추가 데이터 소스 통합 등 최적화할 수 있는 다른 기능이나 구성을 살펴보세요. 이러한 인사이트를 체계적으로 적용하면 더 견고하고 효과적인 대화형 에이전트를 만들어 뛰어난 사용자 경험을 제공할 수 있습니다.

6

지속적인 반복

초기 테스트 및 개선 주기를 완료한 후에는 광범위한 가능한 시나리오를 다루기 위해 종합적인 테스트 모음을 구축하는 것이 좋습니다. 이 모음에서는 다양한 시뮬레이션 사용자 프롬프트와 시작 조건을 사용해 여러 시뮬레이션 대화를 살펴볼 수 있습니다. 접근 방식을 지속적으로 반복하고 개선하면 에이전트가 변화하는 사용자 요구에 효과적으로 대응할 수 있습니다.

전문가 팁

상세한 프롬프트와 기준

상세하고 충분한 정보를 담은 시뮬레이션 사용자 프롬프트와 평가 기준을 작성하면 시뮬레이션 테스트의 효과를 높일 수 있습니다. 더 많은 맥락과 구체적인 정보를 제공할수록 에이전트가 복잡한 상호작용을 더 잘 이해하고 대응할 수 있습니다.

모의 도구 구성

모의 도구 구성을 사용하여 에이전트의 의사 결정 과정을 테스트하세요. 이를 통해 에이전트가 도구 호출을 결정하는 방식과 다양한 도구 호출 결과에 반응하는 방식을 관찰할 수 있습니다. 자세한 내용은 API 레퍼런스의 tool_mock_config 입력 매개변수를 확인하세요.

부분 대화 기록

부분 대화 기록을 사용하여 에이전트가 특정 시점부터의 상호작용을 처리하는 방식을 평가하세요. 이는 사용자가 이미 특정 방식으로 질문을 설정했거나, 특정 도구 호출이 성공 또는 실패한 상황에서 대화를 관리하는 에이전트의 능력을 평가하는 데 특히 유용합니다. 자세한 내용은 API 레퍼런스의 partial_conversation_history 입력 매개변수를 확인하세요.