模拟对话

了解如何通过模拟对话测试和评估 ElevenLabs 智能体

本指南及其使用的端点已弃用。请改用 Simulation 智能体测试类型。

概述

ElevenLabs Agents API 支持你模拟和评估与 AI 智能体进行的文本对话。本指南将介绍如何使用模拟对话端点(批量 和 流式)实现端到端模拟测试 workflow,帮助你精细测试并改进智能体表现,确保其达到预期的交互目标。

前提条件

实现模拟测试 Workflow

1

确定初始评估参数

查看智能体的对话历史,找出表现不佳的案例。基于这些对话,为将与你的智能体互动的模拟用户创建不同提示词。此外,为特定模拟用户定义智能体配置中尚未指定的其他评估标准,以测试你期望的结果。

2

通过 SDK 模拟对话

使用 ElevenLabs SDK 向模拟端点创建请求。

from dotenv import load_dotenv
from elevenlabs import (
ElevenLabs,
ConversationSimulationSpecification,
AgentConfig,
PromptAgent,
PromptEvaluationCriteria
)
load_dotenv()
api_key = os.getenv("ELEVENLABS_API_KEY")
elevenlabs = ElevenLabs(api_key=api_key)
response = elevenlabs.conversational_ai.agents.simulate_conversation(
agent_id="YOUR_AGENT_ID",
simulation_specification=ConversationSimulationSpecification(
simulated_user_config=AgentConfig(
prompt=PromptAgent(
prompt="Your goal is to be a really difficult user.",
llm="gpt-4o",
temperature=0.5
)
)
),
extra_evaluation_criteria=[
PromptEvaluationCriteria(
id="politeness_check",
name="Politeness Check",
conversation_goal_prompt="The agent was polite.",
use_knowledge_base=False
)
]
)
print(response)

这是一个基础示例。如需完整的输入参数列表,请参阅 模拟对话 和 流式模拟对话 端点的 API 参考文档。

3

分析响应

SDK 提供一个完整的 JSON 对象,其中包含完整对话记录和详细分析。

模拟对话:记录模拟用户与智能体之间的每轮互动,包括消息和工具使用情况。

示例对话历史
[
...
{
"role": "user",
"message": "Maybe a little. I'll think about it, but I'm still not convinced it's the right move.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "I understand. If you want to explore more at your own pace, I can direct you to our documentation, which has guides and API references. Would you like me to send you a link?",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "user",
"message": "I guess it wouldn't hurt to take a look. Go ahead and send it over.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"params_as_json": "{\"path\": \"/docs/api-reference/introduction\"}",
"tool_has_been_called": false,
"tool_details": null
}
],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [],
"tool_results": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"result_value": "Tool Called.",
"is_error": false,
"tool_has_been_called": true,
"tool_latency_secs": 0
}
],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "Okay, I've sent you a link to the introduction to our API reference. It provides a good starting point for understanding our different tools and how they can be integrated. Let me know if you have any questions as you explore it.\n",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
}
...
]

分析:提供评估标准结果、数据收集指标以及对话记录摘要等洞察。

示例分析
{
"analysis": {
"evaluation_criteria_results": {
"politeness_check": {
"criteria_id": "politeness_check",
"result": "success",
"rationale": "The agent remained polite and helpful despite the user's challenging attitude."
},
"understood_root_cause": {
"criteria_id": "understood_root_cause",
"result": "success",
"rationale": "The agent acknowledged the user's hesitation and provided relevant information."
},
"positive_interaction": {
"criteria_id": "positive_interaction",
"result": "success",
"rationale": "The user eventually asked for the documentation link, indicating engagement."
}
},
"data_collection_results": {
"issue_type": {
"data_collection_id": "issue_type",
"value": "support_issue",
"rationale": "The user asked for help with integrating ElevenLabs tools."
},
"user_intent": {
"data_collection_id": "user_intent",
"value": "The user is interested in integrating ElevenLabs tools into a project."
}
},
"call_successful": "success",
"transcript_summary": "The user expressed skepticism, but the agent provided useful information and a link to the API documentation."
}
}
4

改进评估标准

全面审查模拟对话,评估评估标准的有效性。找出这些标准在评估智能体表现时可能存在的缺口或不足之处。相应地优化和调整评估标准,确保其符合预期结果,并能准确衡量智能体能力。

5

改进智能体

确认评估标准准确无误后,利用模拟对话中获得的经验提升智能体能力。可以优化系统提示词,更好地引导智能体响应,确保其符合你的目标和用户期望。此外,还可探索其他可优化的功能或配置,例如调整智能体语气、提升其处理特定问题的能力,或整合额外数据源来丰富其响应。系统地应用这些经验,可打造更强大、更有效的对话智能体,提供更出色的用户体验。

6

持续迭代

完成初始测试和改进周期后,建立一套全面的测试套件,可有效覆盖广泛的可能场景。该套件可通过不同的模拟用户提示词和初始条件,探索多种模拟对话。持续迭代和优化方法,确保智能体始终有效,并能响应不断变化的用户需求。

专业提示

详细的提示词和标准

编写详细、丰富的模拟用户提示词和评估标准,可提升模拟测试的效果。提供的上下文和细节越充分,智能体就越能理解并响应复杂互动。

模拟工具配置

使用模拟工具配置来测试智能体的决策过程。这样可以观察智能体如何决定发起工具调用,以及如何应对不同的工具调用结果。更多详情,请参阅 API 参考文档中的 tool_mock_config 输入参数。

部分对话历史

使用部分对话历史,评估智能体如何处理从特定节点开始的互动。这对于评估智能体处理以下对话的能力尤其有用:用户已通过特定方式提出问题,或某些工具调用已成功或失败。更多详情,请参阅 API 参考文档中的 partial_conversation_history 输入参数。