बातचीत सिम्युलेट करें

सिम्युलेटेड बातचीत के साथ अपने ElevenLabs एजेंट को टेस्ट और इवैल्यूएट करना सीखें

यह गाइड और इसमें इस्तेमाल किए गए एंडपॉइंट्स अब अप्रचलित हैं। इसके बजाय सिमुलेशन एजेंट टेस्ट टाइप का इस्तेमाल करें।

अवलोकन

ElevenLabs Agents API आपको अपने AI एजेंट के साथ टेक्स्ट-आधारित बातचीत का सिमुलेशन और मूल्यांकन करने देता है। यह गाइड आपको simulate conversation एंडपॉइंट्स (batch और streaming) का इस्तेमाल करके एंड-टू-एंड सिमुलेशन टेस्टिंग workflow लागू करना सिखाएगी। इससे आप अपने एजेंट की परफ़ॉर्मेंस को बारीकी से टेस्ट और बेहतर कर सकते हैं, ताकि यह आपके इंटरैक्शन लक्ष्यों को पूरा करे।

आवश्यकताएं

सिमुलेशन टेस्टिंग workflow लागू करना

1

शुरुआती मूल्यांकन पैरामीटर पहचानें

अपने एजेंट की बातचीत के इतिहास में उन उदाहरणों को खोजें जहां आपके एजेंट ने अच्छा प्रदर्शन नहीं किया। इन बातचीतों का इस्तेमाल एक सिम्युलेटेड यूज़र के लिए अलग-अलग प्रॉम्प्ट बनाने में करें, जो आपके एजेंट के साथ इंटरैक्ट करेगा। साथ ही, किसी खास सिम्युलेटेड यूज़र के लिए मनचाहे परिणामों को टेस्ट करने हेतु, एजेंट कॉन्फ़िगरेशन में पहले से तय न किए गए अतिरिक्त मूल्यांकन मानदंड भी तय करें।

2

SDK के ज़रिए बातचीत का सिमुलेशन करें

ElevenLabs SDK का इस्तेमाल करके सिमुलेशन एंडपॉइंट के लिए एक रिक्वेस्ट बनाएं।

from dotenv import load_dotenv
from elevenlabs import (
ElevenLabs,
ConversationSimulationSpecification,
AgentConfig,
PromptAgent,
PromptEvaluationCriteria
)
load_dotenv()
api_key = os.getenv("ELEVENLABS_API_KEY")
elevenlabs = ElevenLabs(api_key=api_key)
response = elevenlabs.conversational_ai.agents.simulate_conversation(
agent_id="YOUR_AGENT_ID",
simulation_specification=ConversationSimulationSpecification(
simulated_user_config=AgentConfig(
prompt=PromptAgent(
prompt="Your goal is to be a really difficult user.",
llm="gpt-4o",
temperature=0.5
)
)
),
extra_evaluation_criteria=[
PromptEvaluationCriteria(
id="politeness_check",
name="Politeness Check",
conversation_goal_prompt="The agent was polite.",
use_knowledge_base=False
)
]
)
print(response)

यह एक बुनियादी उदाहरण है। इनपुट पैरामीटर की पूरी सूची के लिए, कृपया Simulate conversation और Stream simulate conversation एंडपॉइंट्स का API रेफरेंस देखें।

3

रिस्पॉन्स का विश्लेषण करें

SDK एक विस्तृत JSON ऑब्जेक्ट देता है, जिसमें पूरी बातचीत की ट्रांसक्रिप्ट और विस्तृत विश्लेषण शामिल होता है।

सिम्युलेटेड बातचीत: इसमें सिम्युलेटेड यूज़र और एजेंट के बीच हर इंटरैक्शन टर्न, मैसेज और टूल उपयोग के विवरण के साथ शामिल होता है।

बातचीत के इतिहास का उदाहरण
[
...
{
"role": "user",
"message": "Maybe a little. I'll think about it, but I'm still not convinced it's the right move.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "I understand. If you want to explore more at your own pace, I can direct you to our documentation, which has guides and API references. Would you like me to send you a link?",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "user",
"message": "I guess it wouldn't hurt to take a look. Go ahead and send it over.",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"params_as_json": "{\"path\": \"/docs/api-reference/introduction\"}",
"tool_has_been_called": false,
"tool_details": null
}
],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": null,
"tool_calls": [],
"tool_results": [
{
"type": "client",
"request_id": "redirectToDocs_421d21e4b4354ed9ac827d7600a2d59c",
"tool_name": "redirectToDocs",
"result_value": "Tool Called.",
"is_error": false,
"tool_has_been_called": true,
"tool_latency_secs": 0
}
],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
},
{
"role": "agent",
"message": "Okay, I've sent you a link to the introduction to our API reference. It provides a good starting point for understanding our different tools and how they can be integrated. Let me know if you have any questions as you explore it.\n",
"tool_calls": [],
"tool_results": [],
"feedback": null,
"llm_override": null,
"time_in_call_secs": 0,
"conversation_turn_metrics": null,
"rag_retrieval_info": null,
"llm_usage": null
}
...
]

विश्लेषण: मूल्यांकन मानदंडों के परिणामों, डेटा कलेक्शन मेट्रिक्स और बातचीत की ट्रांसक्रिप्ट के सारांश की जानकारी देता है।

विश्लेषण का उदाहरण
{
"analysis": {
"evaluation_criteria_results": {
"politeness_check": {
"criteria_id": "politeness_check",
"result": "success",
"rationale": "The agent remained polite and helpful despite the user's challenging attitude."
},
"understood_root_cause": {
"criteria_id": "understood_root_cause",
"result": "success",
"rationale": "The agent acknowledged the user's hesitation and provided relevant information."
},
"positive_interaction": {
"criteria_id": "positive_interaction",
"result": "success",
"rationale": "The user eventually asked for the documentation link, indicating engagement."
}
},
"data_collection_results": {
"issue_type": {
"data_collection_id": "issue_type",
"value": "support_issue",
"rationale": "The user asked for help with integrating ElevenLabs tools."
},
"user_intent": {
"data_collection_id": "user_intent",
"value": "The user is interested in integrating ElevenLabs tools into a project."
}
},
"call_successful": "success",
"transcript_summary": "The user expressed skepticism, but the agent provided useful information and a link to the API documentation."
}
}
4

अपने मूल्यांकन मानदंड बेहतर करें

अपने मूल्यांकन मानदंडों की प्रभावशीलता का आकलन करने के लिए सिम्युलेटेड बातचीतों को अच्छी तरह से देखें। ऐसी कमियों या क्षेत्रों की पहचान करें जहां मानदंड एजेंट की परफ़ॉर्मेंस का मूल्यांकन करने में अपर्याप्त हो सकते हैं। अपने मनचाहे परिणामों के अनुरूप होने और एजेंट की क्षमताओं को सटीक रूप से मापने के लिए, मूल्यांकन मानदंडों को बेहतर बनाएं और समायोजित करें।

5

अपने एजेंट को बेहतर बनाएं

जब आपको अपने मूल्यांकन मानदंडों की सटीकता पर भरोसा हो जाए, तो एजेंट की क्षमताओं को बेहतर बनाने के लिए सिम्युलेटेड बातचीतों से मिली सीख का इस्तेमाल करें। एजेंट के रिस्पॉन्स को बेहतर दिशा देने के लिए सिस्टम प्रॉम्प्ट को बेहतर बनाने पर विचार करें, ताकि वे आपके उद्देश्यों और यूज़र की अपेक्षाओं के अनुरूप हों। साथ ही, अन्य फ़ीचर्स या कॉन्फ़िगरेशन देखें जिन्हें ऑप्टिमाइज़ किया जा सकता है, जैसे एजेंट के टोन को समायोजित करना, खास क्वेरीज़ संभालने की इसकी क्षमता बेहतर करना या इसके रिस्पॉन्स को समृद्ध बनाने के लिए अतिरिक्त डेटा सोर्स जोड़ना। इन सीखों को व्यवस्थित रूप से लागू करके, आप अधिक मज़बूत और प्रभावी कन्वर्सेशनल AI एजेंट बना सकते हैं, जो बेहतर यूज़र अनुभव देता है।

6

निरंतर सुधार

शुरुआती टेस्टिंग और सुधार चक्र पूरा करने के बाद, एक व्यापक टेस्टिंग सुइट बनाना संभावित स्थितियों की विस्तृत रेंज को कवर करने का अच्छा तरीका हो सकता है। यह सुइट अलग-अलग सिम्युलेटेड यूज़र प्रॉम्प्ट और शुरुआती स्थितियों के साथ कई सिम्युलेटेड बातचीतों का परीक्षण कर सकता है। अपने तरीके में लगातार सुधार और बदलाव करके, आप यह सुनिश्चित कर सकते हैं कि आपका एजेंट यूज़र की बदलती ज़रूरतों के लिए प्रभावी और उत्तरदायी बना रहे।

प्रो टिप्स

विस्तृत प्रॉम्प्ट और मानदंड

विस्तृत और वर्णनात्मक सिम्युलेटेड यूज़र प्रॉम्प्ट और मूल्यांकन मानदंड तैयार करने से सिमुलेशन टेस्ट की प्रभावशीलता बढ़ सकती है। आप जितना अधिक संदर्भ और विशिष्टता देंगे, एजेंट उतना ही बेहतर तरीके से जटिल इंटरैक्शन समझ और उनका जवाब दे सकेगा।

मॉक टूल कॉन्फ़िगरेशन

अपने एजेंट की निर्णय लेने की प्रक्रिया को टेस्ट करने के लिए मॉक टूल कॉन्फ़िगरेशन का इस्तेमाल करें। इससे आप देख सकते हैं कि एजेंट टूल कॉल करने का फैसला कैसे करता है और अलग-अलग टूल कॉल परिणामों पर कैसे प्रतिक्रिया देता है। अधिक जानकारी के लिए, API रेफरेंस में tool_mock_config इनपुट पैरामीटर देखें।

आंशिक बातचीत का इतिहास

किसी खास बिंदु से एजेंट इंटरैक्शन को कैसे संभालते हैं, इसका मूल्यांकन करने के लिए आंशिक बातचीत के इतिहास का इस्तेमाल करें। यह खास तौर पर उस स्थिति में एजेंट की बातचीत प्रबंधित करने की क्षमता का आकलन करने के लिए उपयोगी है, जहां यूज़र पहले ही किसी खास तरीके से सवाल रख चुका है या कुछ टूल कॉल सफल अथवा विफल हो चुके हैं। अधिक जानकारी के लिए, API रेफरेंस में partial_conversation_history इनपुट पैरामीटर देखें।