탐색으로 건너뛰기

WebSocket

Stream expressive dialogue audio over a WebSocket by sending incremental text segments per registered voice.

The connection uses Eleven v3 dialogue models only (model_id must start with eleven_v3). The default model is eleven_v3_conversational.

Session setup

  • After connecting, the first JSON message must include voices (voice IDs to register for the session) and credentials if not already sent via headers or query string.
  • Optional voice_settings and pronunciation_dictionary_locators are only accepted on the first message.
  • For eleven_v3_conversational, only one voice ID may be registered. For eleven_v3, you may register up to 10 voices.

Streaming text

  • Send inputs: an array of { "text", "voice_id", "new_turn"? }. Text for the same turn is buffered until the server has enough context (at least ~40 characters and 8 words), then partial audio chunks are emitted.
  • Set new_turn to true (or switch voice_id) to finalize the current prosody segment and start a new speaker turn.

Control messages

  • flush: force generation of any buffered text without closing the socket.
  • close_socket: flush remaining audio, send a final message, and close the connection.
  • keep_alive: reset the 20 second receive timeout (no generation).

Authentication

Use the xi-api-key or Authorization header, single_use_token query parameter, or include xi_api_key, authorization, or single_use_token in the first message body (same pattern as Text to Speech WebSocket). Anonymous sessions are rejected.

For non-streaming dialogue over HTTP, see Create dialogue and Stream dialogue.

핸드셰이크

WSS
/v1/text-to-dialogue/stream-input

헤더

xi-api-keystring선택 사항

쿼리 매개변수

model_idstring선택 사항기본값 eleven_v3_conversational

Identifier of the model that will be used, you can query them using GET /v1/models. Must be a v3 model.

output_formatany선택 사항

Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32.

language_codestring선택 사항

Language code (ISO 639-1) used to enforce a language for the model and text normalization. If the model does not support the provided language code, it will be ignored. This parameter is not supported for multilingual_v2 models.

sync_alignmentboolean선택 사항기본값 false

When true, character timing from the model may be attached to audio chunks as alignment (snake_case field names).

apply_text_normalizationany선택 사항
A string parameter to control text normalization.
seedinteger선택 사항1-4294967295
If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed.
enable_loggingboolean선택 사항기본값 true

When enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.

보내기

text_to_dialogue_websocket_clientobject필수

JSON text frame from client. The first frame must register voices; later frames stream inputs and control flags.

수신

text_to_dialogue_websocket_audio_chunkobject필수

Partial audio for the current generation, base64-encoded using output_format.

OR
text_to_dialogue_websocket_final_audio_for_turnobject필수

Sent after the final audio for a given turn has been sent. For streaming codecs (MP3, Opus, etc.), the persistent encoder buffers a small amount of audio across turn boundaries, so a few bytes belonging to the marked turn may arrive interleaved with the next turn's first audio chunk. Clients that need exact per-turn boundaries should request a PCM output format, which has no buffering.

OR
text_to_dialogue_websocket_finalobject필수

Sent after close_socket once all audio has been flushed.

OR
text_to_dialogue_websocket_errorobject필수
Error payload sent before the socket is closed with a WebSocket close code.