탐색으로 건너뛰기

WebSocket

The Text-to-Speech WebSockets API is designed to generate audio from partial text input while ensuring consistency throughout the generated audio. Although highly flexible, the WebSockets API isn’t a one-size-fits-all solution. It’s well-suited for scenarios where:

  • The input text is being streamed or generated in chunks.
  • Word-to-audio alignment information is required.

However, it may not be the best choice when:

  • The entire input text is available upfront. Given that the generations are partial, some buffering is involved, which could potentially result in slightly higher latency compared to a standard HTTP request.
  • You want to quickly experiment or prototype. Working with WebSockets can be harder and more complex than using a standard HTTP API, which might slow down rapid development and testing.

핸드셰이크

WSS
/v1/text-to-speech/:voice_id/stream-input

헤더

xi-api-keystring선택 사항

경로 매개변수

voice_idstring필수
The unique identifier for the voice to use in the TTS process.

쿼리 매개변수

authorizationstring선택 사항
Your authorization bearer token.
single_use_tokenstring선택 사항

Your single use token. Use this if you want to initiate a session from the client. When providing this parameter, xi-api-key is no longer required for authentication.

model_idstring선택 사항기본값 eleven_multilingual_v2

Identifier of the model that will be used, you can query them using GET /v1/models.

language_codestring선택 사항

Language code (ISO 639-1) used to enforce a language for the model and text normalization. If the model does not support the provided language code, it will be ignored. This parameter is not supported for multilingual_v2 models.

enable_loggingboolean선택 사항기본값 true

When enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.

output_formatany선택 사항

Output format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pro tier or above. Note that the μ-law format (sometimes written mu-law, often approximated as u-law) is commonly used for Twilio audio inputs.

inactivity_timeoutinteger선택 사항<=180기본값 20
The number of seconds that the connection can be inactive before it is automatically closed. The default timeout is set to 20, with a maximum allowed value of 180.
sync_alignmentboolean선택 사항기본값 false
Sync the text alignment to every returned response
auto_modeboolean선택 사항기본값 false
Whether to use auto mode for this request. This setting focuses on reducing the latency by disabling the chunk schedule and all buffers. It is only recommended when sending full sentences, sending partial sentences will result in highly reduced quality.
apply_text_normalizationany선택 사항

This parameter controls text normalization with three modes: 'auto', 'on', and 'off'. When set to 'auto', the system will automatically decide whether to apply text normalization (e.g., spelling out numbers). With 'on', text normalization will always be applied, while with 'off', it will be skipped.

seedinteger선택 사항0-4294967295
If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed.
enable_ssml_parsingboolean선택 사항기본값 false

Whether to enable/disable parsing of SSML tags within the provided text. For best results, we recommend sending SSML tags as fully contained messages to the websockets endpoint, otherwise this may result in additional latency. Please note that rendered text, in normalizedAlignment, will be altered in support of SSML tags. The rendered text will use a . as a placeholder for breaks, and syllables will be reported using the CMU arpabet alphabet where SSML phoneme tags are used to specify pronunciation. IMPORTANT: When using phoneme-based pronunciation dictionaries (IPA/CMU), SSML parsing is automatically enabled if this parameter is not set. Setting this to false with phoneme dictionaries is deprecated and will be ignored in a future release, as phoneme dictionaries require SSML parsing to work correctly.

보내기

initializeConnectionobject필수
OR
sendTextobject필수
OR
closeConnectionobject필수

수신

audioOutputobject필수
OR
finalOutputobject필수