Realtime
Realtime speech-to-text transcription service. This WebSocket API enables streaming audio input and receiving transcription results.
Event Flow
- Audio chunks are sent as
input_audio_chunkmessages - Transcription results are streamed back as
partial_transcript(interim) andcommitted_transcript(stable/final for that segment) - Supports manual commit or VAD-based automatic commit strategies
Authentication is done either by providing a valid API key in the xi-api-key header or by providing a valid token in the token query parameter. Tokens can be generated from the single use token endpoint. Use tokens if you want to transcribe audio from the client side.
핸드셰이크
헤더
쿼리 매개변수
The ID of the model to use for speech-to-text transcription.
Single use token for authentication. Only used when initiating a session from the client. If provided, xi-api-key is no longer required for authentication.
The encoding format of the audio. Supported formats: pcm_8000, pcm_16000, pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000.
An ISO-639-1 or ISO-639-3 language_code corresponding to the language of the audio file. Can sometimes improve transcription performance if known beforehand. Defaults to null, in this case the language is predicted automatically.
Additional ISO-639-1 or ISO-639-3 language codes that may be present in the audio. Providing them makes language identification more reliable by only focusing on a certain set of languages. Each code is validated the same way as language_code.
Commit strategy for speech segmentation. 'manual' requires explicit commits, 'vad' automatically segments speech using silence detection, 'turn_prediction' additionally commits as soon as the model predicts the speaker has finished their turn (only for models that predict turns, where it is the default).
Enable word/character-level timestamps in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with timestamps after each commit. Default: false.
Enable language detection in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with detected language_code after each commit. Default: false.
List of keyterms to bias the model towards. Maximum 50 keyterms. Adds a 20% premium to the base transcription cost.
Detect entities on committed transcripts. Can be 'all', a single entity type or category, or a list of types/categories ('pii', 'phi', 'pci', 'other', 'offensive_language'). When enabled, detected entities are delivered in a separate 'committed_transcript_entities' event with their text, type, and character positions.
Natural-language instruction applied to each committed transcript (max 2000 characters). The edited text is delivered in a separate 'edited_transcript' event containing the committed text and the edited text. Cannot be combined with entity_detection. Adds a 30% premium to the base transcription cost, billed for at least 10 seconds of audio per committed transcript.
Enable background speech filtering to reduce false activations from nearby conversations and ambient noise. When enabled without an explicit vad_threshold, a lower default threshold is applied. Cannot be combined with include_timestamps.
Opt-in keepalive interval in milliseconds (500-10000). While the audio being streamed contains no speech, the server sends a partial_transcript about this often, so clients that enforce a read timeout on incoming frames do not drop the connection during long pauses. The keepalive has empty text (text: ""), or repeats the latest partial text if the current segment is not committed yet. Keepalives are only sent after the transcription models have processed silent audio, so audio must keep streaming. Cadence is rounded to the audio processing granularity (about one second). Disabled by default.
When enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.
보내기
수신
Committed transcription result with word-level timestamps.
Detected entities for a committed transcript. Only sent when the entity_detection query parameter is set.
Edited version of a committed transcript, delivered as a separate event when the transcript_edit query parameter is set. text is the committed text the instruction was applied to (use it to correlate with the committed_transcript event, since edits may arrive out of commit order). If no edits were made, edited_text is identical to text.
Non-fatal notice. Sent after session_started when enable_logging=false was requested but zero retention mode was not applied; the session continues and is still logged.
The connection parameters were rejected; the session is closed afterwards.