탐색으로 건너뛰기

Realtime

Realtime speech-to-text transcription service. This WebSocket API enables streaming audio input and receiving transcription results.

Event Flow

  • Audio chunks are sent as input_audio_chunk messages
  • Transcription results are streamed back as partial_transcript (interim) and committed_transcript (stable/final for that segment)
  • Supports manual commit or VAD-based automatic commit strategies

Authentication is done either by providing a valid API key in the xi-api-key header or by providing a valid token in the token query parameter. Tokens can be generated from the single use token endpoint. Use tokens if you want to transcribe audio from the client side.

핸드셰이크

WSS
/v1/speech-to-text/realtime

헤더

xi-api-keystring선택 사항

쿼리 매개변수

model_idenum필수기본값 scribe_v2_realtime

The ID of the model to use for speech-to-text transcription.

허용된 값:
tokenstring선택 사항

Single use token for authentication. Only used when initiating a session from the client. If provided, xi-api-key is no longer required for authentication.

audio_formatany선택 사항

The encoding format of the audio. Supported formats: pcm_8000, pcm_16000, pcm_22050, pcm_24000, pcm_44100, pcm_48000, ulaw_8000.

language_codestring선택 사항

An ISO-639-1 or ISO-639-3 language_code corresponding to the language of the audio file. Can sometimes improve transcription performance if known beforehand. Defaults to null, in this case the language is predicted automatically.

secondary_languageslist of strings선택 사항

Additional ISO-639-1 or ISO-639-3 language codes that may be present in the audio. Providing them makes language identification more reliable by only focusing on a certain set of languages. Each code is validated the same way as language_code.

commit_strategyenum선택 사항

Commit strategy for speech segmentation. 'manual' requires explicit commits, 'vad' automatically segments speech using silence detection, 'turn_prediction' additionally commits as soon as the model predicts the speaker has finished their turn (only for models that predict turns, where it is the default).

허용된 값:
vad_thresholddouble선택 사항
VAD sensitivity threshold for detecting speech activity. Lower values are more sensitive to speech.
vad_silence_threshold_secsdouble선택 사항
Duration of silence in seconds required to trigger a commit when VAD commit strategy is enabled. Longer values result in fewer commits but longer segments.
min_speech_duration_msinteger선택 사항
Minimum duration of speech in milliseconds required to be considered valid speech by VAD.
min_silence_duration_msinteger선택 사항
Minimum duration of silence in milliseconds required to be considered a speech break by VAD.
include_timestampsboolean선택 사항기본값 false

Enable word/character-level timestamps in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with timestamps after each commit. Default: false.

include_language_detectionboolean선택 사항기본값 false

Enable language detection in a delayed committed_transcript_with_timestamps message. When enabled, you'll receive an additional message with detected language_code after each commit. Default: false.

keytermslist of strings선택 사항

List of keyterms to bias the model towards. Maximum 50 keyterms. Adds a 20% premium to the base transcription cost.

no_verbatimboolean선택 사항기본값 false
If true, removes filler words, false starts and disfluencies from the transcript.
entity_detectionstring or list of strings선택 사항

Detect entities on committed transcripts. Can be 'all', a single entity type or category, or a list of types/categories ('pii', 'phi', 'pci', 'other', 'offensive_language'). When enabled, detected entities are delivered in a separate 'committed_transcript_entities' event with their text, type, and character positions.

transcript_editstring선택 사항

Natural-language instruction applied to each committed transcript (max 2000 characters). The edited text is delivered in a separate 'edited_transcript' event containing the committed text and the edited text. Cannot be combined with entity_detection. Adds a 30% premium to the base transcription cost, billed for at least 10 seconds of audio per committed transcript.

filter_background_audioboolean선택 사항기본값 false

Enable background speech filtering to reduce false activations from nearby conversations and ambient noise. When enabled without an explicit vad_threshold, a lower default threshold is applied. Cannot be combined with include_timestamps.

keepalive_interval_msinteger선택 사항500-10000

Opt-in keepalive interval in milliseconds (500-10000). While the audio being streamed contains no speech, the server sends a partial_transcript about this often, so clients that enforce a read timeout on incoming frames do not drop the connection during long pauses. The keepalive has empty text (text: ""), or repeats the latest partial text if the current segment is not committed yet. Keepalives are only sent after the transcription models have processed silent audio, so audio must keep streaming. Cadence is rounded to the audio processing granularity (about one second). Disabled by default.

enable_loggingboolean선택 사항기본값 true

When enable_logging is set to false zero retention mode will be used for the request. This will mean history features are unavailable for this request, including request stitching. Zero retention mode may only be used by enterprise customers.

보내기

inputAudioChunkobject필수
Audio data chunk sent from client to server for transcription.

수신

sessionStartedobject필수
Sent when the transcription session is successfully started.
OR
partialTranscriptobject필수
Interim transcription result that may change.
OR
committedTranscriptobject필수
Committed transcription result that will not change.
OR
committedTranscriptWithTimestampsobject필수

Committed transcription result with word-level timestamps.

OR
committedTranscriptEntitiesobject필수

Detected entities for a committed transcript. Only sent when the entity_detection query parameter is set.

OR
Edited Transcriptobject필수

Edited version of a committed transcript, delivered as a separate event when the transcript_edit query parameter is set. text is the committed text the instruction was applied to (use it to correlate with the committed_transcript event, since edits may arrive out of commit order). If no edits were made, edited_text is identical to text.

OR
Scribe Warningobject필수

Non-fatal notice. Sent after session_started when enable_logging=false was requested but zero retention mode was not applied; the session continues and is still logged.

OR
scribeErrorobject필수
Error event during transcription.
OR
scribeAuthErrorobject필수
Authentication error during transcription session.
OR
scribeQuotaExceededErrorobject필수
Quota exceeded error during transcription session.
OR
scribeThrottledErrorobject필수
Throttled error during transcription session.
OR
scribeUnacceptedTermsErrorobject필수
Unaccepted terms error during transcription session.
OR
scribeRateLimitedErrorobject필수
Rate limited error during transcription session.
OR
scribeQueueOverflowErrorobject필수
Queue overflow error during transcription session.
OR
scribeResourceExhaustedErrorobject필수
Resource exhausted error during transcription session.
OR
scribeSessionTimeLimitExceededErrorobject필수
Session time limit exceeded error during transcription session.
OR
scribeInputErrorobject필수
Input error during transcription session.
OR
Scribe Invalid Request Errorobject필수

The connection parameters were rejected; the session is closed afterwards.

OR
scribeChunkSizeExceededErrorobject필수
Chunk size exceeded error during transcription session.
OR
scribeInsufficientAudioActivityErrorobject필수
Insufficient audio activity error during transcription session.
OR
scribeTranscriberErrorobject필수
Transcriber error during transcription session.