Audio realtime
The audio realtime endpoint measures emotional expression in speech while you stream the audio. Send audio as binary WebSocket frames, and while each utterance continues, the server sends a measurement.result at a fixed interval. To measure a recording in one request, use Audio upload.
Audio format
The endpoint accepts one audio format, sent as binary WebSocket frames.
A frame must hold whole samples, so its length is an even number of bytes. It can hold up to 10 seconds of audio (320 KB). Any WebSocket message larger than 1 MiB (1,048,576 bytes), audio or text, ends the session with message_too_large. Frames of about 100 ms (3,200 bytes) work well. The server joins frames into one continuous stream, so a frame may begin or end anywhere in the audio.
Send rate
Stream audio in close to real time, no faster than 1 second of audio per second. Audio from a microphone arrives at this rate. To stream a recording, send one 100 ms frame every 100 ms. To measure a recording faster than it plays, send it to Audio upload instead.
A frame sent faster than the send rate can be rejected with rate_limited, and none of its audio is measured. Wait briefly and send the same frame again, so that the stream stays continuous.
Utterances
An utterance is one continuous stretch of speech. The server marks its start with utterance.start, sends one or more measurement.result messages while it continues, and marks its end with utterance.end. Silence between utterances produces no messages. A session that ends with an error can leave its last utterance without an utterance.end.
utterance_id starts at 0 and increases through the session. Every measurement.result names the utterance it belongs to.
Measurements
While an utterance is in progress, the server measures it every measurement_timer_ms of speech, 3000 by default. A measurement covers the most recent 10 seconds of the utterance, or the whole utterance if it is shorter, so consecutive measurements overlap. At the maximum interval of 10000 ms they are back to back. An interval measurement taken as speech stops can end shortly after the utterance does, and no final measurement follows it.
When an utterance ends, the server sends a final measurement covering up to its last 10 seconds, unless less than one second of speech has passed since the previous measurement ended. Every utterance produces at least one measurement.
Sequential within the session, starting at 0.
The utterance this measurement belongs to.
The window measured, in milliseconds from the start of the audio the server has accepted. The window never starts before the utterance does, but an interval measurement taken as speech stops can end shortly after it.
interval when the measurement interval elapsed while speech continued. speech_ended when the utterance ended, or when the session was closed mid-speech.
The 414 expression names are listed in Audio measurements.
How the voice and the recording sound. The 190 names are listed in Audio measurements.
Because the final measurement can be skipped, the last measurement of an utterance can have a trigger of interval. Detect the end of an utterance with utterance.end, not with trigger.
Configuration
To change the measurement interval, send session.update before the first audio frame. The configuration locks at the first audio frame the server accepts; a rejected frame does not lock it. The rules for session.update, including that omitted fields return to their defaults, are in Sessions.
The interval between measurements, in milliseconds of speech, from 3000 to 10000.
The audio format. Only the values shown are supported; anything else is rejected with config_invalid. May be omitted or sent as {}.
Message flow
A session that streams 26 seconds of continuous speech with the default configuration exchanges the messages below. Times are positions in the audio.
Session summary
received.frames counts accepted frames, so subtracting it from the number you sent gives the number rejected. received.audio_duration_ms counts all accepted audio, whether or not it contained speech. produced.measurements counts measurement.result messages.
Errors
The audio endpoint sends the shared codes, plus the code below. The internal_error row describes how the shared code behaves on this endpoint.

