Stream audio from a microphone or a file and receive expression measurements as speech happens. You do not need to detect speech or split it into segments: the server detects each stretch of speech and measures it while it is still being spoken.
Send audio as binary messages of signed 16-bit little-endian PCM, mono, 16 kHz, with no header. Frames of around 100 ms work well. You can cut the audio at any sample boundary; the server joins the frames into one continuous stream.
session.update before your first audio frame. A session.update after your first audio frame ends the session with config_invalid.invalid_audio_frame, and the session continues. A message over 1MiB ends the session with message_too_large. Accepted frames are not acknowledged.rate_limited, and none of its audio is measured. Wait briefly and resend the same frame, so that the stream stays continuous. To measure a recording faster than it plays, use the audio file endpoint.internal_error and the session ends. Open a new connection.The server detects when speech starts and stops. Each continuous stretch of speech is an utterance, bracketed by utterance.start and utterance.end.
measurement.result arrives every measurement_timer_ms milliseconds of speech, with trigger set to interval.trigger set to speech_ended, unless the previous measurement ended less than one second earlier. Every utterance gets at least one measurement.audio_start_ms and audio_end_ms give the span, counted from the start of the audio the server has accepted.session.close during speech ends the utterance. You receive its final measurement, if one is due, then utterance.end, then session.closed. A session that ends with an error can leave its last utterance without an utterance.end.A session at the default measurement_timer_ms of 3000, with one utterance: speech begins 98 ms into the audio and is still in progress at 26000 ms, when session.close ends it.
The format of the audio you send. Only one format is supported; any other value is rejected with config_invalid.
The interval between measurements while speech is in progress, in milliseconds of speech. Lower values produce measurements more often; at the maximum, interval measurements are back to back rather than overlapping.
Sent when the connection opens, before you send anything. Carries the session ID and the default configuration, which stays in effect unless you change it with session.update.
Unique identifier for the session. It is also the run_id of the run the session is recorded as. Include it when contacting support about a session.
Sent when a session.update is accepted. Carries the full configuration now in effect, including defaults for any fields you omitted.
Sent when speech starts. Every measurement.result until the matching utterance.end belongs to this utterance.
Sent every measurement_timer_ms of speech while an utterance is in progress, and once more when the utterance ends unless the previous measurement ended less than one second earlier. Each measurement covers up to the most recent 10 seconds of the utterance.
The utterance this measurement belongs to, as given in utterance.start.
Why the measurement was produced.
interval: the measurement interval (measurement_timer_ms) elapsed while speech continued.speech_ended: the utterance ended, or the audio ended while speech was in progress.A final measurement is skipped when the previous one ended less than one second earlier, so the last measurement of an utterance may carry the trigger interval. Where an utterance ended is reported separately: as an utterance.end message in a realtime session, and in utterances in an audio file response.
Expressions detected in the speech, such as joy, anger and chuckling, in descending order of probability. Scores with equal probability are ordered by the model’s underlying score, then by name. Only expressions judged present are included, each against a threshold set for that expression, so the array may be empty. probability is the measured chance that the expression is truly present.
Qualities of the voice and the audio, such as husky, fast and background music, in descending order of probability. Scores with equal probability are ordered by the model’s underlying score, then by name. Only scores of 0.725 or above are included, so the array may be empty. probability is calibrated: the expected intensity a human rater would give the quality.
Sent when speech stops, after the utterance’s final measurement.result, so every measurement for the utterance has arrived by the time you receive it.
The utterance that ended, as given in utterance.start.
Where the utterance started, in milliseconds from the start of the session’s audio. Matches the utterance.start message.
Identifies what went wrong and whether the session continues.
Recoverable. The message is rejected and the session continues.
invalid_audio_frame: the audio frame was empty, had an odd number of bytes, or held more than 10 seconds of audio. Check how you slice audio and keep sending. Audio only.invalid_image_frame: the frame was empty, was not a JPEG, or exceeded the pixel limit. Check the image and keep sending. Video only.message_invalid: a text message was not valid JSON, had no type, or had an unrecognized type. Correct the message and keep sending.rate_limited: media arrived faster than the send rate and was not processed. Audio: wait briefly and send the same frame again, so that the stream stays continuous. Video: from a live camera, drop the frame and send the next; from a recording, wait briefly and send the same frame again.Depends on the endpoint.
internal_error: the server failed while processing media. Audio: the session ends as for a session-ending error; open a new connection. Video: the frame is skipped and the session continues; resend it if you need a result for it. Five consecutive failures end the session; frames that are rejected or have no faces to measure do not reset the count.Session-ending. error is followed by session.closed with reason error, and the connection closes with code 1008.
config_invalid: a session.update contained an unrecognized field or an unsupported value, or arrived after you started sending media. Correct the message and open a new connection.message_too_large: a message exceeded the size limit. Reduce the frame size and open a new connection.A human-readable description of the error, intended for logging and debugging. The wording may change; use code for programmatic handling.
false for every error that ends the session. true for every error the session survives, except internal_error on the video endpoint, where it reports whether the failure was temporary and the frame is worth resending.
Sent as the last message of the session, after any final measurement and utterance.end. Carries the reason the session ended and totals of what was received and produced. The connection closes immediately after it, with the close code for the reason.
Unique identifier for the session. It is also the run_id of the run the session is recorded as. Include it when contacting support about a session.
Why the session ended. session.closed is the last message, and the WebSocket close code that follows it depends on the reason.
client_request: you sent session.close. Close code 1000.server_shutdown: the server is restarting; reconnect at once to start a new session. Close code 1012.idle_timeout: the server received no media frame it could use for 30 seconds; open a new session to continue. Close code 1000.error: a session-ending error, reported in the preceding error message. Close code 1008.Replaces the whole configuration; any field you omit returns to its default. Send it before your first audio frame, because once audio has been accepted a session.update ends the session with config_invalid. The server confirms with session.updated, or ends the session with config_invalid if the message contains an unrecognized field or an unsupported value.
Ends the session. The server sends any final results, then session.closed, and closes the connection with code 1000; on the audio endpoint, if the final measurement fails, error and session.closed with reason error follow instead. Prefer this to closing the WebSocket yourself, which ends the session immediately and discards any final results and the totals.
A binary message of signed 16-bit little-endian PCM audio, mono, 16 kHz, with no header. The server appends it to the session’s audio and measures speech as it is detected; accepted frames are not acknowledged. A frame that is empty, does not contain whole samples, or exceeds the frame limit is rejected with invalid_audio_frame, and a frame that arrives faster than the send rate can be rejected with rate_limited; in both cases the session continues.
Sent when a message you sent is rejected or the server fails to process it. For a recoverable error the message is rejected and the session continues; for a session-ending error, session.closed with reason error follows. ErrorCode lists each code, its class, and what to do.