Video realtime
The video realtime endpoint measures emotional expression on faces in video you stream, such as a camera feed. Send the video one frame at a time, each as a JPEG in a binary WebSocket frame, and each frame the server measures produces one measurement.result listing the faces found in it. To measure a set of images in one request, use Image upload.
Frames
Each binary WebSocket frame holds one complete JPEG image. To measure a camera or video feed, capture frames, encode each one as a JPEG, and send them one at a time, up to the send rate.
A frame that is empty, is not a valid JPEG, or exceeds the pixel limit is rejected with invalid_image_frame, and the session continues. A WebSocket message larger than 2 MB cannot be read at all and ends the session with message_too_large, so check the encoded size before sending.
Send rate
Send frames as they are captured or played, up to 3 frames per second. To measure a recorded video faster than it plays, extract its frames and send them to Image upload, passing face_state from one request to the next to keep face IDs.
A frame sent faster than the send rate can be rejected with rate_limited and is not measured. From a live camera, drop the frame and send the next one. From a recording, wait briefly and send the same frame again.
A frame within the send rate is never rejected because the server is busy. It waits until the server can measure it, so under heavy load replies take longer instead.
Matching results to frames
Replies arrive in the order the frames were sent, one reply per frame, so pair them by position. Keep a queue of the frames you have sent, and remove the oldest for each measurement.result and for each error whose code is invalid_image_frame, rate_limited, or internal_error. Other errors, such as message_invalid, reply to a text message and do not consume a frame. An error that ends the session needs no pairing. Frames sent before session.close receive their replies before session.closed.
frame_id counts results. It starts at 0 and advances only when a frame produces a measurement.result, so pair replies by position rather than by frame_id.
Detection
The server scans each frame for faces and excludes any whose detection confidence is below face.threshold or whose bounding box is shorter than face.min_size pixels on its shorter side. It measures the largest of the remaining faces, up to 2 by default, and lists them first in faces, largest first. The faces that were not measured follow, most confident first, up to 32 faces in all.
A face that was not measured carries its bbox and confidence, with face_id, expressions, and descriptions set to null. The limit on measured faces is set for your organization and does not appear in any message. To raise it, contact support.
Sequential within the session, starting at 0.
Identifies the same face from frame to frame within the session. null for a face that was not measured. See Face tracking.
The face’s location in the submitted image as [x0, y0, x1, y1]: the top-left and bottom-right corners in pixels, measured from the top-left of the image. Coordinates are in the pixels as stored in the file: EXIF orientation is not applied.
The detector’s confidence that the box contains a face, from 0 to 1. Never below face.threshold. Unrelated to the expression scores.
The 48 expression names are listed in Video measurements. null for a face that was not measured.
Visible facial actions such as smile or jaw drop. The 27 names are listed in Video measurements. null for a face that was not measured.
A frame with no faces produces a measurement.result with an empty faces array.
Face tracking
face_id links the same face across frames within a session. IDs are assigned from 0 in the order faces are first measured, largest first within a frame. A face normally keeps its ID while it is measured, and can get it back after it leaves the view or is no longer among the largest faces. The server remembers up to 32 faces, so a face that returns after many others have been measured may receive a new ID. IDs restart in every session and do not identify a person.
Configuration
To change the detection settings, send session.update before the first frame. The configuration locks at the first frame that passes the send-rate check, even if that frame is then rejected with invalid_image_frame because it is not a JPEG. An empty frame, or one rejected with rate_limited, does not lock it. The rules for session.update, including that omitted fields return to their defaults, are in Sessions.
The minimum detection confidence for a face to be measured, from 0 to 1. Lower values include more faces at the risk of false detections. The scale is specific to this API, so thresholds tuned for other face detection APIs do not carry over. Reported rounded to four decimal places.
The smallest face to measure, as the shorter side of its bounding box in pixels of the submitted image. At least 1.
The image format. Only image/jpeg is supported; anything else is rejected with config_invalid. May be omitted or sent as {}.
Message flow
A session that sends four frames with the default configuration exchanges the messages below. The third frame is a PNG sent by mistake. Face 0 appears in the first two frames, and face 1 in the second and fourth.
The rejected frame receives no frame_id, so the last result is frame_id 2. In session.closed, received.frames is 3 because it counts only frames accepted for measurement, and produced.measurements is 4 because it counts measured faces rather than frames.
Session summary
received.frames counts frames accepted for detection, including any whose measurement then failed with internal_error, so it can exceed the number of results. produced.measurements counts only measured faces across all frames, so a frame with two measured faces counts as two and a frame with none as zero. The audio endpoint’s field of the same name counts measurement.result messages instead.
Errors
The video endpoint sends the shared codes, plus the code below. The internal_error and message_too_large rows describe how those shared codes behave on this endpoint.

