Expression Measurement API
The API is currently in limited release; contact our team to request access.
The Expression Measurement API measures emotional expression in speech and in faces, with dedicated audio and video endpoints. The audio endpoints detect utterances in audio, and the video endpoints detect faces in images and video frames. Every score is a probability from 0 to 1, and Scores explains what it denotes in each array.
What it measures
From speech, the API returns expressions such as amusement and vocal qualities such as monotone. From an image, it returns expressions and visible facial actions, such as smile, for the largest faces it detects, 2 by default.
Learn more about the science behind expression measurement in Hume’s research and publications.
Endpoints
The endpoints fall into three groups. Upload endpoints measure media you already have, realtime endpoints measure media as it streams, and run endpoints return the record of each request and session.
Upload
The upload endpoints take audio or images you already have and return every measurement in one response. Use them for recorded media, for example to label a dataset or to evaluate generated audio.
Send the request
Send a multipart/form-data request with the media in file parts and, to change the defaults, settings in a JSON config part. Put your API key in the X-Hume-Api-Key header. See Authentication.
Read the response
The response lists the measurements: one for each window of speech in the audio, or one for each image with the faces found in it. Scores explains how to read the scores. The request_id identifies the request, and the run_id identifies the run it was recorded as.
Realtime
The realtime endpoints measure media streamed over a WebSocket and send each measurement as soon as it is produced. Stream audio to the audio endpoint from a source such as a microphone, or stream video to the video endpoint one JPEG frame at a time, such as from a camera. Use them to act on media as it happens, for example to route or monitor a live call.
Both endpoints share one protocol, described in Sessions.
Connect
Open a WebSocket to the audio or video realtime endpoint with your API key in the X-Hume-Api-Key header. The server immediately sends session.created, which contains the session ID and the default configuration.
Configure, if needed
To change a setting, such as the measurement interval or the face detection threshold, send session.update before the first media frame. The configuration locks when the server takes the first frame. Skip this step to use the defaults.
Stream media
Send audio or JPEG images as binary WebSocket frames. The server detects speech in the audio, or faces in each image, and measures them.
Runs
Each measured upload request and each realtime session is recorded as a run. The run endpoints return a run’s status, configuration, and event log. Measurements come only from the upload response or the session, so store the results you need.
SDKs
Hume publishes SDKs for Python and Node.js. Both cover every endpoint and include audio helpers that record from a microphone, read WAV files, and convert audio to the format the audio endpoints accept.
Compatibility
The API version is the path prefix, /v1/. Within a version:
- The server may add fields to the responses and messages it sends. Ignore fields you do not recognize.
- New prediction names may be added. Treat a name you do not recognize as a new prediction, not as an error.
- Requests and messages the client sends are validated strictly. An unrecognized field in a
configpart or asession.updateis rejected withconfig_invalid.

