Python quickstart
Measure expressions in speech and faces from Python. Work through the parts in order or jump to the one you need.
Prerequisites
- Python 3.10 or later.
- ffmpeg, to convert the audio and read frames from the video.
- An API key from the API keys page of the Hume Platform. See Authentication.
- Speech audio shorter than 30 seconds, in any format ffmpeg reads, and a JPEG image of one or more faces.
- A microphone, to stream live audio.
- A short video of one or more faces, in any format ffmpeg reads.
Install
uv
pip
Set your API key
The client reads the key from the HUME_API_KEY environment variable and sends it in the X-Hume-Api-Key header of every request and connection.
To pass the key explicitly instead, construct the client with ExpressionMeasurementClient(api_key="..."). Keep the key on a server. A browser cannot use it; see Browsers.
Prepare the audio
The audio endpoints accept one audio format: signed 16-bit little-endian PCM at 16,000 Hz, mono. Convert your audio with ffmpeg to raw samples with no header, which both the upload and realtime endpoints accept.
Upload and stream files
Upload audio
Create
audio_upload.py. It uploads the audio in one request and prints every measurement in the response.client.audio.measure()sends the file as amultipart/form-datarequest and returns the parsed response. The tuple gives the part a filename and a content type;application/octet-streamlabels it as headerless samples. To upload a WAV file instead, open it and passaudio/wav.The response lists every utterance in
utterancesand every measurement inmeasurements, each with theutterance_idit belongs to. The request holds up to 25 MB, a little over 13 minutes of audio.A failed request raises an
ApiError.status_codeholds the HTTP status, andbodyholds the error body, whosecodesays what went wrong. See HTTP errors.To change the measurement interval, import
AudioFileConfigfrom the package and pass one withmeasurement_timer_ms=5000asconfig, alongsidefile. Audio upload covers the request, the response, and every error.
Upload an image
Create
image_upload.py. It uploads one JPEG and prints each face the server measured in it.fileis a list, so one request can carry several images.measurementsholds one result per image, in the order sent, andframe_idis the image’s position in the list.bboxis the face’s location as[x0, y0, x1, y1]in pixels of the submitted image. Each face also carriesdescriptions, the facial actions visible on it. Image upload covers image limits, detection settings, and tracking faces across requests.The server lists the faces it detects but measures only the largest, 2 by default. A face it did not measure has
face_id,expressionsanddescriptionsset toNone, so the script skips it. Detection explains which faces are measured.
Stream audio
Create
audio_realtime.py. It opens a session, streams the audio in frames of 100 ms at the rate it plays, asks the server to close the session, and prints each measurement as it arrives. Streaming takes as long as the audio lasts; to measure a file sooner, upload it instead.AsyncExpressionMeasurementClienthas the same methods asExpressionMeasurementClient, and you await them.client.audio.connect()opens the WebSocket and yields a socket. Leaving theasync withblock closes the connection. The socket does not reconnect if the server restarts or the connection drops; Reconnecting shows how to connect with one that does.Measurements arrive while audio is still being sent, so
print_measurementsreads the socket in its own task while the script sends frames. Iterating the socket yields one message at a time, each parsed into the model for itstype, such asAudioMeasurementResultformeasurement.result. Check the model withisinstance, which also lets a type checker narrow the message to that model’s fields. The loop ends atsession.closed.send_audio_framesends one binary frame. At 16,000 Hz, 16-bit, mono, 3,200 bytes is 100 ms of audio, so sleeping 100 ms after each frame keeps the stream within the send rate of 1 second of audio per second. The server joins the frames into one stream, so a frame may begin or end anywhere in the audio.send_audio_session_closetells the server that no more audio is coming. It finishes measuring, sendssession.closed, and closes the connection. The script then waits forprint_measurementsto reachsession.closed.The server detects speech in the audio, groups it into utterances, and sends a
measurement.resultevery 3 seconds of speech while an utterance continues, plus a final one when it ends. Each result also carriesvoice_attributes, which describe how the voice and the recording sound. Audio realtime covers every message.
Stream images
Create
video_realtime.py. It sends the same JPEG as a single frame over the realtime endpoint and prints each face the server measured in it.Each measured frame produces one
measurement.resultlisting the faces in it, and the script skips the faces that were not measured, as in the upload.bboxis the face’s location as[x0, y0, x1, y1]in pixels of the submitted image, andface_ididentifies the same face from frame to frame when you send a series of images. Each face also carriesdescriptions, the facial actions visible on it. Video realtime covers frame limits, detection settings, and face tracking.
Stream from a microphone
The SDK’s audio helpers record from a microphone and convert the audio to the format the audio endpoints accept. Install them with the audio extra, which adds numpy, sounddevice, and soxr.
uv
pip
On Linux, sounddevice also needs the PortAudio library, for example sudo apt-get install libportaudio2 on Debian and Ubuntu.
Create microphone.py. It records from your default microphone for 10 seconds, sends the audio as it is recorded, and prints each measurement as it arrives.
Microphone.open()records from the system’s default input device at the device’s own sample rate.async foryields its audio as 100 ms frames in the format the audio endpoints accept, so 100 frames is 10 seconds. To record from another device, passdevicewith its index or part of its name. On Windows, where each device is listed once per host API, such as MME and Windows WASAPI, add the host API to the name, for example"Microphone WASAPI", or pass the index. An unknown device, one that is not an input device, or a name that matches several devices raises aValueErrorthat lists the input devices.- The microphone opens before the connection, so a missing or busy device fails before a session starts.
AsyncExpressionMeasurementClienthas the same methods asExpressionMeasurementClient, and you await them. Measurements arrive while audio is still being sent, soprint_measurementsreads the socket in its own task while the loop sends frames.- After the last frame, the script sends
session.closeand waits forsession.closed. Audio from a microphone arrives in real time, within the send rate, so it needs no pacing.
On macOS, the first run asks you to allow the terminal or app running Python to use the microphone. If access is denied, the microphone records silence, and the session closes with 0 utterances.
hume_expression_measurement.audio_helpers also has iter_wav_frames, which reads an uncompressed WAV file and turns it into frames to send, and Resampler, which converts audio from other sources. Audio helpers in the SDK’s README covers both.
Stream a video file
The SDK’s video helpers read frames from a video file through ffmpeg and send them to the realtime endpoint within its send rate. They work with AsyncExpressionMeasurementClient and need no extra, only the ffmpeg you installed for the audio.
Create video_file.py. It measures video.mp4 at 3 frames per second of video and prints the top expressions of each measured face in every frame. Streaming takes as long as the video plays; to measure a recording sooner, send its frames to Image upload instead.
iter_video_framesruns ffmpeg on the file and yields each frame as a JPEG image, with its position in the video astimestamp_ms. It takes 3 frames per second of video by default, the most the send rate allows; passfpsto take fewer.realtime=Trueyields the frames no faster than the video plays, as a live source would. Frames wider than 1280 pixels are scaled down. If ffmpeg is not installed, the loop raises anFfmpegNotFoundErrorthat says how to install it.stream_videosends the frames asiter_video_framesyields them, so a minute of video takes about a minute to measure. If the server rejects a frame withrate_limited, it slows down and sends the frame again, so every frame is measured.- The loop yields one reply per frame, in the order of the frames. Each reply holds the frame’s
timestamp_msand either theresultthat measured it or theerrorthat rejected it, such asinvalid_image_frame. A frame with no faces is a result with an emptyfaceslist. A face that was not measured hasexpressionsset toNone, and the script skips it. - When the frames run out,
stream_videowaits for the remaining replies and closes the session, the server closes the connection, andstream.closedholds thesession.closedmessage with the session’s totals. If the session ends before the frames do, for example because the connection drops, the loop raises aSessionEndedError; open a new connection to continue. stream_videoreads every message from the socket while it runs, so read messages from the replies instead. The configuration locks at the first frame, so send anysession.updatebefore streaming.
To measure a live source, such as a camera your application already captures, take at most 3 frames per second from it, encode each as a JPEG, and pass live=True. stream_video then drops a frame the server rejects with rate_limited instead of sending it again, so results keep up with the source. JpegSplitter splits a stream of concatenated JPEG images, such as MJPEG, into single frames. Video helpers in the SDK’s README covers both.
Run
uv
pip
The upload and streaming scripts print the same measurements for the same files. Only scores that pass each array’s cutoff are returned, so a measurement may list fewer than three expressions, or none. Scores explains how to read them.
Reconnecting
The scripts above connect with connect(), which does not reconnect, so they stop when the connection ends. To keep measuring through a server restart or a dropped connection, connect with reconnecting from hume_expression_measurement.reconnect instead. When the connection ends unexpectedly, its socket connects again, which starts a new session. The socket sends your last session.update to the new session first, so the new session runs with your configuration.
Create microphone_reconnecting.py. It measures your default microphone until you press Ctrl+C, through any number of reconnects. Run it like the other scripts.
- The socket reconnects after a
session.closedwith reasonserver_shutdown, after a close with code 1001, 1011, 1012, or 1013, after a dropped connection, and after an audio session ends with aninternal_error. Any other close is final, and so is a close after you sendsession.closeor leave the block. - Receiving is what notices a close and reconnects, so keep iterating the socket while the session runs, including past a
session.closed, as the script does. Once the socket has ended for good, iteration ends or raises on its own. - A failed first connection is not retried, since a wrong API key or URL would fail the same way again. Entering the block raises what
connect()raises. - A new session does not continue the previous one. You receive the previous session’s
session.closed, if it sent one, thensession.createdwith a new session ID. IDs and times in the new session’s results start again from 0, and results for media the previous session had not yet measured never arrive. - While the socket reconnects, sending raises
websockets.exceptions.ConnectionClosedand has no effect, except thatsession.closestops the reconnect. A live source can skip frames until sending succeeds again, assend_audiodoes. A recording loses whatever the previous session had not measured, so stop sending, leave the block, and send the recording again on a new connection or upload it.stream_videodoes not reconnect, and raises aTypeErrorwhen given a socket fromreconnecting. - With
ExpressionMeasurementClient, enterreconnecting(client.audio)withwith. Its socket accepts sends from other threads while one thread receives, so send from one thread and iterate in another.
Reconnecting in the SDK’s README covers retry timing and what receiving raises when the socket stops reconnecting.

