Python quickstart

Install the Python SDK, upload audio and an image, and stream audio and video to the realtime endpoints.

Measure expressions in speech and faces from Python. Work through the parts in order or jump to the one you need.

Packagehume-expression-measurement on PyPI
Sourcehume-expression-measurement-python-sdk on GitHub
Python3.10 or later

Prerequisites

  1. Python 3.10 or later.
  2. ffmpeg, to convert the audio and read frames from the video.
  3. An API key from the API keys page of the Hume Platform. See Authentication.
  4. Speech audio shorter than 30 seconds, in any format ffmpeg reads, and a JPEG image of one or more faces.
  5. A microphone, to stream live audio.
  6. A short video of one or more faces, in any format ffmpeg reads.

Install

uv
uv add hume-expression-measurement

Set your API key

The client reads the key from the HUME_API_KEY environment variable and sends it in the X-Hume-Api-Key header of every request and connection.

Shell
export HUME_API_KEY=your_api_key

To pass the key explicitly instead, construct the client with ExpressionMeasurementClient(api_key="..."). Keep the key on a server. A browser cannot use it; see Browsers.

Prepare the audio

The audio endpoints accept one audio format: signed 16-bit little-endian PCM at 16,000 Hz, mono. Convert your audio with ffmpeg to raw samples with no header, which both the upload and realtime endpoints accept.

Shell
ffmpeg -i audio.wav -f s16le -acodec pcm_s16le -ac 1 -ar 16000 speech.pcm

Upload and stream files

  1. Upload audio

    Create audio_upload.py. It uploads the audio in one request and prints every measurement in the response.

    1. client.audio.measure() sends the file as a multipart/form-data request and returns the parsed response. The tuple gives the part a filename and a content type; application/octet-stream labels it as headerless samples. To upload a WAV file instead, open it and pass audio/wav.

    2. The response lists every utterance in utterances and every measurement in measurements, each with the utterance_id it belongs to. The request holds up to 25 MB, a little over 13 minutes of audio.

    3. A failed request raises an ApiError. status_code holds the HTTP status, and body holds the error body, whose code says what went wrong. See HTTP errors.

    4. To change the measurement interval, import AudioFileConfig from the package and pass one with measurement_timer_ms=5000 as config, alongside file. Audio upload covers the request, the response, and every error.

  2. Upload an image

    Create image_upload.py. It uploads one JPEG and prints each face the server measured in it.

    1. file is a list, so one request can carry several images. measurements holds one result per image, in the order sent, and frame_id is the image’s position in the list. bbox is the face’s location as [x0, y0, x1, y1] in pixels of the submitted image. Each face also carries descriptions, the facial actions visible on it. Image upload covers image limits, detection settings, and tracking faces across requests.

    2. The server lists the faces it detects but measures only the largest, 2 by default. A face it did not measure has face_id, expressions and descriptions set to None, so the script skips it. Detection explains which faces are measured.

  3. Stream audio

    Create audio_realtime.py. It opens a session, streams the audio in frames of 100 ms at the rate it plays, asks the server to close the session, and prints each measurement as it arrives. Streaming takes as long as the audio lasts; to measure a file sooner, upload it instead.

    1. AsyncExpressionMeasurementClient has the same methods as ExpressionMeasurementClient, and you await them. client.audio.connect() opens the WebSocket and yields a socket. Leaving the async with block closes the connection. The socket does not reconnect if the server restarts or the connection drops; Reconnecting shows how to connect with one that does.

    2. Measurements arrive while audio is still being sent, so print_measurements reads the socket in its own task while the script sends frames. Iterating the socket yields one message at a time, each parsed into the model for its type, such as AudioMeasurementResult for measurement.result. Check the model with isinstance, which also lets a type checker narrow the message to that model’s fields. The loop ends at session.closed.

    3. send_audio_frame sends one binary frame. At 16,000 Hz, 16-bit, mono, 3,200 bytes is 100 ms of audio, so sleeping 100 ms after each frame keeps the stream within the send rate of 1 second of audio per second. The server joins the frames into one stream, so a frame may begin or end anywhere in the audio.

    4. send_audio_session_close tells the server that no more audio is coming. It finishes measuring, sends session.closed, and closes the connection. The script then waits for print_measurements to reach session.closed.

    5. The server detects speech in the audio, groups it into utterances, and sends a measurement.result every 3 seconds of speech while an utterance continues, plus a final one when it ends. Each result also carries voice_attributes, which describe how the voice and the recording sound. Audio realtime covers every message.

  4. Stream images

    Create video_realtime.py. It sends the same JPEG as a single frame over the realtime endpoint and prints each face the server measured in it.

    1. Each measured frame produces one measurement.result listing the faces in it, and the script skips the faces that were not measured, as in the upload. bbox is the face’s location as [x0, y0, x1, y1] in pixels of the submitted image, and face_id identifies the same face from frame to frame when you send a series of images. Each face also carries descriptions, the facial actions visible on it. Video realtime covers frame limits, detection settings, and face tracking.

from hume_expression_measurement import ExpressionMeasurementClient
from hume_expression_measurement.core import ApiError
client = ExpressionMeasurementClient()
try:
with open("speech.pcm", "rb") as audio:
response = client.audio.measure(
file=("speech.pcm", audio, "application/octet-stream")
)
except ApiError as error:
print(f"{error.status_code}: {error.body}")
else:
for measurement in response.measurements:
expressions = ", ".join(
f"{score.name} {score.probability:.2f}"
for score in measurement.expressions[:3]
)
print(
f"{measurement.audio_start_ms} to "
f"{measurement.audio_end_ms} ms: {expressions}"
)
print(
f"{len(response.utterances)} utterances, "
f"{len(response.measurements)} measurements"
)

Stream from a microphone

The SDK’s audio helpers record from a microphone and convert the audio to the format the audio endpoints accept. Install them with the audio extra, which adds numpy, sounddevice, and soxr.

uv
uv add "hume-expression-measurement[audio]"

On Linux, sounddevice also needs the PortAudio library, for example sudo apt-get install libportaudio2 on Debian and Ubuntu.

Create microphone.py. It records from your default microphone for 10 seconds, sends the audio as it is recorded, and prints each measurement as it arrives.

microphone.py
import asyncio
from hume_expression_measurement import (
AsyncExpressionMeasurementClient,
Error,
SessionClose,
AudioMeasurementResult,
AudioSessionClosed,
)
from hume_expression_measurement.audio_helpers import Microphone
FRAME_COUNT = 100
async def main() -> None:
client = AsyncExpressionMeasurementClient()
async with (
Microphone.open() as microphone,
client.audio.connect() as socket,
):
async def print_measurements() -> None:
async for message in socket:
if isinstance(message, AudioMeasurementResult):
expressions = ", ".join(
f"{score.name} {score.probability:.2f}"
for score in message.expressions[:3]
)
print(
f"{message.audio_start_ms} to "
f"{message.audio_end_ms} ms: {expressions}"
)
elif isinstance(message, Error):
print(f"{message.code}: {message.message}")
elif isinstance(message, AudioSessionClosed):
produced = message.produced
print(
f"{produced.utterances} utterances, "
f"{produced.measurements} measurements"
)
break
receiver = asyncio.create_task(print_measurements())
try:
print(
f"Recording from {microphone.device_name} "
"for 10 seconds"
)
frames_sent = 0
async for frame in microphone:
await socket.send_audio_frame(frame)
frames_sent += 1
if frames_sent == FRAME_COUNT:
break
await socket.send_audio_session_close(
SessionClose(type="session.close")
)
await receiver
finally:
receiver.cancel()
asyncio.run(main())
  1. Microphone.open() records from the system’s default input device at the device’s own sample rate. async for yields its audio as 100 ms frames in the format the audio endpoints accept, so 100 frames is 10 seconds. To record from another device, pass device with its index or part of its name. On Windows, where each device is listed once per host API, such as MME and Windows WASAPI, add the host API to the name, for example "Microphone WASAPI", or pass the index. An unknown device, one that is not an input device, or a name that matches several devices raises a ValueError that lists the input devices.
  2. The microphone opens before the connection, so a missing or busy device fails before a session starts.
  3. AsyncExpressionMeasurementClient has the same methods as ExpressionMeasurementClient, and you await them. Measurements arrive while audio is still being sent, so print_measurements reads the socket in its own task while the loop sends frames.
  4. After the last frame, the script sends session.close and waits for session.closed. Audio from a microphone arrives in real time, within the send rate, so it needs no pacing.

On macOS, the first run asks you to allow the terminal or app running Python to use the microphone. If access is denied, the microphone records silence, and the session closes with 0 utterances.

hume_expression_measurement.audio_helpers also has iter_wav_frames, which reads an uncompressed WAV file and turns it into frames to send, and Resampler, which converts audio from other sources. Audio helpers in the SDK’s README covers both.

Stream a video file

The SDK’s video helpers read frames from a video file through ffmpeg and send them to the realtime endpoint within its send rate. They work with AsyncExpressionMeasurementClient and need no extra, only the ffmpeg you installed for the audio.

Create video_file.py. It measures video.mp4 at 3 frames per second of video and prints the top expressions of each measured face in every frame. Streaming takes as long as the video plays; to measure a recording sooner, send its frames to Image upload instead.

video_file.py
import asyncio
from hume_expression_measurement import AsyncExpressionMeasurementClient
from hume_expression_measurement.video_helpers import (
SessionEndedError,
iter_video_frames,
stream_video,
)
async def main() -> None:
client = AsyncExpressionMeasurementClient()
async with client.video.connect() as socket:
frames = iter_video_frames("video.mp4", realtime=True)
stream = stream_video(socket, frames)
try:
async for reply in stream:
if reply.error is not None:
print(
f"{reply.timestamp_ms:.0f} ms: "
f"{reply.error.code}"
)
continue
for face in reply.result.faces:
if face.expressions is None:
continue
expressions = ", ".join(
f"{score.name} {score.probability:.2f}"
for score in face.expressions[:3]
)
print(
f"{reply.timestamp_ms:.0f} ms, "
f"face {face.face_id}: {expressions}"
)
closed = stream.closed
if closed is not None:
print(
f"{closed.received.frames} frames, "
f"{closed.produced.measurements} faces measured"
)
except SessionEndedError as error:
print(error)
asyncio.run(main())
  1. iter_video_frames runs ffmpeg on the file and yields each frame as a JPEG image, with its position in the video as timestamp_ms. It takes 3 frames per second of video by default, the most the send rate allows; pass fps to take fewer. realtime=True yields the frames no faster than the video plays, as a live source would. Frames wider than 1280 pixels are scaled down. If ffmpeg is not installed, the loop raises an FfmpegNotFoundError that says how to install it.
  2. stream_video sends the frames as iter_video_frames yields them, so a minute of video takes about a minute to measure. If the server rejects a frame with rate_limited, it slows down and sends the frame again, so every frame is measured.
  3. The loop yields one reply per frame, in the order of the frames. Each reply holds the frame’s timestamp_ms and either the result that measured it or the error that rejected it, such as invalid_image_frame. A frame with no faces is a result with an empty faces list. A face that was not measured has expressions set to None, and the script skips it.
  4. When the frames run out, stream_video waits for the remaining replies and closes the session, the server closes the connection, and stream.closed holds the session.closed message with the session’s totals. If the session ends before the frames do, for example because the connection drops, the loop raises a SessionEndedError; open a new connection to continue.
  5. stream_video reads every message from the socket while it runs, so read messages from the replies instead. The configuration locks at the first frame, so send any session.update before streaming.

To measure a live source, such as a camera your application already captures, take at most 3 frames per second from it, encode each as a JPEG, and pass live=True. stream_video then drops a frame the server rejects with rate_limited instead of sending it again, so results keep up with the source. JpegSplitter splits a stream of concatenated JPEG images, such as MJPEG, into single frames. Video helpers in the SDK’s README covers both.

Run

uv
uv run audio_upload.py
uv run image_upload.py
uv run audio_realtime.py
uv run video_realtime.py
uv run microphone.py
uv run video_file.py

The upload and streaming scripts print the same measurements for the same files. Only scores that pass each array’s cutoff are returned, so a measurement may list fewer than three expressions, or none. Scores explains how to read them.

Reconnecting

The scripts above connect with connect(), which does not reconnect, so they stop when the connection ends. To keep measuring through a server restart or a dropped connection, connect with reconnecting from hume_expression_measurement.reconnect instead. When the connection ends unexpectedly, its socket connects again, which starts a new session. The socket sends your last session.update to the new session first, so the new session runs with your configuration.

Create microphone_reconnecting.py. It measures your default microphone until you press Ctrl+C, through any number of reconnects. Run it like the other scripts.

microphone_reconnecting.py
import asyncio
import contextlib
from hume_expression_measurement import (
AsyncExpressionMeasurementClient,
AudioMeasurementResult,
AudioSessionClosed,
AudioSessionCreated,
AudioSessionUpdate,
Error,
)
from hume_expression_measurement.audio_helpers import Microphone
from hume_expression_measurement.reconnect import reconnecting
from websockets.exceptions import ConnectionClosed
async def main() -> None:
client = AsyncExpressionMeasurementClient()
async with (
Microphone.open() as microphone,
reconnecting(client.audio) as socket,
):
await socket.send_audio_session_update(
AudioSessionUpdate(
type="session.update", measurement_timer_ms=5000
)
)
async def send_audio() -> None:
async for frame in microphone:
try:
await socket.send_audio_frame(frame)
except ConnectionClosed:
continue
sender = asyncio.create_task(send_audio())
try:
async for message in socket:
if isinstance(message, AudioSessionCreated):
print(f"session {message.session_id} created")
elif isinstance(message, AudioMeasurementResult):
expressions = ", ".join(
f"{score.name} {score.probability:.2f}"
for score in message.expressions[:3]
)
print(
f"{message.audio_start_ms} to "
f"{message.audio_end_ms} ms: {expressions}"
)
elif isinstance(message, Error):
print(f"{message.code}: {message.message}")
elif isinstance(message, AudioSessionClosed):
print(
f"session {message.session_id} closed: "
f"{message.reason}"
)
finally:
sender.cancel()
with contextlib.suppress(KeyboardInterrupt):
asyncio.run(main())
  1. The socket reconnects after a session.closed with reason server_shutdown, after a close with code 1001, 1011, 1012, or 1013, after a dropped connection, and after an audio session ends with an internal_error. Any other close is final, and so is a close after you send session.close or leave the block.
  2. Receiving is what notices a close and reconnects, so keep iterating the socket while the session runs, including past a session.closed, as the script does. Once the socket has ended for good, iteration ends or raises on its own.
  3. A failed first connection is not retried, since a wrong API key or URL would fail the same way again. Entering the block raises what connect() raises.
  4. A new session does not continue the previous one. You receive the previous session’s session.closed, if it sent one, then session.created with a new session ID. IDs and times in the new session’s results start again from 0, and results for media the previous session had not yet measured never arrive.
  5. While the socket reconnects, sending raises websockets.exceptions.ConnectionClosed and has no effect, except that session.close stops the reconnect. A live source can skip frames until sending succeeds again, as send_audio does. A recording loses whatever the previous session had not measured, so stop sending, leave the block, and send the recording again on a new connection or upload it. stream_video does not reconnect, and raises a TypeError when given a socket from reconnecting.
  6. With ExpressionMeasurementClient, enter reconnecting(client.audio) with with. Its socket accepts sends from other threads while one thread receives, so send from one thread and iterate in another.

Reconnecting in the SDK’s README covers retry timing and what receiving raises when the socket stops reconnecting.

Next steps