Node.js quickstart
Measure expressions in speech and faces from Node.js. Every script below runs with node and nothing else, so work through the parts in order or jump to the one you need.
The package ships TypeScript types for every message. The scripts contain no type annotations, so on Node.js earlier than 22.18 they also run as .js files.
Prerequisites
- Node.js 22.18 or later, which runs TypeScript directly, or 18 or later to run the scripts as JavaScript.
- ffmpeg, to convert the audio and read frames from the video.
- An API key from the API keys page of the Hume Platform. See Authentication.
- Speech audio shorter than 30 seconds, in any format ffmpeg reads, and a JPEG image of one or more faces.
- A short video of one or more faces, in any format ffmpeg reads.
Install
bun
pnpm
npm
yarn
The scripts below are ES modules, so set "type": "module" in the project’s package.json.
To type-check the TypeScript scripts, also install the Node.js type definitions.
Set your API key
The client reads the key from the HUME_API_KEY environment variable and sends it in the X-Hume-Api-Key header of every request and connection.
To pass the key explicitly instead, construct the client with new ExpressionMeasurementClient({ apiKey: "..." }). This package is for servers. A browser cannot use the key; see Browsers.
Prepare the audio
The audio endpoints accept one audio format: signed 16-bit little-endian PCM at 16,000 Hz, mono. Convert your audio with ffmpeg to raw samples with no header, which both the upload and realtime endpoints accept.
Upload and stream files
Upload audio
Create
audio-upload.ts. It uploads the audio in one request and prints every measurement in the response.client.audio.measure()sends the file as amultipart/form-datarequest and resolves to the parsed response.contentTypelabels the part as headerless samples; to upload a WAV file instead, pass its path withaudio/wav.The response lists every utterance in
utterancesand every measurement inmeasurements, each with theutteranceIdit belongs to. The request holds up to 25 MB, a little over 13 minutes of audio.A failed request throws an
ExpressionMeasurementError. Itsmessageincludes the HTTP status and the error body, whosecodesays what went wrong. To branch on them, readstatusCodeandbody. See HTTP errors.To change the measurement interval, pass
configalongsidefile, set to{ measurementTimerMs: 5000 }. Audio upload covers the request, the response, and every error.
Upload an image
Create
image-upload.ts. It uploads one JPEG and prints each face the server measured in it.fileis an array, so one request can carry several images.measurementsholds one result per image, in the order sent, andframeIdis the image’s position in the array.bboxis the face’s location as[x0, y0, x1, y1]in pixels of the submitted image. Each face also carriesdescriptions, the facial actions visible on it. Image upload covers image limits, detection settings, and tracking faces across requests.The server lists the faces it detects but measures only the largest, 2 by default. A face it did not measure has
faceId,expressionsanddescriptionsset tonull, so the script skips it. Detection explains which faces are measured.
Stream audio
Create
audio-realtime.ts. It opens a session, registers a message handler, streams the audio in frames of 100 ms at the rate it plays, and asks the server to close the session. Streaming takes as long as the audio lasts; to measure a file sooner, upload it instead.client.audio.connect()returns a socket before the handshake completes, so the script registers its handler and then waits forwaitForOpen()before sending. If the server rejects the handshake, for example because the key is invalid,waitForOpen()throws. If the server restarts or the connection drops before the script sendssession.close, the socket reconnects and starts a new session; see Reconnecting.The
messagehandler receives one typed message at a time. Every message has atype, so aswitchonmessage.typeis enough to handle them, and in TypeScript it narrows the message to the matching type.sendAudioFramesends one binary frame. At 16,000 Hz, 16-bit, mono, 3,200 bytes is 100 ms of audio, so waiting 100 ms after each frame keeps the stream within the send rate of 1 second of audio per second. The server joins the frames into one stream, so a frame may begin or end anywhere in the audio.sendAudioSessionClosetells the server that no more audio is coming. It finishes measuring, sendssession.closed, and closes the connection. The socket does not reconnect aftersession.close, so the script exits.The server detects speech in the audio, groups it into utterances, and sends a
measurement.resultevery 3 seconds of speech while an utterance continues, plus a final one when it ends. Each result also carriesvoiceAttributes, which describe how the voice and the recording sound. Audio realtime covers every message.
Stream images
Create
video-realtime.ts. It sends the same JPEG as a single frame over the realtime endpoint and prints each face the server measured in it.Each measured frame produces one
measurement.resultlisting the faces in it, and the script skips the faces that were not measured, as in the upload.bboxis the face’s location as[x0, y0, x1, y1]in pixels of the submitted image, andfaceIdidentifies the same face from frame to frame when you send a series of images. Each face also carriesdescriptions, the facial actions visible on it. Video realtime covers frame limits, detection settings, and face tracking.
Stream from a microphone
The SDK’s audio helpers record from a microphone and convert the audio to the format the audio endpoints accept. They record through decibri, an optional peer dependency, so install it alongside the SDK.
bun
pnpm
npm
yarn
decibri provides native addons for Windows and Linux (glibc) on x64 and arm64, and for macOS on arm64. On Linux it also needs the ALSA library: sudo apt-get install libasound2t64 on Ubuntu 24.04, Debian 13 and later, or sudo apt-get install libasound2 on earlier releases of Debian and Ubuntu.
Create microphone.ts. It records from your default microphone for 10 seconds, sends the audio as it is recorded, and prints each measurement as it arrives.
Microphone.open()records from the system’s default input device at the device’s own sample rate.for awaityields its audio as 100 ms frames in the format the audio endpoints accept, so 100 frames is 10 seconds. To record from another device, passdevicewith its index or part of its name. An unknown device, or a name that matches several devices, throws an error that lists the input devices.- The microphone opens before the connection, so a missing or busy device fails before a session starts. The
finallyblock closes it, which releases the device. - If the server restarts or the connection drops, the socket reconnects and starts a new session. While it reconnects it is not open, and sending throws, so the loop skips frames until
readyStateisReadyState.OPENagain and still stops after 10 seconds of recording. See Reconnecting. - After the last frame, the script sends
session.close, and the server sendssession.closedand closes the connection. If the socket is still reconnecting, the script callsclose()instead, which stops the reconnect. Audio from a microphone arrives in real time, within the send rate, so it needs no pacing.
On macOS, the first run asks you to allow the terminal or app running Node to use the microphone. If access is denied, the script fails or records silence.
@humeai/expression-measurement-node/audio-helpers also has iterWavFrames, which reads an uncompressed WAV file and turns it into frames to send, and Resampler, which converts audio from other sources. Audio helpers in the SDK’s README covers both.
Stream a video file
The SDK’s video helpers read frames from a video file through ffmpeg and send them to the realtime endpoint within its send rate. They need nothing beyond the ffmpeg you installed for the audio.
Create video-file.ts. It measures video.mp4 at 3 frames per second of video and prints the top expressions of each measured face in every frame. Streaming takes as long as the video plays; to measure a recording sooner, send its frames to Image upload instead.
iterVideoFramesruns ffmpeg on the file and yields each frame as a JPEG image, with its position in the video astimestampMs. It takes 3 frames per second of video by default, the most the send rate allows; pass{ fps }to take fewer.{ realtime: true }yields the frames no faster than the video plays, as a live source would. Frames wider than 1280 pixels are scaled down. If ffmpeg is not installed, the loop throws anFfmpegNotFoundErrorthat says how to install it.streamVideosends the frames asiterVideoFramesyields them, so a minute of video takes about a minute to measure. If the server rejects a frame withrate_limited, it slows down and sends the frame again, so every frame is measured.- The loop yields one reply per frame, in the order of the frames. Each reply holds the frame’s
timestampMsand either theresultthat measured it or theerrorthat rejected it, such asinvalid_image_frame. A frame with no faces is a result with an emptyfacesarray. A face that was not measured hasexpressionsset tonull, and the script skips it. - When the frames run out,
streamVideocloses the session and the socket, andstream.closedholds thesession.closedmessage with the session’s totals. If the session ends before the frames do, for example because the server restarts or the connection drops, the loop throws aSessionEndedError.streamVideocloses the socket instead of reconnecting, so open a new connection to continue. streamVideotakes over the socket’smessagehandler, so read messages from the replies instead. The configuration locks at the first frame, so send anysession.updatebefore streaming.
To measure a live source, such as a camera your application already captures, take at most 3 frames per second from it, encode each as a JPEG, and pass { live: true }. streamVideo then drops a frame the server rejects with rate_limited instead of sending it again, so results keep up with the source. JpegSplitter splits a stream of concatenated JPEG images, such as MJPEG, into single frames. Video helpers in the SDK’s README covers both.
Run
The upload and streaming scripts print the same measurements for the same files. Only scores that pass each array’s cutoff are returned, so a measurement may list fewer than three expressions, or none. Scores explains how to read them.
Reconnecting
When a realtime connection ends unexpectedly, the socket connects again, which starts a new session. The socket sends your latest session.update to each new session first, so the new session runs with your configuration.
- The socket reconnects after a
session.closedwith reasonserver_shutdown, after a close with code 1001, 1011, 1012, or 1013, after a dropped connection, and after an audio session ends with aninternal_error. Any other close is final, and so is a close after you sendsession.closeor callclose(). - A failed first connection is not retried, since a wrong API key or URL would fail the same way again.
waitForOpen()throws, and the socket closes. - A new session does not continue the previous one. It has a new session ID, and IDs and times in its results start again from 0. Results for media the previous session had not yet measured never arrive.
- While the socket reconnects it is not open, and sending throws. A live source should skip frames until
socket.readyStateisReadyState.OPEN, as the microphone script does. A recording loses whatever the previous session had not measured, so stop sending, callclose(), and send the recording again on a new connection or upload it.streamVideostops on its own: it closes the socket instead of reconnecting and throws aSessionEndedError. - To turn reconnecting off, pass
reconnectAttempts: 0toconnect(). To choose which closes reconnect, passshouldReconnect, which receives the close event. A close aftersession.closeis final either way.
Reconnecting in the SDK’s README covers retry timing, the socket’s events while it reconnects, and both options.

