Node.js quickstart

Install the Node.js SDK, upload audio and an image, and stream audio and video to the realtime endpoints.

Measure expressions in speech and faces from Node.js. Every script below runs with node and nothing else, so work through the parts in order or jump to the one you need.

Package@humeai/expression-measurement-node on npm
Sourcehume-expression-measurement-node-sdk on GitHub
Node.js18 or later, or 22.18 or later for TypeScript

The package ships TypeScript types for every message. The scripts contain no type annotations, so on Node.js earlier than 22.18 they also run as .js files.

Prerequisites

  1. Node.js 22.18 or later, which runs TypeScript directly, or 18 or later to run the scripts as JavaScript.
  2. ffmpeg, to convert the audio and read frames from the video.
  3. An API key from the API keys page of the Hume Platform. See Authentication.
  4. Speech audio shorter than 30 seconds, in any format ffmpeg reads, and a JPEG image of one or more faces.
  5. A short video of one or more faces, in any format ffmpeg reads.

Install

bun
bun add @humeai/expression-measurement-node

The scripts below are ES modules, so set "type": "module" in the project’s package.json.

Shell
npm pkg set type=module

To type-check the TypeScript scripts, also install the Node.js type definitions.

Shell
npm install --save-dev @types/node

Set your API key

The client reads the key from the HUME_API_KEY environment variable and sends it in the X-Hume-Api-Key header of every request and connection.

Shell
export HUME_API_KEY=your_api_key

To pass the key explicitly instead, construct the client with new ExpressionMeasurementClient({ apiKey: "..." }). This package is for servers. A browser cannot use the key; see Browsers.

Prepare the audio

The audio endpoints accept one audio format: signed 16-bit little-endian PCM at 16,000 Hz, mono. Convert your audio with ffmpeg to raw samples with no header, which both the upload and realtime endpoints accept.

Shell
ffmpeg -i audio.wav -f s16le -acodec pcm_s16le -ac 1 -ar 16000 speech.pcm

Upload and stream files

  1. Upload audio

    Create audio-upload.ts. It uploads the audio in one request and prints every measurement in the response.

    1. client.audio.measure() sends the file as a multipart/form-data request and resolves to the parsed response. contentType labels the part as headerless samples; to upload a WAV file instead, pass its path with audio/wav.

    2. The response lists every utterance in utterances and every measurement in measurements, each with the utteranceId it belongs to. The request holds up to 25 MB, a little over 13 minutes of audio.

    3. A failed request throws an ExpressionMeasurementError. Its message includes the HTTP status and the error body, whose code says what went wrong. To branch on them, read statusCode and body. See HTTP errors.

    4. To change the measurement interval, pass config alongside file, set to { measurementTimerMs: 5000 }. Audio upload covers the request, the response, and every error.

  2. Upload an image

    Create image-upload.ts. It uploads one JPEG and prints each face the server measured in it.

    1. file is an array, so one request can carry several images. measurements holds one result per image, in the order sent, and frameId is the image’s position in the array. bbox is the face’s location as [x0, y0, x1, y1] in pixels of the submitted image. Each face also carries descriptions, the facial actions visible on it. Image upload covers image limits, detection settings, and tracking faces across requests.

    2. The server lists the faces it detects but measures only the largest, 2 by default. A face it did not measure has faceId, expressions and descriptions set to null, so the script skips it. Detection explains which faces are measured.

  3. Stream audio

    Create audio-realtime.ts. It opens a session, registers a message handler, streams the audio in frames of 100 ms at the rate it plays, and asks the server to close the session. Streaming takes as long as the audio lasts; to measure a file sooner, upload it instead.

    1. client.audio.connect() returns a socket before the handshake completes, so the script registers its handler and then waits for waitForOpen() before sending. If the server rejects the handshake, for example because the key is invalid, waitForOpen() throws. If the server restarts or the connection drops before the script sends session.close, the socket reconnects and starts a new session; see Reconnecting.

    2. The message handler receives one typed message at a time. Every message has a type, so a switch on message.type is enough to handle them, and in TypeScript it narrows the message to the matching type.

    3. sendAudioFrame sends one binary frame. At 16,000 Hz, 16-bit, mono, 3,200 bytes is 100 ms of audio, so waiting 100 ms after each frame keeps the stream within the send rate of 1 second of audio per second. The server joins the frames into one stream, so a frame may begin or end anywhere in the audio.

    4. sendAudioSessionClose tells the server that no more audio is coming. It finishes measuring, sends session.closed, and closes the connection. The socket does not reconnect after session.close, so the script exits.

    5. The server detects speech in the audio, groups it into utterances, and sends a measurement.result every 3 seconds of speech while an utterance continues, plus a final one when it ends. Each result also carries voiceAttributes, which describe how the voice and the recording sound. Audio realtime covers every message.

  4. Stream images

    Create video-realtime.ts. It sends the same JPEG as a single frame over the realtime endpoint and prints each face the server measured in it.

    1. Each measured frame produces one measurement.result listing the faces in it, and the script skips the faces that were not measured, as in the upload. bbox is the face’s location as [x0, y0, x1, y1] in pixels of the submitted image, and faceId identifies the same face from frame to frame when you send a series of images. Each face also carries descriptions, the facial actions visible on it. Video realtime covers frame limits, detection settings, and face tracking.

import {
ExpressionMeasurementClient,
ExpressionMeasurementError,
} from "@humeai/expression-measurement-node";
const client = new ExpressionMeasurementClient();
try {
const response = await client.audio.measure({
file: {
path: "speech.pcm",
contentType: "application/octet-stream",
},
});
for (const measurement of response.measurements) {
const expressions = measurement.expressions
.slice(0, 3)
.map((score) => `${score.name} ${score.probability.toFixed(2)}`)
.join(", ");
console.log(
`${measurement.audioStartMs} to ` +
`${measurement.audioEndMs} ms: ${expressions}`,
);
}
console.log(
`${response.utterances.length} utterances, ` +
`${response.measurements.length} measurements`,
);
} catch (error) {
if (!(error instanceof ExpressionMeasurementError)) throw error;
console.error(error.message);
}

Stream from a microphone

The SDK’s audio helpers record from a microphone and convert the audio to the format the audio endpoints accept. They record through decibri, an optional peer dependency, so install it alongside the SDK.

bun
bun add decibri

decibri provides native addons for Windows and Linux (glibc) on x64 and arm64, and for macOS on arm64. On Linux it also needs the ALSA library: sudo apt-get install libasound2t64 on Ubuntu 24.04, Debian 13 and later, or sudo apt-get install libasound2 on earlier releases of Debian and Ubuntu.

Create microphone.ts. It records from your default microphone for 10 seconds, sends the audio as it is recorded, and prints each measurement as it arrives.

microphone.ts
import {
ExpressionMeasurementClient,
ReadyState,
} from "@humeai/expression-measurement-node";
import {
Microphone,
} from "@humeai/expression-measurement-node/audio-helpers";
const FRAME_COUNT = 100;
const client = new ExpressionMeasurementClient();
const microphone = await Microphone.open();
const socket = await client.audio.connect();
socket.on("message", (message) => {
switch (message.type) {
case "measurement.result": {
const expressions = message.expressions
.slice(0, 3)
.map(
(score) => `${score.name} ${score.probability.toFixed(2)}`,
)
.join(", ");
console.log(
`${message.audioStartMs} to ` +
`${message.audioEndMs} ms: ${expressions}`,
);
break;
}
case "error":
console.error(`${message.code}: ${message.message}`);
break;
case "session.closed":
console.log(
`${message.produced.utterances} utterances, ` +
`${message.produced.measurements} measurements`,
);
break;
}
});
await socket.waitForOpen();
console.log(`Recording from ${microphone.deviceName} for 10 seconds`);
try {
let framesRecorded = 0;
for await (const frame of microphone) {
if (socket.readyState === ReadyState.OPEN) {
socket.sendAudioFrame(frame);
}
framesRecorded++;
if (framesRecorded === FRAME_COUNT) {
break;
}
}
} finally {
await microphone.close();
}
if (socket.readyState === ReadyState.OPEN) {
socket.sendAudioSessionClose({ type: "session.close" });
} else {
socket.close();
}
  1. Microphone.open() records from the system’s default input device at the device’s own sample rate. for await yields its audio as 100 ms frames in the format the audio endpoints accept, so 100 frames is 10 seconds. To record from another device, pass device with its index or part of its name. An unknown device, or a name that matches several devices, throws an error that lists the input devices.
  2. The microphone opens before the connection, so a missing or busy device fails before a session starts. The finally block closes it, which releases the device.
  3. If the server restarts or the connection drops, the socket reconnects and starts a new session. While it reconnects it is not open, and sending throws, so the loop skips frames until readyState is ReadyState.OPEN again and still stops after 10 seconds of recording. See Reconnecting.
  4. After the last frame, the script sends session.close, and the server sends session.closed and closes the connection. If the socket is still reconnecting, the script calls close() instead, which stops the reconnect. Audio from a microphone arrives in real time, within the send rate, so it needs no pacing.

On macOS, the first run asks you to allow the terminal or app running Node to use the microphone. If access is denied, the script fails or records silence.

@humeai/expression-measurement-node/audio-helpers also has iterWavFrames, which reads an uncompressed WAV file and turns it into frames to send, and Resampler, which converts audio from other sources. Audio helpers in the SDK’s README covers both.

Stream a video file

The SDK’s video helpers read frames from a video file through ffmpeg and send them to the realtime endpoint within its send rate. They need nothing beyond the ffmpeg you installed for the audio.

Create video-file.ts. It measures video.mp4 at 3 frames per second of video and prints the top expressions of each measured face in every frame. Streaming takes as long as the video plays; to measure a recording sooner, send its frames to Image upload instead.

video-file.ts
import {
ExpressionMeasurementClient,
} from "@humeai/expression-measurement-node";
import {
iterVideoFrames,
SessionEndedError,
streamVideo,
} from "@humeai/expression-measurement-node/video-helpers";
const client = new ExpressionMeasurementClient();
const socket = await client.video.connect();
const frames = iterVideoFrames("video.mp4", { realtime: true });
const stream = streamVideo(socket, frames);
try {
for await (const reply of stream) {
if (reply.error) {
console.error(`${reply.timestampMs} ms: ${reply.error.code}`);
continue;
}
for (const face of reply.result.faces) {
if (!face.expressions) continue;
const expressions = face.expressions
.slice(0, 3)
.map((score) => `${score.name} ${score.probability.toFixed(2)}`)
.join(", ");
console.log(
`${reply.timestampMs} ms, face ${face.faceId}: ${expressions}`,
);
}
}
console.log(
`${stream.closed?.received.frames} frames, ` +
`${stream.closed?.produced.measurements} faces measured`,
);
} catch (error) {
if (!(error instanceof SessionEndedError)) throw error;
console.error(error.message);
}
  1. iterVideoFrames runs ffmpeg on the file and yields each frame as a JPEG image, with its position in the video as timestampMs. It takes 3 frames per second of video by default, the most the send rate allows; pass { fps } to take fewer. { realtime: true } yields the frames no faster than the video plays, as a live source would. Frames wider than 1280 pixels are scaled down. If ffmpeg is not installed, the loop throws an FfmpegNotFoundError that says how to install it.
  2. streamVideo sends the frames as iterVideoFrames yields them, so a minute of video takes about a minute to measure. If the server rejects a frame with rate_limited, it slows down and sends the frame again, so every frame is measured.
  3. The loop yields one reply per frame, in the order of the frames. Each reply holds the frame’s timestampMs and either the result that measured it or the error that rejected it, such as invalid_image_frame. A frame with no faces is a result with an empty faces array. A face that was not measured has expressions set to null, and the script skips it.
  4. When the frames run out, streamVideo closes the session and the socket, and stream.closed holds the session.closed message with the session’s totals. If the session ends before the frames do, for example because the server restarts or the connection drops, the loop throws a SessionEndedError. streamVideo closes the socket instead of reconnecting, so open a new connection to continue.
  5. streamVideo takes over the socket’s message handler, so read messages from the replies instead. The configuration locks at the first frame, so send any session.update before streaming.

To measure a live source, such as a camera your application already captures, take at most 3 frames per second from it, encode each as a JPEG, and pass { live: true }. streamVideo then drops a frame the server rejects with rate_limited instead of sending it again, so results keep up with the source. JpegSplitter splits a stream of concatenated JPEG images, such as MJPEG, into single frames. Video helpers in the SDK’s README covers both.

Run

Shell
node audio-upload.ts
node image-upload.ts
node audio-realtime.ts
node video-realtime.ts
node microphone.ts
node video-file.ts

The upload and streaming scripts print the same measurements for the same files. Only scores that pass each array’s cutoff are returned, so a measurement may list fewer than three expressions, or none. Scores explains how to read them.

Reconnecting

When a realtime connection ends unexpectedly, the socket connects again, which starts a new session. The socket sends your latest session.update to each new session first, so the new session runs with your configuration.

  1. The socket reconnects after a session.closed with reason server_shutdown, after a close with code 1001, 1011, 1012, or 1013, after a dropped connection, and after an audio session ends with an internal_error. Any other close is final, and so is a close after you send session.close or call close().
  2. A failed first connection is not retried, since a wrong API key or URL would fail the same way again. waitForOpen() throws, and the socket closes.
  3. A new session does not continue the previous one. It has a new session ID, and IDs and times in its results start again from 0. Results for media the previous session had not yet measured never arrive.
  4. While the socket reconnects it is not open, and sending throws. A live source should skip frames until socket.readyState is ReadyState.OPEN, as the microphone script does. A recording loses whatever the previous session had not measured, so stop sending, call close(), and send the recording again on a new connection or upload it. streamVideo stops on its own: it closes the socket instead of reconnecting and throws a SessionEndedError.
  5. To turn reconnecting off, pass reconnectAttempts: 0 to connect(). To choose which closes reconnect, pass shouldReconnect, which receives the close event. A close after session.close is final either way.

Reconnecting in the SDK’s README covers retry timing, the socket’s events while it reconnects, and both options.

Next steps