> This page is for Expression Measurement.

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://dev.hume.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://dev.hume.ai/_mcp/server.

# Expression Measurement API

> **Info**
>
> The API is currently in limited release; [contact our team](https://www.hume.ai/sales-form) to request access.

**The Expression Measurement API measures emotional expression in speech and in faces**, with dedicated audio and video endpoints. The audio endpoints detect utterances in audio, and the video endpoints detect faces in images and video frames. Every score is a probability from 0 to 1, and [Scores](/expression-measurement/docs/scores) explains what it denotes in each array.

#### [Python quickstart](/expression-measurement/docs/quickstart/python)

Label datasets and evaluate audio from a script, notebook, or backend service.

#### [Node.js quickstart](/expression-measurement/docs/quickstart/nodejs)

Add expression measurement to a Node.js service or command-line tool.

## What it measures

From speech, the API returns expressions such as `amusement` and vocal qualities such as `monotone`. From an image, it returns expressions and visible facial actions, such as `smile`, for the largest faces it detects, 2 by default.

<table>
  <thead>
    <tr>
      <th width="12%">
        Modality
      </th>

      <th>
        Input
      </th>

      <th>
        Output
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        [Audio](/expression-measurement/docs/audio/measurements)
      </td>

      <td>
        Speech audio
      </td>

      <td>
        For each window of speech, scores for the 

        `expressions`

         and 

        `voice_attributes`

         that pass each array's cutoff, out of 414 and 190 names.
      </td>
    </tr>

    <tr>
      <td>
        [Video](/expression-measurement/docs/video/measurements)
      </td>

      <td>
        Images or video frames containing faces
      </td>

      <td>
        For each face, its bounding box. For the largest faces, 2 by default, scores for the 

        `expressions`

         and 

        `descriptions`

         that pass each array's cutoff, out of 48 and 27 names.
      </td>
    </tr>
  </tbody>
</table>

Learn more about the science behind expression measurement in Hume's [research](https://www.hume.ai/research) and [publications](https://www.hume.ai/publications).

## Endpoints

**The endpoints fall into three groups.** Upload endpoints measure media you already have, realtime endpoints measure media as it streams, and run endpoints return the record of each request and session.

### Upload

The upload endpoints take audio or images you already have and return every measurement in one response. Use them for recorded media, for example to label a dataset or to evaluate generated audio.

<table>
  <thead>
    <tr>
      <th width="20%">
        Endpoint
      </th>

      <th>
        URL
      </th>

      <th>
        Guide
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        Audio upload
      </td>

      <td>
        `POST https://api.cloud.hume.ai/v1/expression/audio/file`
      </td>

      <td>
        [Audio upload](/expression-measurement/docs/audio/upload)
      </td>
    </tr>

    <tr>
      <td>
        Image upload
      </td>

      <td>
        `POST https://api.cloud.hume.ai/v1/expression/video/file`
      </td>

      <td>
        [Image upload](/expression-measurement/docs/video/upload)
      </td>
    </tr>
  </tbody>
</table>

#### Send the request

Send a `multipart/form-data` request with the media in `file` parts and, to change the defaults, settings in a JSON `config` part. Put your API key in the `X-Hume-Api-Key` header. See [Authentication](/expression-measurement/docs/authentication).

#### Read the response

The response lists the measurements: one for each window of speech in the audio, or one for each image with the faces found in it. [Scores](/expression-measurement/docs/scores) explains how to read the scores. The `request_id` identifies the request, and the `run_id` identifies the run it was recorded as.

### Realtime

The realtime endpoints measure media streamed over a WebSocket and send each measurement as soon as it is produced. Stream audio to the audio endpoint from a source such as a microphone, or stream video to the video endpoint one JPEG frame at a time, such as from a camera. Use them to act on media as it happens, for example to route or monitor a live call.

<table>
  <thead>
    <tr>
      <th width="20%">
        Endpoint
      </th>

      <th>
        URL
      </th>

      <th>
        Guide
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        Audio realtime
      </td>

      <td>
        `wss://api.cloud.hume.ai/v1/expression/audio/realtime`
      </td>

      <td>
        [Audio realtime](/expression-measurement/docs/audio/realtime)
      </td>
    </tr>

    <tr>
      <td>
        Video realtime
      </td>

      <td>
        `wss://api.cloud.hume.ai/v1/expression/video/realtime`
      </td>

      <td>
        [Video realtime](/expression-measurement/docs/video/realtime)
      </td>
    </tr>
  </tbody>
</table>

Both endpoints share one protocol, described in [Sessions](/expression-measurement/docs/sessions).

#### Connect

Open a WebSocket to the audio or video realtime endpoint with your API key in the `X-Hume-Api-Key` header. The server immediately sends `session.created`, which contains the session ID and the default configuration.

#### Configure, if needed

To change a setting, such as the measurement interval or the face detection threshold, send `session.update` before the first media frame. The configuration locks when the server takes the first frame. Skip this step to use the defaults.

#### Stream media

Send audio or JPEG images as binary WebSocket frames. The server detects speech in the audio, or faces in each image, and measures them.

#### Read results

Each `measurement.result` contains the scores for one window of speech or for the faces in one image.

#### Close

Send `session.close`, then keep reading until `session.closed` arrives. It states why the session ended and totals what the server received and produced.

### Runs

Each measured upload request and each realtime session is recorded as a run. The run endpoints return a run's status, configuration, and event log. Measurements come only from the upload response or the session, so store the results you need.

#### [Runs guide](/expression-measurement/docs/runs)

Look up the status, configuration, and event log of past requests and sessions.

#### [API reference](/expression-measurement/reference/runs/list)

Every parameter and field of the run endpoints.

## SDKs

**Hume publishes SDKs for Python and Node.js.** Both cover every endpoint and include audio helpers that record from a microphone, read WAV files, and convert audio to the format the audio endpoints accept.

<table>
  <thead>
    <tr>
      <th width="15%">
        SDK
      </th>

      <th>
        Package
      </th>

      <th>
        Source
      </th>

      <th>
        Guide
      </th>
    </tr>
  </thead>

  <tbody>
    <tr>
      <td>
        Python
      </td>

      <td>
        [`hume-expression-measurement`](https://pypi.org/project/hume-expression-measurement/)

         on PyPI
      </td>

      <td>
        [GitHub](https://github.com/HumeAI/hume-expression-measurement-python-sdk)
      </td>

      <td>
        [Python quickstart](/expression-measurement/docs/quickstart/python)
      </td>
    </tr>

    <tr>
      <td>
        Node.js
      </td>

      <td>
        [`@humeai/expression-measurement-node`](https://www.npmjs.com/package/@humeai/expression-measurement-node)

         on npm
      </td>

      <td>
        [GitHub](https://github.com/HumeAI/hume-expression-measurement-node-sdk)
      </td>

      <td>
        [Node.js quickstart](/expression-measurement/docs/quickstart/nodejs)
      </td>
    </tr>
  </tbody>
</table>

## Compatibility

The API version is the path prefix, `/v1/`. Within a version:

1. **The server may add fields to the responses and messages it sends.** Ignore fields you do not recognize.
2. **New prediction names may be added.** Treat a name you do not recognize as a new prediction, not as an error.
3. **Requests and messages the client sends are validated strictly.** An unrecognized field in a `config` part or a `session.update` is rejected with `config_invalid`.

## Next steps

#### [Audio](/expression-measurement/docs/audio)

Upload or stream audio and measure expression in speech.

#### [Video](/expression-measurement/docs/video)

Upload images or stream video and measure expression in faces.

#### [Authentication](/expression-measurement/docs/authentication)

API keys, access tokens, and calling the API from a browser.

#### [Scores](/expression-measurement/docs/scores)

How to read probabilities.