Audio realtime

Stream audio and receive measurements for each utterance while it is still being spoken.

The audio realtime endpoint measures emotional expression in speech while you stream the audio. Send audio as binary WebSocket frames, and while each utterance continues, the server sends a measurement.result at a fixed interval. To measure a recording in one request, use Audio upload.

Endpointwss://api.cloud.hume.ai/v1/expression/audio/realtime
ReferenceStream audio
ProtocolSessions
RunRecorded under the session_id. See Runs.

Audio format

The endpoint accepts one audio format, sent as binary WebSocket frames.

PropertyValue
EncodingSigned 16-bit little-endian PCM (s16le)
Sample rate16,000 Hz
Channels1
ContainerNone. Raw samples, no header.

A frame must hold whole samples, so its length is an even number of bytes. It can hold up to 10 seconds of audio (320 KB). Any WebSocket message larger than 1 MiB (1,048,576 bytes), audio or text, ends the session with message_too_large. Frames of about 100 ms (3,200 bytes) work well. The server joins frames into one continuous stream, so a frame may begin or end anywhere in the audio.

Send rate

Stream audio in close to real time, no faster than 1 second of audio per second. Audio from a microphone arrives at this rate. To stream a recording, send one 100 ms frame every 100 ms. To measure a recording faster than it plays, send it to Audio upload instead.

A frame sent faster than the send rate can be rejected with rate_limited, and none of its audio is measured. Wait briefly and send the same frame again, so that the stream stays continuous.

Utterances

An utterance is one continuous stretch of speech. The server marks its start with utterance.start, sends one or more measurement.result messages while it continues, and marks its end with utterance.end. Silence between utterances produces no messages. A session that ends with an error can leave its last utterance without an utterance.end.

utterance.start
{
"type": "utterance.start",
"utterance_id": 0,
"audio_start_ms": 98
}
utterance.end
{
"type": "utterance.end",
"utterance_id": 0,
"audio_start_ms": 98,
"audio_end_ms": 26000
}

utterance_id starts at 0 and increases through the session. Every measurement.result names the utterance it belongs to.

Measurements

While an utterance is in progress, the server measures it every measurement_timer_ms of speech, 3000 by default. A measurement covers the most recent 10 seconds of the utterance, or the whole utterance if it is shorter, so consecutive measurements overlap. At the maximum interval of 10000 ms they are back to back. An interval measurement taken as speech stops can end shortly after the utterance does, and no final measurement follows it.

When an utterance ends, the server sends a final measurement covering up to its last 10 seconds, unless less than one second of speech has passed since the previous measurement ended. Every utterance produces at least one measurement.

measurement.result
{
"type": "measurement.result",
"measurement_id": 0,
"utterance_id": 0,
"audio_start_ms": 98,
"audio_end_ms": 3098,
"trigger": "interval",
"expressions": [
{
"name": "interest",
"probability": 0.6231
},
{
"name": "amusement",
"probability": 0.4187
},
{
"name": "joy",
"probability": 0.2054
}
],
"voice_attributes": [
{
"name": "clear",
"probability": 0.9612
},
{
"name": "calm",
"probability": 0.9148
}
]
}
measurement_id

Sequential within the session, starting at 0.

utterance_id

The utterance this measurement belongs to.

audio_start_ms, audio_end_ms

The window measured, in milliseconds from the start of the audio the server has accepted. The window never starts before the utterance does, but an interval measurement taken as speech stops can end shortly after it.

trigger

interval when the measurement interval elapsed while speech continued. speech_ended when the utterance ended, or when the session was closed mid-speech.

expressions

The 414 expression names are listed in Audio measurements.

voice_attributes

How the voice and the recording sound. The 190 names are listed in Audio measurements.

Because the final measurement can be skipped, the last measurement of an utterance can have a trigger of interval. Detect the end of an utterance with utterance.end, not with trigger.

Configuration

To change the measurement interval, send session.update before the first audio frame. The configuration locks at the first audio frame the server accepts; a rejected frame does not lock it. The rules for session.update, including that omitted fields return to their defaults, are in Sessions.

session.update
{
"type": "session.update",
"audio": {
"type": "audio/pcm",
"encoding": "s16le",
"sample_rate": 16000,
"channels": 1
},
"measurement_timer_ms": 10000
}
Defaults to 3000

The interval between measurements, in milliseconds of speech, from 3000 to 10000.

Defaults to the values shown

The audio format. Only the values shown are supported; anything else is rejected with config_invalid. May be omitted or sent as {}.

Message flow

A session that streams 26 seconds of continuous speech with the default configuration exchanges the messages below. Times are positions in the audio.

Time in audioClientServer
receivesession.created
sendsession.update (optional, before any audio)
receivesession.updated
0 ms onwardsendBinary audio frames
98 msreceiveutterance.start (utterance 0)
3098 msreceivemeasurement.result (98 to 3098 ms, interval)
6098 msreceivemeasurement.result (98 to 6098 ms, interval)
9098 to 24098 msreceivemeasurement.result every 3000 ms of speech (six more, each covering at most the most recent 10 seconds, interval)
26000 mssendsession.close
26000 msreceivemeasurement.result (16000 to 26000 ms, speech_ended)
26000 msreceiveutterance.end (utterance 0)
26000 msreceivesession.closed

Session summary

received.frames counts accepted frames, so subtracting it from the number you sent gives the number rejected. received.audio_duration_ms counts all accepted audio, whether or not it contained speech. produced.measurements counts measurement.result messages.

session.closed
{
"type": "session.closed",
"session_id": "01926f3a-5b1c-7d2e-8f40-3a9b7c1d2e5f",
"reason": "client_request",
"received": {
"audio_duration_ms": 26000,
"frames": 260
},
"produced": {
"utterances": 1,
"measurements": 9
}
}

Errors

The audio endpoint sends the shared codes, plus the code below. The internal_error row describes how the shared code behaves on this endpoint.

CodeSessionDescription
invalid_audio_frameContinuesThe frame was empty, ended mid-sample (an odd number of bytes), or held more than 10 seconds of audio. The frame is discarded.
internal_errorEndsProcessing the audio failed unexpectedly. Start a new session. If the failure recurs, contact Hume with the session_id.

Next steps