Audio upload

Upload an audio file and receive measurements for every utterance in it.

The audio upload endpoint measures emotional expression in a recorded audio file. Send the file in one request, and the response lists every utterance the server detects, with its measurements. A realtime session produces the same measurements for the same audio.

EndpointPOST https://api.cloud.hume.ai/v1/expression/audio/file
ReferenceUpload audio
Content typemultipart/form-data
RunRecorded under the run_id. See Runs.

Request

The request body has one required part and one optional part. Authenticate with a credential header, as described in Authentication.

Required

The audio, in the format described in Audio format and size. Send exactly one file part.

Settings as a JSON object, sent with the content type application/json. See Configuration. Omit it to use the defaults.

The request below uploads a WAV file and sets the measurement interval to 5000 ms.

curl https://api.cloud.hume.ai/v1/expression/audio/file \
-H "X-Hume-Api-Key: $HUME_API_KEY" \
-F "[email protected];type=audio/wav" \
-F 'config={
"measurement_timer_ms": 5000
};type=application/json'

Audio format and size

The endpoint accepts one audio format, sent as a WAV file or as headerless samples, in a request of up to 25 MB.

PropertyValue
EncodingSigned 16-bit little-endian PCM (s16le)
Sample rate16,000 Hz
Channels1
ContainerA WAV file (audio/wav), or headerless samples (application/octet-stream)
Request size25 MB for the whole multipart body, a little over 13 minutes of audio
  1. MP3, FLAC, Ogg, MP4, and RIFF containers other than WAV are rejected with unsupported_media_type. Convert them before sending.
  2. Files without a recognized container are read as headerless samples. An empty file, or one with an odd number of bytes, is rejected with audio_decode_failed.
  3. Audio is not resampled or downmixed. A WAV file at another sample rate, channel count, or sample format is rejected with audio_decode_failed, as is a WAV file that is truncated or malformed, or that uses the extensible WAV header rather than plain PCM.
  4. A request over the size limit is rejected with audio_too_long.

To convert a recording in any format ffmpeg reads, run the command below.

Shell
ffmpeg -i recording.mp3 -acodec pcm_s16le -ac 1 -ar 16000 speech.wav

Utterances and measurements

An utterance is one continuous stretch of speech. The response lists every utterance in utterances with its span in the audio, and every measurement in measurements with the utterance_id it belongs to. All times are in milliseconds from the start of the audio.

  1. Measurements are taken at a fixed interval. While an utterance continues, the server measures it every measurement_timer_ms of speech, 3000 by default, with trigger set to interval. A measurement covers the most recent 10 seconds of the utterance, or the whole utterance if it is shorter, so consecutive measurements overlap. At the maximum interval of 10000 ms they are back to back. An interval measurement taken as speech stops can end shortly after the utterance does, and no final measurement follows it.
  2. Each utterance ends with a final measurement. When an utterance ends, the server produces a measurement with trigger set to speech_ended, unless less than one second of speech has passed since the previous measurement ended.
  3. Every utterance produces at least one measurement. An utterance shorter than the interval has no interval measurements, so its final measurement is always produced.
  4. The end of the audio ends the utterance. Audio that stops mid-speech closes the utterance in progress, and its final measurement follows the rule in item 2.

Because the final measurement can be skipped, the last measurement of an utterance can have a trigger of interval. Use utterances to find where each utterance ends, not trigger.

Response

The response below is for 5 seconds of audio containing one utterance.

200 OK
{
"request_id": "3f9a1c07e2b84d56a0c7e19b5d2f8a64",
"run_id": "01926f3c-d33c-7c9f-9c91-4502d80834d5",
"received": {
"audio_duration_ms": 5000
},
"utterances": [
{
"utterance_id": 0,
"audio_start_ms": 98,
"audio_end_ms": 4510
}
],
"measurements": [
{
"measurement_id": 0,
"utterance_id": 0,
"audio_start_ms": 98,
"audio_end_ms": 3098,
"trigger": "interval",
"expressions": [
{
"name": "interest",
"probability": 0.4286
},
{
"name": "curiosity",
"probability": 0.308
}
],
"voice_attributes": [
{
"name": "clear",
"probability": 0.9612
},
{
"name": "calm",
"probability": 0.9148
}
]
},
{
"measurement_id": 1,
"utterance_id": 0,
"audio_start_ms": 98,
"audio_end_ms": 4510,
"trigger": "speech_ended",
"expressions": [
{
"name": "amusement",
"probability": 0.5
},
{
"name": "joy",
"probability": 0.3593
}
],
"voice_attributes": [
{
"name": "cheerful",
"probability": 0.9275
},
{
"name": "bright",
"probability": 0.8864
}
]
}
]
}
request_id

Identifies the request in Hume’s records. Quote it when contacting support.

run_id

Identifies the run the request was recorded as. Pass it to the run endpoints to look the run up.

received.audio_duration_ms

The duration of the audio, whether or not it contained speech.

utterances

Every utterance, in order, each with its utterance_id and its span as audio_start_ms and audio_end_ms. utterance_id starts at 0.

measurements[].measurement_id

Sequential within the request, starting at 0.

measurements[].utterance_id

The utterance this measurement belongs to.

measurements[].audio_start_ms, measurements[].audio_end_ms

The window measured. The window never starts before the utterance does, but an interval measurement taken as speech stops can end shortly after it.

measurements[].trigger

interval when the measurement interval elapsed while speech continued. speech_ended when the utterance ended, or when the audio ended mid-speech.

measurements[].expressions

The 414 expression names are listed in Audio measurements.

measurements[].voice_attributes

How the voice and the recording sound. The 190 names are listed in Audio measurements.

Configuration

To change the measurement interval, send a config part. Fields you omit keep their defaults.

config
{
"measurement_timer_ms": 10000
}
Defaults to 3000

The interval between measurements, in milliseconds of speech, from 3000 to 10000.

Errors

A failed request returns an error body with a code, a message, and a request_id. A 401 carries only a message, and a request_id when an API key was rejected. HTTP errors describes the format.

StatusCodeDescription
400message_invalidThe multipart body could not be read, has no file part or more than one, or has a part with another name.
400config_invalidThe config part is not valid JSON, names an unrecognized field, holds a value outside its range, or appears more than once.
400audio_decode_failedThe audio could not be decoded: a WAV file that is truncated or malformed, is at another sample rate, channel count, or sample format, or uses the extensible WAV header, or a file that is empty or is not a whole number of 16-bit samples.
401NoneThe request carried no valid API key or access token.
403forbiddenYour organization is not permitted to use this endpoint. Contact support.
405method_not_allowedThe endpoint does not accept this HTTP method. The Allow header lists the methods it does accept.
413audio_too_longThe request body is over 25 MB.
415unsupported_media_typeThe audio is in a format the server recognizes but does not accept, such as MP3 or FLAC. Convert it before sending.
500internal_errorThe server failed while serving the request. Retry it. If the failure recurs, contact Hume with the request_id.
503server_busyThe service is too busy to serve the request now. Wait the number of seconds in the Retry-After header, then retry.

Next steps