Audio upload
The audio upload endpoint measures emotional expression in a recorded audio file. Send the file in one request, and the response lists every utterance the server detects, with its measurements. A realtime session produces the same measurements for the same audio.
Request
The request body has one required part and one optional part. Authenticate with a credential header, as described in Authentication.
The audio, in the format described in Audio format and size. Send exactly one file part.
Settings as a JSON object, sent with the content type application/json. See Configuration. Omit it to use the defaults.
The request below uploads a WAV file and sets the measurement interval to 5000 ms.
Audio format and size
The endpoint accepts one audio format, sent as a WAV file or as headerless samples, in a request of up to 25 MB.
- MP3, FLAC, Ogg, MP4, and RIFF containers other than WAV are rejected with
unsupported_media_type. Convert them before sending. - Files without a recognized container are read as headerless samples. An empty file, or one with an odd number of bytes, is rejected with
audio_decode_failed. - Audio is not resampled or downmixed. A WAV file at another sample rate, channel count, or sample format is rejected with
audio_decode_failed, as is a WAV file that is truncated or malformed, or that uses the extensible WAV header rather than plain PCM. - A request over the size limit is rejected with
audio_too_long.
To convert a recording in any format ffmpeg reads, run the command below.
Utterances and measurements
An utterance is one continuous stretch of speech. The response lists every utterance in utterances with its span in the audio, and every measurement in measurements with the utterance_id it belongs to. All times are in milliseconds from the start of the audio.
- Measurements are taken at a fixed interval. While an utterance continues, the server measures it every
measurement_timer_msof speech, 3000 by default, withtriggerset tointerval. A measurement covers the most recent 10 seconds of the utterance, or the whole utterance if it is shorter, so consecutive measurements overlap. At the maximum interval of 10000 ms they are back to back. An interval measurement taken as speech stops can end shortly after the utterance does, and no final measurement follows it. - Each utterance ends with a final measurement. When an utterance ends, the server produces a measurement with
triggerset tospeech_ended, unless less than one second of speech has passed since the previous measurement ended. - Every utterance produces at least one measurement. An utterance shorter than the interval has no
intervalmeasurements, so its final measurement is always produced. - The end of the audio ends the utterance. Audio that stops mid-speech closes the utterance in progress, and its final measurement follows the rule in item 2.
Because the final measurement can be skipped, the last measurement of an utterance can have a trigger of interval. Use utterances to find where each utterance ends, not trigger.
Response
The response below is for 5 seconds of audio containing one utterance.
Identifies the request in Hume’s records. Quote it when contacting support.
Identifies the run the request was recorded as. Pass it to the run endpoints to look the run up.
The duration of the audio, whether or not it contained speech.
Every utterance, in order, each with its utterance_id and its span as audio_start_ms and audio_end_ms. utterance_id starts at 0.
Sequential within the request, starting at 0.
The utterance this measurement belongs to.
The window measured. The window never starts before the utterance does, but an interval measurement taken as speech stops can end shortly after it.
interval when the measurement interval elapsed while speech continued. speech_ended when the utterance ended, or when the audio ended mid-speech.
The 414 expression names are listed in Audio measurements.
How the voice and the recording sound. The 190 names are listed in Audio measurements.
Configuration
To change the measurement interval, send a config part. Fields you omit keep their defaults.
The interval between measurements, in milliseconds of speech, from 3000 to 10000.
Errors
A failed request returns an error body with a code, a message, and a request_id. A 401 carries only a message, and a request_id when an API key was rejected. HTTP errors describes the format.

