Upload audio

Send the audio as the `file` part and, to change the defaults, settings as a JSON `config` part. The response holds every utterance found in the audio and every measurement taken, the same results a realtime session would stream for the same audio. ## How measurements are produced The server detects when speech starts and stops. Each continuous stretch of speech is an **utterance**, listed in `utterances` with its span in the audio. - **During speech:** a measurement is taken every `measurement_timer_ms` milliseconds of speech, with `trigger` set to `interval`. - **When speech ends:** a final measurement is taken with `trigger` set to `speech_ended`, unless the previous measurement ended less than one second earlier. Every utterance gets at least one measurement, and audio that ends during speech ends the utterance. - **Span:** each measurement covers the most recent 10 seconds of the utterance, or the whole utterance if it is shorter, so consecutive measurements can overlap. `audio_start_ms` and `audio_end_ms` give the span, counted from the start of the audio. ## Limits | Limit | Value | |---|---| | Request size | 25MB, the whole multipart body | | Audio format | 16-bit PCM, 16 kHz, mono, as a WAV file or headerless samples | The server does not convert audio, so send it in the format above. - **Too large:** a request over the size limit is rejected with `audio_too_long`. - **Unsupported formats:** MP3, FLAC, Ogg, MP4 or M4A, and RIFF files other than WAV, are rejected with `unsupported_media_type`. Convert them before sending. - **Wrong WAV format:** a WAV file at another sample rate, channel count or sample format, one that uses the extensible WAV header, or one that is truncated or malformed is rejected with `audio_decode_failed`. - **Anything else** is read as headerless samples rather than rejected, so check the format before sending. An empty file, or one with an odd number of bytes, is rejected with `audio_decode_failed`.

Authentication

X-Hume-Api-Keystring
API Key authentication via header

Request

This endpoint expects a multipart form containing a file.
configobjectOptional

Settings for an audio request, sent as the JSON config part.

measurement_timer_msintegerOptional3000-10000Defaults to 3000
How often to measure within an utterance, in milliseconds of speech.
filefileRequired

The audio, as 16-bit PCM at 16 kHz, mono: either a WAV file or the headerless little-endian samples. A WAV file at another sample rate, channel count or sample format is rejected rather than converted.

Response

Every utterance and measurement in the audio.
request_idstring
Identifies this request in Hume's records. Quote it when contacting support.
run_idstringformat: "uuid"
Identifies the run this request was recorded as. Pass it to the run endpoints to look the run up.
receivedobject
What the server took from the request.
audio_duration_msinteger>=0
The duration of the audio, in milliseconds.
utteranceslist of objects
Every utterance found, in order.
utterance_idinteger>=0
Identifies the utterance within the run, counting from 0.
audio_start_msinteger>=0
Where the utterance started, in milliseconds from the start of the audio.
audio_end_msinteger>=0
Where the utterance ended, in milliseconds from the start of the audio.
measurementslist of objects

Every measurement, in order. utterance_id says which utterance each belongs to.

measurement_idinteger>=0
Sequential identifier of the measurement within the run, starting at 0.
utterance_idinteger>=0
The utterance this measurement belongs to.
audio_start_msinteger>=0
Start of the window measured, in milliseconds from the start of the audio. Never earlier than the start of the utterance.
audio_end_msinteger>=0
End of the window measured, in milliseconds from the start of the audio.
triggerenum

Why the measurement was produced.

  • interval: the measurement interval (measurement_timer_ms) elapsed while speech continued.
  • speech_ended: the utterance ended, or the audio ended while speech was in progress.

A final measurement is skipped when the previous one ended less than one second earlier, so the last measurement of an utterance may carry the trigger interval. Where an utterance ended is reported separately: as an utterance.end message in a realtime session, and in utterances in an audio file response.

Allowed values:
expressionslist of objects

Expressions detected in the speech, such as joy, anger and chuckling, in descending order of probability. Scores with equal probability are ordered by the model’s underlying score, then by name. Only expressions judged present are included, each against a threshold set for that expression, so the array may be empty. probability is the measured chance that the expression is truly present.

voice_attributeslist of objects

Qualities of the voice and the audio, such as husky, fast and background music, in descending order of probability. Scores with equal probability are ordered by the model’s underlying score, then by name. Only scores of 0.725 or above are included, so the array may be empty. probability is calibrated: the expected intensity a human rater would give the quality.

Errors

400
Bad Request Error
401
Unauthorized Error
403
Forbidden Error
405
Method Not Allowed Error
413
Content Too Large Error
415
Unsupported Media Type Error
500
Internal Server Error
503
Service Unavailable Error