Audio

Measure emotional expression in speech, from an audio file or a live stream.

The audio endpoints measure emotional expression in speech. The server detects speech, groups it into utterances, and measures each utterance at a fixed interval. Every measurement has two lists of scores: expressions, such as amusement, and voice_attributes, such as monotone, which describe how the voice and the recording sound.

Upload or realtime

Both endpoints return the same measurements for the same audio. They differ in how the audio arrives and how the results come back.

UploadRealtime
EndpointPOST /v1/expression/audio/fileWSS /v1/expression/audio/realtime
AudioA WAV file or headerless samples, up to 25 MBHeaderless samples, streamed as binary frames
ResultsEvery utterance and every measurement in one responseutterance.start, measurement.result, and utterance.end messages while speech continues
ConfigurationA config partsession.update, before the first audio frame
Recorded asA run under the run_idA run under the session_id
Use forRecorded media, for example to label a dataset or to evaluate generated audioMedia as it happens, for example to route or monitor a live call

Both endpoints accept one audio format: signed 16-bit little-endian PCM at 16,000 Hz, mono.

Supported languages

The model behind the audio endpoints measures expression in speech in more than 50 languages, including 16 core languages: Arabic, Bengali, English, French, German, Hebrew, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Russian, Spanish, Turkish, and Vietnamese. Requests have no language parameter, so the same request works across languages.

Next steps