Audio
The audio endpoints measure emotional expression in speech. The server detects speech, groups it into utterances, and measures each utterance at a fixed interval. Every measurement has two lists of scores: expressions, such as amusement, and voice_attributes, such as monotone, which describe how the voice and the recording sound.
Upload or realtime
Both endpoints return the same measurements for the same audio. They differ in how the audio arrives and how the results come back.
Both endpoints accept one audio format: signed 16-bit little-endian PCM at 16,000 Hz, mono.
Supported languages
The model behind the audio endpoints measures expression in speech in more than 50 languages, including 16 core languages: Arabic, Bengali, English, French, German, Hebrew, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Russian, Spanish, Turkish, and Vietnamese. Requests have no language parameter, so the same request works across languages.

