Stream audio

Stream audio from a microphone or a file and receive expression measurements as speech happens. You do not need to detect speech or split it into segments: the server detects each stretch of speech and measures it while it is still being spoken. ## Sending audio Send audio as binary messages of signed 16-bit little-endian PCM, mono, 16 kHz, with no header. Frames of around 100 ms work well. You can cut the audio at any sample boundary; the server joins the frames into one continuous stream. - **Configuration:** send `session.update` before your first audio frame. A `session.update` after your first audio frame ends the session with `config_invalid`. - **Rejected frames:** a frame that is empty, has an odd number of bytes, or holds more than 10 seconds of audio gets `invalid_audio_frame`, and the session continues. A message over 1MiB ends the session with `message_too_large`. Accepted frames are not acknowledged. - **Send rate:** stream audio in close to real time, no faster than 1 second of audio per second. A frame sent faster can get `rate_limited`, and none of its audio is measured. Wait briefly and resend the same frame, so that the stream stays continuous. To measure a recording faster than it plays, use the audio file endpoint. - **Failures:** if measurement fails, you receive `internal_error` and the session ends. Open a new connection. ## How measurements are produced The server detects when speech starts and stops. Each continuous stretch of speech is an **utterance**, bracketed by `utterance.start` and `utterance.end`. - **During speech:** a `measurement.result` arrives every `measurement_timer_ms` milliseconds of speech, with `trigger` set to `interval`. - **When speech ends:** a final measurement arrives with `trigger` set to `speech_ended`, unless the previous measurement ended less than one second earlier. Every utterance gets at least one measurement. - **Span:** each measurement covers the most recent 10 seconds of the utterance, or the whole utterance if it is shorter, so consecutive measurements can overlap. `audio_start_ms` and `audio_end_ms` give the span, counted from the start of the audio the server has accepted. - **Ending the session:** `session.close` during speech ends the utterance. You receive its final measurement, if one is due, then `utterance.end`, then `session.closed`. A session that ends with an error can leave its last utterance without an `utterance.end`. ## Limits | Limit | Value | |---|---| | Audio per frame | 10 seconds (320KB) | | Message size | 1MiB (1,048,576 bytes) | | Send rate | 1 second of audio per second | ## Message flow A session at the default `measurement_timer_ms` of 3000, with one utterance: speech begins 98 ms into the audio and is still in progress at 26000 ms, when `session.close` ends it. | Time in audio | Direction | Message | |---|---|---| | on connect | receive | `session.created` | | before audio | send | `session.update` (if you change the configuration) | | before audio | receive | `session.updated` (if you sent `session.update`) | | 0 ms onward | send | audio frames | | 98 ms | receive | `utterance.start` (utterance 0) | | 3098 ms | receive | `measurement.result` (98 to 3098 ms, `interval`) | | 6098 ms | receive | `measurement.result` (98 to 6098 ms, `interval`) | | 9098 to 24098 ms | receive | `measurement.result` every 3000 ms of speech (six more, each covering at most the most recent 10 seconds) | | 26000 ms | send | `session.close` | | 26000 ms | receive | `measurement.result` (16000 to 26000 ms, `speech_ended`) | | 26000 ms | receive | `utterance.end` (utterance 0) | | 26000 ms | receive | `session.closed` (`produced.utterances` 1, `produced.measurements` 9) |