If enabled, the audio for all the chunks of a generation, once concatenated together, will constitute a single audio file. Otherwise, if disabled, each chunk’s audio will be its own audio file, each with its own headers (if applicable).
Selects the Octave model version used to synthesize speech for this request. If you omit this field, Hume automatically routes the request to the most appropriate model. Setting a specific version ensures stable and repeatable behavior across requests.
Use 2 to opt into the latest Octave capabilities. When you specify version 2, you must also provide a
voice. Requests that set version: 2 without a voice will be rejected.
For a comparison of Octave versions, see the Octave versions section in the TTS overview.
API key used for authenticating the client. If not provided, an access_token must be provided to authenticate.
For more details, refer to the Authentication Strategies Guide.
Natural language instructions describing how the synthesized speech should sound, including but not limited to tone, intonation, pacing, and accent.
This field behaves differently depending on whether a voice is specified:
Duration of trailing silence (in seconds) to add to this utterance
The name or id associated with a Voice from the Voice Library to be used as the speaker for this and all subsequent utterances, until the voice field is updated again.
See our voices guide for more details on generating and specifying Voices.
Access token used for authenticating the client. If not provided, an api_key must be provided to authenticate.
The access token is generated using both an API key and a Secret key, which provides an additional layer of security compared to using just an API key.
For more details, refer to the Authentication Strategies Guide.
Enables ultra-low latency streaming, significantly reducing the time until the first audio chunk is received. Recommended for real-time applications requiring immediate audio playback. For further details, see our documentation on instant mode.
Sampling temperature for the speech generation model. Higher values increase variation; lower values increase consistency.
This is an experimental parameter. It is recommended to use the default values for most use cases.
Defaults when omitted:
0.90.80.75