Voice Conversion (Streamed JSON)

Authentication

X-Hume-Api-Keystring
API Key authentication via header

Request

This endpoint expects a multipart form containing an optional file.
strip_headersbooleanOptional

If enabled, the audio for all the chunks of a generation, once concatenated together, will constitute a single audio file. Otherwise, if disabled, each chunk’s audio will be its own audio file, each with its own headers (if applicable).

audiofileOptional

Audio file containing speech to be converted to the target voice. Supported formats include MP3, WAV, M4A, and OGG.

contextobject or nullOptional
Utterances to use as context for generating consistent speech style and prosody across multiple requests. These will not be converted to speech output.
ContextGenerationIdobjectRequired
OR
ContextUtterancesobjectRequired
voiceobjectOptional
VoiceIdobjectRequired
OR
VoiceNameobjectRequired
formatobjectOptionalDefaults to {"type":"mp3"}
Specifies the output audio file format.
mp3objectRequired
OR
pcmobjectRequired
OR
wavobjectRequired
include_timestamp_typeslist of enumsOptionalDefaults to []

The set of timestamp types to include in the response. When used in multipart/form-data, specify each value using bracket notation: include_timestamp_types[0]=word&include_timestamp_types[1]=phoneme. Only supported for Octave 2 requests.

Allowed values:

Response

Successful Response
audioobject
Metadata for a chunk of generated audio.
type"audio"
audiostring
The generated audio output chunk in the requested format.
audio_formatenum
The generated audio output format.
Allowed values:
chunk_indexinteger
The index of the audio chunk in the snippet.
generation_idstringformat: "uuid4"
The generation ID of the parent snippet that this chunk corresponds to.
is_last_chunkboolean
Whether or not this is the last chunk streamed back from the decoder for one input snippet.
request_idstring
ID of the initiating request.
snippet_idstringformat: "uuid4"
The ID of the parent snippet that this chunk corresponds to.
textstring
The text of the parent snippet that this chunk corresponds to.
transcribed_textstring or null

The transcribed text of the generated audio of the parent snippet that this chunk corresponds to. It is only present if instant_mode is set to false.

utterance_indexinteger or null
The index of the utterance in the request that the parent snippet of this chunk corresponds to.
snippetobjectOptional
OR
timestampobject
A word or phoneme level timestamp for the generated audio.
type"timestamp"
generation_idstringformat: "uuid4"
The generation ID of the parent snippet that this chunk corresponds to.
request_idstring
ID of the initiating request.
snippet_idstringformat: "uuid4"
The ID of the parent snippet that this chunk corresponds to.
timestampobject
A word or phoneme level timestamp for the generated audio.

Errors

422
Unprocessable Entity Error