Chat
Chat with Empathic Voice Interface (EVI)
Chat with Empathic Voice Interface (EVI)
Allows external connections to this chat via the /connect endpoint.
The maximum number of chat events to return from chat history. By default, the system returns up to 300 events (100 events per page × 3 pages). Set this parameter to a smaller value to limit the number of events returned.
A flag to enable verbose transcription. Set this query parameter to "true" to have unfinalized user transcripts be sent to the client as interim UserMessage messages.
Configuration details for the audio input used during the session. Ensures the audio is being correctly set up for processing.
This optional field is only required when the audio input is encoded in PCM Linear 16 (16-bit, little-endian, signed PCM WAV data). For detailed instructions on how to configure session settings for PCM Linear 16 audio, please refer to the Session Settings section on the EVI Configuration page.
Field for injecting additional context into the conversation, which is appended to the end of user messages for the session.
When included in a Session Settings message, the provided context can be used to remind the LLM of its role in every user message, prevent it from forgetting important details, or add new relevant information to the conversation.
Set to null to clear injected context.
Unique identifier for the session. Used to manage conversational state, correlate frontend and backend data, and persist conversations across EVI sessions.
If included, the response sent from Hume to your backend will include this ID. This allows you to correlate frontend users with their incoming messages.
It is recommended to pass a custom_session_id if you are using a Custom Language Model. Please see our guide to using a custom language model with EVI to learn more.
The maximum number of chat events to return from chat history. By default, the system returns up to 300 events (100 events per page × 3 pages). Set this parameter to a smaller value to limit the number of events returned.
Instructions used to shape EVI’s behavior, responses, and style for the session.
When included in a Session Settings message, the provided Prompt overrides the existing one specified in the EVI configuration. If no Prompt was defined in the configuration, this Prompt will be the one used for the session.
You can use the Prompt to define a specific goal or role for EVI, specifying how it should act or what it should focus on during the conversation. For example, EVI can be instructed to act as a customer support representative, a fitness coach, or a travel advisor, each with its own set of behaviors and response styles.
For help writing a system prompt, see our Prompting Guide.
This field allows you to assign values to dynamic variables referenced in your system prompt.
Each key represents the variable name, and the corresponding value is the specific content you wish to assign to that variable within the session. While the values for variables can be strings, numbers, or booleans, the value will ultimately be converted to a string when injected into your system prompt.
When used in query parameters, specify each variable using bracket notation: session_settings[variables][key]=value. For example: session_settings[variables][name]=John&session_settings[variables][age]=30.
Using this field, you can personalize responses based on session-specific details. For more guidance, see our guide on using dynamic variables.
Base64 encoded audio input to insert into the conversation.
The content of an Audio Input message is treated as the user’s speech to EVI and must be streamed continuously. Pre-recorded audio files are not supported.
For optimal transcription quality, the audio data should be transmitted in small chunks.
Hume recommends streaming audio with a buffer window of 20 milliseconds (ms), or 100 milliseconds (ms) for web applications.
The type of message sent through the socket; must be audio_input for our server to correctly identify and process it as an Audio Input message.
This message is used for sending audio input data to EVI for processing and expression measurement. Audio data should be sent as a continuous stream, encoded in Base64.
The type of message sent through the socket; must be session_settings for our server to correctly identify and process it as a Session Settings message.
Session settings are temporary and apply only to the current Chat session. These settings can be adjusted dynamically based on the requirements of each session to ensure optimal performance and user experience.
For more information, please refer to the Session Settings section on the EVI Configuration page.
Configuration details for the audio input used during the session. Ensures the audio is being correctly set up for processing.
This optional field is only required when the audio input is encoded in PCM Linear 16 (16-bit, little-endian, signed PCM WAV data). For detailed instructions on how to configure session settings for PCM Linear 16 audio, please refer to the Session Settings section on the EVI Configuration page.
List of built-in tools to enable for the session.
Tools are resources used by EVI to perform various tasks, such as searching the web or calling external APIs. Built-in tools, like web search, are natively integrated, while user-defined tools are created and invoked by the user. To learn more, see our Tool Use Guide.
Currently, the only built-in tool Hume provides is Web Search. When enabled, Web Search equips EVI with the ability to search the web for up-to-date information.
Field for injecting additional context into the conversation, which is appended to the end of user messages for the session.
When included in a Session Settings message, the provided context can be used to remind the LLM of its role in every user message, prevent it from forgetting important details, or add new relevant information to the conversation.
Set to null to clear injected context.
Unique identifier for the session. Used to manage conversational state, correlate frontend and backend data, and persist conversations across EVI sessions.
If included, the response sent from Hume to your backend will include this ID. This allows you to correlate frontend users with their incoming messages.
It is recommended to pass a custom_session_id if you are using a Custom Language Model. Please see our guide to using a custom language model with EVI to learn more.
Instructions used to shape EVI’s behavior, responses, and style for the session.
When included in a Session Settings message, the provided Prompt overrides the existing one specified in the EVI configuration. If no Prompt was defined in the configuration, this Prompt will be the one used for the session.
You can use the Prompt to define a specific goal or role for EVI, specifying how it should act or what it should focus on during the conversation. For example, EVI can be instructed to act as a customer support representative, a fitness coach, or a travel advisor, each with its own set of behaviors and response styles.
For help writing a system prompt, see our Prompting Guide.
List of user-defined tools to enable for the session.
Tools are resources used by EVI to perform various tasks, such as searching the web or calling external APIs. Built-in tools, like web search, are natively integrated, while user-defined tools are created and invoked by the user. To learn more, see our Tool Use Guide.
This field allows you to assign values to dynamic variables referenced in your system prompt.
Each key represents the variable name, and the corresponding value is the specific content you wish to assign to that variable within the session. While the values for variables can be strings, numbers, or booleans, the value will ultimately be converted to a string when injected into your system prompt.
Using this field, you can personalize responses based on session-specific details. For more guidance, see our guide on using dynamic variables.
The type of message sent through the socket; must be user_input for our server to correctly identify and process it as a User Input message.
Assistant text to synthesize into spoken audio and insert into the conversation.
EVI uses this text to generate spoken audio using our proprietary expressive text-to-speech model. Our model adds appropriate emotional inflections and tones to the text based on the user’s expressions and the context of the conversation. The synthesized audio is streamed back to the user as an Assistant Message.
The type of message sent through the socket; must be assistant_input for our server to correctly identify and process it as an Assistant Input message.
The unique identifier for a specific tool call instance.
This ID is used to track the request and response of a particular tool invocation, ensuring that the correct response is linked to the appropriate request. The specified tool_call_id must match the one received in the Tool Call message.
The type of message sent through the socket; for a Tool Response message, this must be tool_response.
Upon receiving a Tool Call message and successfully invoking the function, this message is sent to convey the result of the function call back to EVI.
The unique identifier for a specific tool call instance.
This ID is used to track the request and response of a particular tool invocation, ensuring that the Tool Error message is linked to the appropriate tool call request. The specified tool_call_id must match the one received in the Tool Call message.
The type of message sent through the socket; for a Tool Error message, this must be tool_error.
Upon receiving a Tool Call message and failing to invoke the function, this message is sent to notify EVI of the tool’s failure.
Indicates the severity of an error; for a Tool Error message, this must be warn to signal an unexpected event.
Type of tool called. Either builtin for natively implemented tools, like web search, or function for user-defined tools.
The type of message sent through the socket; must be pause_assistant_message for our server to correctly identify and process it as a Pause Assistant message.
Once this message is sent, EVI will not respond until a Resume Assistant message is sent. When paused, EVI won’t respond, but transcriptions of your audio inputs will still be recorded.
The type of message sent through the socket; must be resume_assistant_message for our server to correctly identify and process it as a Resume Assistant message.
Upon resuming, if any audio input was sent during the pause, EVI will retain context from all messages sent but only respond to the last user message. (e.g., If you ask EVI two questions while paused and then send a resume_assistant_message, EVI will respond to the second question and have added the first question to its conversation context.)
Indicates the conclusion of the assistant’s response, signaling that the assistant has finished speaking for the current conversational turn.
Transcript of the assistant’s message. Contains the message role, content, and optionally tool call information including the tool name, parameters, response requirement status, tool call ID, and tool type.
Expression measurement predictions of the assistant’s audio output. Contains inference model results including prosody scores for 48 emotions within the detected expression of the assistant’s audio sample.
Base64 encoded audio output. This encoded audio is transmitted to the client, where it can be decoded and played back as part of the user interaction. The returned audio format is WAV and the sample rate is 48kHz.
Contains the audio data, an ID to track and reference the audio output, and an index indicating the chunk position relative to the whole audio segment. See our Audio Guide for more details on preparing and processing audio.
The first message received after establishing a connection with EVI, containing important identifiers for the current Chat session.
Includes the Chat ID (which allows the Chat session to be tracked and referenced) and the Chat Group ID (used to resume a Chat when passed in the resumed_chat_group_id query parameter of a subsequent connection request, allowing EVI to continue the conversation from where it left off within the Chat Group).
Indicates a disruption in the WebSocket connection, such as an unexpected disconnection, protocol error, or data transmission issue.
Contains an error code identifying the type of error encountered, a detailed description of the error, and a short, human-readable identifier and description (slug) for the error.
Indicates the user has interrupted the assistant’s response. EVI detects the interruption in real-time and sends this message to signal the interruption event.
This message allows the system to stop the current audio playback, clear the audio queue, and prepare to handle new user input. Contains a Unix timestamp of when the user interruption was detected. For more details, see our Interruptibility Guide
Transcript of the user’s message. Contains the message role and content, along with a from_text field indicating if this message was inserted into the conversation as text from a UserInput message.
Includes an interim field indicating whether the transcript is provisional (words may be repeated or refined in subsequent UserMessage responses as additional audio is processed) or final and complete. Interim transcripts are only sent when the verbose_transcription query parameter is set to true in the initial handshake.
Indicates that the supplemental LLM has detected a need to invoke the specified tool. This message is only received for user-defined function tools.
Contains the tool name, parameters (as a stringified JSON schema), whether a response is required from the developer (either in the form of a ToolResponseMessage or a ToolErrorMessage), the unique tool call ID for tracking the request and response, and the tool type. See our Tool Use Guide for further details.
Return value of the tool call. Contains the output generated by the tool to pass back to EVI. Upon receiving a Tool Call message and successfully invoking the function, this message is sent to convey the result of the function call back to EVI.
For built-in tools implemented on the server, you will receive this message type rather than a ToolCallMessage. See our Tool Use Guide for further details.
Error message from the tool call, not exposed to the LLM or user. Upon receiving a Tool Call message and failing to invoke the function, this message is sent to notify EVI of the tool’s failure.
For built-in tools implemented on the server, you will receive this message type rather than a ToolCallMessage if the tool fails. See our Tool Use Guide for further details.
Settings for this chat session. Session settings are temporary and apply only to the current Chat session.
These settings can be adjusted dynamically based on the requirements of each session to ensure optimal performance and user experience. See our Session Settings Guide for a complete list of configurable settings.
Access token used for authenticating the client. If not provided, an api_key must be provided to authenticate.
The access token is generated using both an API key and a Secret key, which provides an additional layer of security compared to using just an API key.
For more details, refer to the Authentication Strategies Guide.
The unique identifier for an EVI configuration.
Include this ID in your connection request to equip EVI with the Prompt, Language Model, Voice, and Tools associated with the specified configuration. If omitted, EVI will apply default configuration settings.
For help obtaining this ID, see our Configuration Guide.
The version number of the EVI configuration specified by the config_id.
Configs, as well as Prompts and Tools, are versioned. This versioning system supports iterative development, allowing you to progressively refine configurations and revert to previous versions if needed.
Include this parameter to apply a specific version of an EVI configuration. If omitted, the latest version will be applied.
The unique identifier for a Chat Group. Use this field to preserve context from a previous Chat session.
A Chat represents a single session from opening to closing a WebSocket connection. In contrast, a Chat Group is a series of resumed Chats that collectively represent a single conversation spanning multiple sessions. Each Chat includes a Chat Group ID, which is used to preserve the context of previous Chat sessions when starting a new one.
Including the Chat Group ID in the resumed_chat_group_id query parameter is useful for seamlessly resuming a Chat after unexpected network disconnections and for picking up conversations exactly where you left off at a later time. This ensures preserved context across multiple sessions.
There are three ways to obtain the Chat Group ID:
Chat Metadata: Upon establishing a WebSocket connection with EVI, the user receives a Chat Metadata message. This message contains a chat_group_id, which can be used to resume conversations within this chat group in future sessions.
List Chats endpoint: Use the GET /v0/evi/chats endpoint to obtain the Chat Group ID of individual Chat sessions. This endpoint lists all available Chat sessions and their associated Chat Group ID.
List Chat Groups endpoint: Use the GET /v0/evi/chat_groups endpoint to obtain the Chat Group IDs of all Chat Groups associated with an API key. This endpoint returns a list of all available chat groups.
Base64 encoded audio input to insert into the conversation. The content is treated as the user’s speech to EVI and must be streamed continuously. Pre-recorded audio files are not supported.
For optimal transcription quality, the audio data should be transmitted in small chunks. Hume recommends streaming audio with a buffer window of 20 milliseconds (ms), or 100 milliseconds (ms) for web applications. See our Audio Guide for more details on preparing and processing audio.
Settings for this chat session. Session settings are temporary and apply only to the current Chat session.
These settings can be adjusted dynamically based on the requirements of each session to ensure optimal performance and user experience. See our Session Settings Guide for a complete list of configurable settings.
User text to insert into the conversation. Text sent through a User Input message is treated as the user’s speech to EVI. EVI processes this input and provides a corresponding response.
Expression measurement results are not available for User Input messages, as the prosody model relies on audio input and cannot process text alone.
Assistant text to synthesize into spoken audio and insert into the conversation. EVI uses this text to generate spoken audio using our proprietary expressive text-to-speech model.
Our model adds appropriate emotional inflections and tones to the text based on the user’s expressions and the context of the conversation. The synthesized audio is streamed back to the user as an Assistant Message.
Return value of the tool call. Contains the output generated by the tool to pass back to EVI. Upon receiving a Tool Call message and successfully invoking the function, this message is sent to convey the result of the function call back to EVI.
For built-in tools implemented on the server, you will receive this message type rather than a ToolCallMessage. See our Tool Use Guide for further details.
Error message from the tool call, not exposed to the LLM or user. Upon receiving a Tool Call message and failing to invoke the function, this message is sent to notify EVI of the tool’s failure.
For built-in tools implemented on the server, you will receive this message type rather than a ToolCallMessage if the tool fails. See our Tool Use Guide for further details.
Pause responses from EVI. Chat history is still saved and sent after resuming. Once this message is sent, EVI will not respond until a Resume Assistant message is sent.
When paused, EVI won’t respond, but transcriptions of your audio inputs will still be recorded. See our Pause Response Guide for further details.
Resume responses from EVI. Chat history sent while paused will now be sent.
Upon resuming, if any audio input was sent during the pause, EVI will retain context from all messages sent but only respond to the last user message. See our Pause Response Guide for further details.