Speech-to-Speech (EVI)
Speech-to-Speech (EVI)
Hume’s Empathic Voice Interface (EVI) is an advanced, real-time emotionally intelligent voice AI.
Hume is sunsetting the TTS and EVI APIs. Access ends November 13, 2026 at 12:01 a.m. EST, and both APIs remain fully supported until then. Any data in your account is permanently deleted after November 13, 2026. To compare alternative voice models, see Real World VoiceEQ. For questions, contact Hume support or email [email protected].
Hume’s Empathic Voice Interface (EVI) is an advanced, real-time emotionally intelligent voice AI. EVI measures users’ nuanced vocal modulations and responds to them using a speech-language model, which guides language and speech generation.
By processing the tune, rhythm, and timbre of speech, EVI unlocks a variety of new capabilities, like knowing when to speak and generating more empathic language with the right tone of voice.
These features enable smoother and more satisfying voice-based interactions between humans and AI, opening new possibilities for personal AI, customer service, accessibility, robotics, immersive gaming, VR experiences, and much more.
To try EVI in your browser, use the EVI Playground in the Hume platform.
EVI features
Version comparison
Basic capabilities
Empathic AI Features
Quickstart
Kickstart your integration with our quickstart guides for Next.js, TypeScript, and Python. Each guide walks you through integrating the EVI API, capturing user audio, and playing back EVI’s response so you can get up and running quickly.
Building with EVI
EVI chat sessions run over a real-time WebSocket connection, enabling fluid, interactive dialogue. Users speak naturally while EVI analyzes their vocal expression and responds with emotionally intelligent speech.
Authentication
REST endpoints support the API key authentication strategy.
specify your API key in the X-HUME-API-KEY header of your request.
The EVI WebSocket endpoint supports both the API key and Token authentication strategies, specify your API key or Access token in the query parameters of your request.
Configuration
Before starting a session, you’ll need a voice and a configuration.
- Design a voice, clone an existing one, or select one from Hume’s extensive Voice Library.
- Build an EVI configuration to define system behavior, voice selection, and other settings.
Connection
The EVI Playground is the easiest way to test your configuration. It lets you speak directly with EVI using your selected voice and settings, without writing any code.
To begin a conversation, connect using the EVI WebSocket URL start streaming the user’s audio input, via audio_input messages. EVI responds in real time with a sequence of structured messages:
- user_message: Message containing a transcript of the user’s message along with their vocal expression measures
- assistant_message: Message containing EVI’s response content.
- audio_output: EVI’s response audio
corresponding with the
assistant_message - assistant_end: Message denoting the end of EVI’s response.
Developer tools
Hume provides a suite of developer tools to integrate and customize EVI.
API limits
The following limits apply to Hume’s Speech-to-Speech (EVI) API.
The EVI API supports thousands of concurrent sessions. To increase limits:
- Upgrade your account to Business or Enterprise.
- Submit the Sales & Partnerships form.

