TTS Python Quickstart Guide

Step-by-step guide for integrating the TTS API using Hume’s Python SDK.

This guide shows how to use Hume’s Text-to-Speech API using Hume’s Python SDK. It assumes you are running on a system with access to speakers for playback and with the PortAudio library installed.

It demonstrates:

  1. Using an existing voice.
  2. Create a new voice via a prompt.
  3. Continuing from previous speech.
  4. Providing “acting instructions” to modulate the voice.
  5. Generating speech from live input.

The complete code for the example in this guide is available on GitHub.

Environment Setup

Set up a Python virtual environment and install the required packages. We recommend uv or Poetry for managing your environment and dependencies, but you can also use venv and pip.

uv
uv init
uv add hume[microphone] python-dotenv

Authenticating the HumeClient

You must authenticate to use the Hume TTS API. Your API key can be retrieved from the Hume AI platform.

This example uses python-dotenv. Place your API key in a .env file at the root of your project.

.env
echo "HUME_API_KEY=your_api_key_here" > .env

First, use your API key to instantiate the AsyncHumeClient, importing as necessary.

# app.py
import os
from hume import AsyncHumeClient
from dotenv import load_dotenv
load_dotenv()
api_key = os.getenv("HUME_API_KEY")
if not api_key:
raise EnvironmentError("HUME_API_KEY not found in environment variables.")
hume = AsyncHumeClient(api_key=api_key)

Using a pre-existing voice

Use this method if you want to synthesize speech with a high-quality voice from Hume’s Voice Library, or specify provider='CUSTOM_VOICE' to use a voice that you created previously via the Hume Platform or the API.

import base64
from hume.empathic_voice.chat.audio.audio_utilities import play_audio_streaming
from hume.tts import PostedUtterance, PostedUtteranceVoiceWithName
utterance = PostedUtterance(
text="Dogs became domesticated between 23,000 and 30,000 years ago.",
voice=PostedUtteranceVoiceWithName(name='Ava Song', provider='HUME_AI')
)
stream = hume.tts.synthesize_json_streaming(
utterances=[utterance],
strip_headers=True,
version="1"
)
await play_audio_streaming(base64.b64decode(chunk.audio) async for chunk in stream)

Create a new voice via a prompt

The Voice Creation API allows you to create custom voices programmatically, via prompting. There are two steps to creating a voice:

  1. Send a description of the voice, along with sample text that is characteristic of the voice, to the standard tts endpoint without specifying a voice.
  2. Take the generation_id from one of the resulting audio samples, and use it to create a new voice with the Voice Creation API.
import base64
import time
from hume.empathic_voice.chat.audio.audio_utilities import play_audio
from hume.tts import PostedUtterance
result1 = await hume.tts.synthesize_json(
utterances=[PostedUtterance(
description="Crisp, upper-class British accent with impeccably articulated consonants and perfectly placed vowels. Authoritative and theatrical, as if giving a lecture.",
text="The science of speech. That\'s my profession; also my hobby. Happy is the man who can make a living by his hobby!"
)],
num_generations=2,
)
sample_number = 1
for generation in result1.generations:
print(f'Playing option {sample_number}... সন')
audio_data = base64.b64decode(generation.audio)
await play_audio(audio_data)
sample_number += 1
# Prompt user to select which voice they prefer
print('\nWhich voice did you prefer?')
print('1. First voice (generation ID:', result1.generations[0].generation_id, ') সন')
print('2. Second voice (generation ID:', result1.generations[1].generation_id, ') সন')
try:
user_choice = input('Enter your choice (1 or 2): ').strip()
except EOFError:
user_choice = '1'
print('No input available, selecting option 1')
selected_index = int(user_choice) - 1
if selected_index not in [0, 1]:
raise ValueError('Invalid choice. Please select 1 or 2.')
selected_generation_id = result1.generations[selected_index].generation_id
print(f'Selected voice option {selected_index + 1} (generation ID: {selected_generation_id})')
# Save the selected voice
voice_name = f'higgins-{int(time.time() * 1000)}'
await hume.tts.voices.create(
name=voice_name,
generation_id=selected_generation_id,
)
print(f'Created voice: {voice_name}')

Continuing previous speech

You can make new speech sound like a natural continuation from previous speech by providing the generation_id of the previous audio in the context parameter. This helps maintain consistency in tone, pacing, and emotional state.

Additionally, you can provide “acting instructions” using the description field alongside an existing voice. When you specify both a voice and a description, the description modulates the voice’s tone, emotion, and delivery style while maintaining the core voice characteristics.

import base64
from hume.empathic_voice.chat.audio.audio_utilities import play_audio_streaming
from hume.tts import PostedUtterance, PostedUtteranceVoiceWithName, PostedContextWithGenerationId
stream = hume.tts.synthesize_json_streaming(
utterances=[PostedUtterance(
voice=PostedUtteranceVoiceWithName(name=voice_name),
text="YOU can spot an Irishman or a Yorkshireman by his brogue. I can place any man within six miles. I can place him within two miles in London. Sometimes within two streets.",
description="Bragging about his abilities"
)],
context=PostedContextWithGenerationId(
generation_id=selected_generation_id
),
strip_headers=True
)
await play_audio_streaming(base64.b64decode(chunk.audio) async for chunk in stream)

Generating speech from live input

If you need to generate speech from text that is being produced in real-time, you can use the bidirectional streaming WebSocket endpoint at /v0/tts/stream/input.

Support for connecting to the WebSocket directly is coming soon to the Python SDK, for the time being, this example shows how you can implement a simple WebSocket client yourself.

First, create a streaming.py file with the StreamingTtsClient:

streaming.py
import asyncio
import json
from typing import AsyncGenerator, Dict, Any
import websockets
from hume.tts import PublishTts, SnippetAudioChunk
class StreamingTtsClient:
def __init__(self, websocket: websockets.WebSocketClientProtocol):
self._websocket: websockets.WebSocketClientProtocol = websocket
self._message_queue = asyncio.Queue()
@classmethod
async def connect(cls, api_key: str) -> "StreamingTtsClient":
client = await websockets.connect(
f"wss://api.hume.ai/v0/tts/stream/input?api_key={api_key}&instant_mode=true&strip_headers=true&no_binary=true"
)
ret = cls(client)
try:
asyncio.create_task(ret._message_handler())
except (websockets.exceptions.InvalidURI, websockets.exceptions.InvalidHandshake) as e:
raise RuntimeError(f"Failed to connect to WebSocket: {e}") from e
return ret
async def _message_handler(self):
try:
while True:
message = await self._websocket.recv()
try:
parsed_json = json.loads(message)
chunk = SnippetAudioChunk.model_validate(parsed_json)
await self._message_queue.put(chunk)
except Exception as parse_error:
print(f"Error parsing message: {parse_error}")
print(f"Raw message was: {message}")
except websockets.exceptions.ConnectionClosed:
print("WebSocket connection closed")
await self._message_queue.put(None) # Signal end of stream
except Exception as e:
print(f"Error in message handler: {e}")
await self._message_queue.put(None)
async def __aiter__(self) -> AsyncGenerator[SnippetAudioChunk, None]:
while True:
message = await self._message_queue.get()
if message is None:
break
yield message
def send(self, tts: PublishTts):
message = tts.json()
print(f"Sending TTS message: {message}")
asyncio.create_task(self._websocket.send(message))
async def _send_dict(self, message: Dict[str, Any]):
await self._websocket.send(json.dumps(message))
async def close(self):
if self._websocket and not self._websocket.closed:
await self._websocket.close()

You can use the client as follows:

import asyncio
import base64
from streaming import StreamingTtsClient
from hume.tts import PublishTts
from hume.empathic_voice.chat.audio.audio_utilities import play_audio_streaming
stream = await StreamingTtsClient.connect(api_key)
# Helper functions for flushing and closing the stream
def send_flush():
asyncio.create_task(stream._send_dict({"flush": True}))
def send_close():
asyncio.create_task(stream._send_dict({"close": True}))
async def send_input():
print("Sending TTS messages...")
stream.send(PublishTts(text="Hello world."))
send_flush()
print('Waiting 8 seconds...')
await asyncio.sleep(8)
stream.send(PublishTts(text="Goodbye, world."))
send_flush()
print("Closing stream...")
send_close()
async def handle_messages():
await play_audio_streaming(base64.b64decode(chunk.audio) async for chunk in stream)
await asyncio.gather(handle_messages(), send_input())

Running the Example

uv run app.py