Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Run Breeze TTS 2 for Voice Cloning, Design and Streaming

Breeze TTS 2 is an open-weight English and Chinese TTS model with voice cloning, text-described voices and a streaming API. Weights are non-commercial.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Breeze TTS 2 is an open-weight text-to-speech model from BreezeBlue. It speaks English and Chinese and is built for real-time use. BreezeBlue says it ranks #1 among open-weight models on the Artificial Analysis TTS leaderboard. What sets it apart is that you steer the voice with plain-language instructions. You can describe a voice from scratch (“a warm, thoughtful young woman…”) without any reference audio. Or you can clone a speaker from clean reference audio and its exact transcript, then tell it how to deliver the line. The weights are free to download, but the license only allows research and non-commercial use.

Key specs

Languages English, Chinese (one model)
Modes Voice Clone, Voice Design (no reference audio), Voice Direction (clone + instruction)
Vocal events (laugh), (cough), (clears throat), (sigh) in English; [笑], [咳嗽], [清嗓子], [叹气] in Chinese
Time to first audio Under 40 ms (warmed-up fast path, H100)
Real-time factor 0.32, about 3.1x real time (warmed-up fast path, H100)
GPU memory ~7.7 GiB eager, ~14.4 GiB with --fast-all
Recommended GPU 12 GB for eager, 24 GB for the fast path
Output (API) Streaming mono 24 kHz signed 16-bit little-endian PCM

The latency numbers come from an H100 with the fast path warmed up. The card gives no latency figures for eager mode or for other GPUs.

Install

You need Linux, Python 3.10 or newer, and a CUDA-capable NVIDIA GPU. Clone the inference code and install its dependencies:

Terminal window
git clone https://github.com/breezeblue-ai/breeze-tts.git
cd breeze-tts
python -m pip install -r requirements.txt

Download the Breeze TTS 2 checkpoint from the Hugging Face repo. All the card’s examples expect it at ../breeze-tts-2, next to the breeze-tts folder. The card says every model component ships in that checkpoint.

The GitHub repo also includes a Docker image for the tested CUDA environment. According to the card, it targets H100/Hopper (sm90) by default:

Terminal window
bash docker/build.sh
# For A100:
FLASH_ATTN_CUDA_ARCHS=80 bash docker/build.sh

Run it

Voice Design is the quickest test because it needs no reference audio:

Terminal window
python infer.py ../breeze-tts-2 \
--text "(sigh) Welcome aboard. Your journey begins now." \
--instruction "A warm, thoughtful young woman with a clear voice and a calm, reflective delivery." \
--cfg-scale 4 \
--output outputs/voice_design_en.wav

To clone a voice, pass a clean reference clip and its exact transcript:

Terminal window
python infer.py ../breeze-tts-2 \
--ref-audio reference_en.wav \
--ref-text "This is the exact transcript of the English reference audio." \
--text "(sigh) It is good to hear your voice again after all this time." \
--output outputs/voice_clone_en.wav

Add --instruction (and --cfg-scale 4) to a clone command to get Voice Direction. That keeps the speaker’s identity while changing tone, emotion or pace.

Streaming speech for a voice agent

To use the model in a voice agent, run the bundled HTTP server. It streams raw PCM as the audio is generated, so playback can start before the whole sentence is done. Start it with --fast-all for the low-latency path. This needs about 14.4 GiB of VRAM and adds cold-start time:

Terminal window
python -m breeze_infer.api ../breeze-tts-2 --host 0.0.0.0 --port 7860 --fast-all

The client below sends a Voice Direction request using the form fields from the card. It reports when the first audio bytes arrive and wraps the PCM in a WAV header using the documented format (mono, 24 kHz, 16-bit). It needs requests on the client side (pip install requests).

stream_client.py
import time
import wave
import requests
URL = "http://127.0.0.1:7860/v1/audio/speech"
fields = {
"cfg_scale": "4",
"ref_text": "This is the exact transcript of the reference audio.",
"text": "(clears throat) We need to discuss what happened last night.",
"instruction": "Speak slowly with a restrained, serious tone.",
"seed": "42",
}
start = time.perf_counter()
first_chunk_at = None
with open("reference.wav", "rb") as ref, \
requests.post(URL, data=fields, files={"ref_audio": ref}, stream=True) as resp, \
wave.open("voice_direction.wav", "wb") as out:
resp.raise_for_status()
out.setnchannels(1) # mono
out.setsampwidth(2) # signed 16-bit
out.setframerate(24000) # 24 kHz
for chunk in resp.iter_content(chunk_size=4096):
if not chunk:
continue
if first_chunk_at is None:
first_chunk_at = time.perf_counter() - start
# In a real agent, push `chunk` straight to your audio output here.
out.writeframes(chunk)
if first_chunk_at is None:
print("no audio received")
else:
print(f"first audio after {first_chunk_at * 1000:.0f} ms (includes HTTP and upload)")

The server handles one request at a time, so a multi-user agent will need a queue or several processes.

Batch narration with a cloned voice

For audiobooks, course material or product demos (non-commercial only), you can loop over a script and clone one narrator for every line. This uses only the CLI flags from the card.

Run it from inside the breeze-tts repo folder, with the checkpoint in ../breeze-tts-2, because infer.py and the checkpoint are given as relative paths. It uses sys.executable so the subprocess runs under the same Python (and venv) you launched it with. narrator.wav and REF_TEXT are placeholders: replace them with your own clean clip and its exact transcript.

batch_narrate.py
import subprocess
import sys
from pathlib import Path
CHECKPOINT = "../breeze-tts-2"
REF_AUDIO = "narrator.wav" # placeholder: your clean reference clip
REF_TEXT = "This is the exact transcript of the narrator reference audio." # placeholder: its exact transcript
lines = [
"Chapter one. The storm arrived a day early.",
"(sigh) Nobody in the village had prepared for it.",
"By morning, the river had climbed past the old stone bridge.",
]
out_dir = Path("outputs/narration")
out_dir.mkdir(parents=True, exist_ok=True)
for i, line in enumerate(lines, start=1):
out_path = out_dir / f"line_{i:03d}.wav"
subprocess.run(
[
sys.executable, "infer.py", CHECKPOINT,
"--ref-audio", REF_AUDIO,
"--ref-text", REF_TEXT,
"--text", line,
"--output", str(out_path),
],
check=True,
)
print("wrote", out_path)

Each call starts a new process, which likely means reloading the model each time. For long scripts, keeping the API server running and sending each line to it may be faster, though we haven’t measured this. The API returns raw PCM, so you would need to add the WAV header to each line as the streaming client above does.

Bilingual character voices without reference audio

Voice Design is useful for prototyping games and interactive fiction, where you want distinct characters but have no voice actors. The card says to write the instruction in the same language as the target text. Vocal events also use a different syntax in each language.

As with the batch script, run this from inside the breeze-tts folder with the checkpoint at ../breeze-tts-2.

design_cast.py
import subprocess
import sys
from pathlib import Path
CHECKPOINT = "../breeze-tts-2"
cast = [
{
"name": "guide_en",
"text": "(laugh) Welcome aboard. Your journey begins now.",
"instruction": "A warm, thoughtful young woman with a clear voice and a calm, reflective delivery.",
},
{
"name": "host_zh",
"text": "[笑] 欢迎来到今晚的故事时间,让我们一起开始吧。",
"instruction": "一位温柔自信的年轻女性,声音清晰,语气亲切,表达轻快而富有感染力。",
},
]
Path("outputs/cast").mkdir(parents=True, exist_ok=True)
for role in cast:
subprocess.run(
[
sys.executable, "infer.py", CHECKPOINT,
"--text", role["text"],
"--instruction", role["instruction"],
"--cfg-scale", "4",
"--output", f"outputs/cast/{role['name']}.wav",
],
check=True,
)

The card doesn’t cover keeping a designed voice consistent across many lines. One untested idea: save an output you like, then use it as --ref-audio, with its exact text as --ref-text, in clone or direction mode.

Gotchas

  • The license blocks commercial use of self-hosted output. The code is Apache 2.0. The weights, derivative models and anything you generate on your own hardware fall under the BreezeBlue Research and Non-Commercial License. Only output generated through BreezeBlue’s paid hosted platform can be used commercially, and a subscription does not unlock the open weights.
  • Linux and NVIDIA CUDA only. The card lists no CPU, macOS or AMD path.
  • The card shows no transformers API. The repo carries a transformers library tag, but the card’s examples all run through infer.py or breeze_infer.api from the GitHub repo.
  • The fast path costs memory and startup time. --fast-all roughly doubles VRAM use (7.7 to 14.4 GiB) and adds graph warmup at cold start. The per-stage flags (--fast-codec, --fast-depth-decoder, etc.) are meant for profiling and debugging.
  • Docker defaults to H100. On an A100, build with FLASH_ATTN_CUDA_ARCHS=80.
  • Reference quality matters. --ref-text must be the exact transcript of the clip, and the clip should be clean speech with little background noise.
  • Language-specific syntax. Vocal events use (laugh) in English and [笑] in Chinese. Write instructions in the same language as the text.
  • The API returns raw PCM, not WAV. You need to add a header yourself (mono, 24 kHz, s16le), as the streaming client above does.
  • One request at a time. The streaming server runs a single request at a time.
  • Consent. The license prohibits cloning or impersonating someone without authorization.

When to pick it

Pick Breeze TTS 2 for research, prototypes or personal projects that need expressive English or Chinese speech. It fits best when you want to control delivery with plain-language instructions or build voices from a text description. It runs in eager mode on a single 12 GB NVIDIA GPU. For the low-latency streaming numbers BreezeBlue reports, you need the fast path, which calls for a 24 GB card (the published figures are from an H100). Voice Direction, which clones a speaker and then changes the delivery, is its most distinctive feature.

Skip it if you plan to ship a commercial product on self-hosted inference, because the license does not allow it. In that case, use BreezeBlue’s paid hosted API or a model with a permissive weights license. It is also the wrong choice if you need languages other than English and Chinese, have no NVIDIA GPU, or need a server that handles many concurrent requests out of the box.

Related