Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Run Kokoro-82M: Apache-Licensed Text-to-Speech in a Few Lines

Kokoro-82M is an 82M-parameter open-weight TTS model. How to install it, generate speech, and use it for narration, pronunciation fixes, and captions.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Kokoro-82M is an open-weight text-to-speech model with 82 million parameters. It was trained by @rzvzn on the StyleTTS 2 architecture by Li et al. The authors say it delivers “comparable quality to larger models” while being significantly faster and more cost-efficient. It is small, and its weights are Apache 2.0 licensed. The card says it was trained only on permissive or non-copyrighted audio. Together, those points make it a good candidate for commercial products. The card puts hosted API pricing (as of April 2025) at under $1 per million input characters.

Key specs

Parameters 82 million
Current release v1.0, published 2025 Jan 27
Languages and voices (v1.0) 8 languages, 54 voices (see VOICES.md)
Architecture StyleTTS 2 + ISTFTNet, decoder only: no diffusion, no encoder release
Output sample rate 24,000 Hz (as used in the card’s example)
Training data A few hundred hours of permissive/non-copyrighted audio plus IPA phoneme labels
Training cost About $1,000 total (1,000 A100 80GB GPU hours across v0.19 and v1.0)
Hosted price Under $1 per million characters as of April 2025; roughly 1,000 characters of input is about 1 minute of audio
License Apache 2.0

The card links to an EVAL.md but doesn’t include benchmark numbers itself. Treat “comparable quality to larger models” as the authors’ claim. Listen to SAMPLES.md or try the demo Space before you commit.

Install

Kokoro ships as the kokoro Python package. According to the card, it uses the misaki G2P library under the hood. The card’s setup also installs espeak-ng as a system package. The card’s install commands are for a Debian or Ubuntu environment such as Colab:

Terminal window
pip install -q "kokoro>=0.9.2" soundfile
sudo apt-get -qq -y install espeak-ng

Colab runs as root, so the card leaves out sudo. On a local Debian or Ubuntu machine, apt-get install needs root, so keep sudo in the command.

A general shell note that the card doesn’t mention: quote kokoro>=0.9.2. Without quotes, the shell reads > as a redirect, and pip installs an unpinned kokoro instead. The card’s own command is unquoted.

Run it

This is the card’s example rewritten as a plain script, without the notebook display code:

hello_kokoro.py
from kokoro import KPipeline
import soundfile as sf
pipeline = KPipeline(lang_code='a')
text = '''
[Kokoro](/kˈOkəɹO/) is an open-weight TTS model with 82 million parameters.
With Apache-licensed weights, [Kokoro](/kˈOkəɹO/) can be deployed anywhere
from production environments to personal projects.
'''
generator = pipeline(text, voice='af_heart')
for i, (gs, ps, audio) in enumerate(generator):
print(i, gs, ps)
sf.write(f'{i}.wav', audio, 24000)

The pipeline returns a generator, not a single clip. Each item unpacks into (gs, ps, audio). The card doesn’t define these fields. Going by the names and how the card uses them, they appear to be:

  • gs: the text segment (graphemes)
  • ps: its phonemes
  • audio: the waveform for that segment, which the card writes at 24,000 Hz

The card writes one file per item. Longer text appears to be split into several items, so you get several files unless you join them yourself.

Batch narration into a single file

A common job is turning a set of articles or docs pages into one audio file each. This script reads every .txt file in a folder and streams all the segments into a single WAV per input:

batch_narrate.py
import sys
from pathlib import Path
import soundfile as sf
from kokoro import KPipeline
SAMPLE_RATE = 24000
def narrate(pipeline, text, out_path, voice='af_heart'):
with sf.SoundFile(out_path, 'w', samplerate=SAMPLE_RATE, channels=1) as f:
for _, _, audio in pipeline(text, voice=voice):
f.write(audio)
def main(in_dir, out_dir):
pipeline = KPipeline(lang_code='a') # load once, reuse for every file
out = Path(out_dir)
out.mkdir(parents=True, exist_ok=True)
for path in sorted(Path(in_dir).glob('*.txt')):
text = path.read_text(encoding='utf-8')
target = out / f'{path.stem}.wav'
narrate(pipeline, text, target)
print(f'{path.name} -> {target}')
if __name__ == '__main__':
main(sys.argv[1], sys.argv[2])

Run it with python batch_narrate.py articles/ audio/. Create the KPipeline once and reuse it, rather than building a new one per file. To estimate cost before you scale up, use the card’s rule of thumb: about 1,000 characters of input per minute of audio.

Fix pronunciation of product names

TTS models often mispronounce brand names, acronyms, and jargon. The card’s example text writes the model’s own name as [Kokoro](/kˈOkəɹO/). The card doesn’t explain this syntax. From the example, it looks like an inline phoneme override written as a Markdown link: [word](/phonemes/). The script below assumes the syntax works the way the example suggests, and applies a glossary of overrides before synthesis:

glossary.py
import re
import soundfile as sf
from kokoro import KPipeline
# Map words to phoneme strings. The Kokoro entry is from the model card;
# add your own entries in the same /.../ format.
GLOSSARY = {
'Kokoro': 'kˈOkəɹO',
}
def apply_glossary(text, glossary):
for word, phonemes in glossary.items():
# Skip words that are already wrapped as [word](/.../).
pattern = rf'(?<!\[)\b{re.escape(word)}\b(?!\]\()'
# A function replacement stops backslashes in phonemes from being
# read as regex escapes.
text = re.sub(pattern, lambda m: f'[{m.group(0)}](/{phonemes}/)', text)
return text
pipeline = KPipeline(lang_code='a')
raw = 'Kokoro is an open-weight TTS model. We use [Kokoro](/kˈOkəɹO/) for our audio docs.'
text = apply_glossary(raw, GLOSSARY)
print(text)
with sf.SoundFile('glossary.wav', 'w', samplerate=24000, channels=1) as f:
for gs, ps, audio in pipeline(text, voice='af_heart'):
print(ps) # confirm the override reached the phoneme string
f.write(audio)

Printing ps is the quickest way to check whether an override worked. If a word still sounds wrong, check the phoneme string before you blame the voice. The card doesn’t document which phoneme symbols are valid. Copy the style of the card’s example and test by ear.

Generate captions alongside the audio

Each segment includes its source text and its audio, so you can compute timestamps from the audio length and write an SRT subtitle file next to the WAV:

captions.py
import soundfile as sf
from kokoro import KPipeline
SAMPLE_RATE = 24000
def fmt(seconds):
ms = int(round(seconds * 1000))
h, ms = divmod(ms, 3_600_000)
m, ms = divmod(ms, 60_000)
s, ms = divmod(ms, 1000)
return f'{h:02}:{m:02}:{s:02},{ms:03}'
text = '''
Kokoro is an open-weight TTS model with 82 million parameters.
It was trained exclusively on permissive and non-copyrighted audio data.
The total training cost was about one thousand dollars.
'''
pipeline = KPipeline(lang_code='a')
cues = []
t = 0.0
with sf.SoundFile('narration.wav', 'w', samplerate=SAMPLE_RATE, channels=1) as f:
for i, (gs, ps, audio) in enumerate(pipeline(text, voice='af_heart'), start=1):
f.write(audio)
duration = len(audio) / SAMPLE_RATE
cues.append(f'{i}\n{fmt(t)} --> {fmt(t + duration)}\n{gs.strip()}\n')
t += duration
with open('narration.srt', 'w', encoding='utf-8') as srt:
srt.write('\n'.join(cues))
print(f'Wrote {len(cues)} cues, {t:.1f}s of audio')

This timing is per segment, not per word. That works for subtitles on narrated videos but not for karaoke-style word highlighting. The card doesn’t document word-level timing.

Gotchas

  • System dependency. The card installs espeak-ng before it uses KPipeline. It only shows the apt-get command, which needs root outside Colab. On other operating systems you’ll need to install it yourself.
  • Output is chunked. pipeline(...) yields one (gs, ps, audio) tuple per segment. If you only grab the first item, you lose most of the audio on longer inputs.
  • Sample rate. The card writes and plays audio at 24,000 Hz. Use that same rate when saving, or playback speed and pitch will be off.
  • Languages and voices. The card’s example only shows lang_code='a' with voice='af_heart'. For the other v1.0 languages, see the voice list in VOICES.md and the Advanced Usage section of the GitHub repo. The model card’s metadata lists only English.
  • No encoder release. The card says there’s no encoder release and doesn’t describe any voice cloning. The voices it documents are the 54 listed in VOICES.md.
  • Training data attribution. The weights are Apache 2.0. Some of the v1.0 training data was CC BY audio (Koniwa tnc under CC BY 3.0, SIWIS under CC BY 4.0). The card lists these under its attribution section.
  • Scam sites. The card warns that sites with “kokoro” in their root domain, such as kokorottsai_com and kokorotts_net, are not affiliated with the model. The official sources are the Hugging Face repo and github.com/hexgrad/kokoro.

When to pick it

Pick Kokoro-82M when you need TTS you can self-host and use commercially. The weights are Apache 2.0, and the card says the model was trained on permissive data. The card makes no promise of legal clearance, though, so review the training-data notes yourself. It fits audio versions of articles, narration for docs or videos, and voice output for apps and agents. It’s small, and the setup is a pip install plus espeak-ng. If you’d rather not run it yourself, the card says hosted APIs charged under $1 per million characters as of April 2025.

Look elsewhere if you need to clone a specific person’s voice. The card documents no cloning and says there’s no encoder release. If you need word-level timestamps, check the kokoro package docs first, because the card doesn’t document them. Kokoro is also a poor fit if your language isn’t among the 8 covered in VOICES.md. The card makes quality claims but publishes no benchmark numbers on the page itself. Listen to the samples and check EVAL.md against your own content before you replace an existing TTS provider.

Related