Run Kokoro-82M: Apache-Licensed Text-to-Speech in a Few Lines
Kokoro-82M is an 82M-parameter open-weight TTS model. How to install it, generate speech, and use it for narration, pronunciation fixes, and captions.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Kokoro-82M is an open-weight text-to-speech model with 82 million parameters. It was trained by @rzvzn on the StyleTTS 2 architecture by Li et al. The authors say it delivers “comparable quality to larger models” while being significantly faster and more cost-efficient. It is small, and its weights are Apache 2.0 licensed. The card says it was trained only on permissive or non-copyrighted audio. Together, those points make it a good candidate for commercial products. The card puts hosted API pricing (as of April 2025) at under $1 per million input characters.
Key specs
| Parameters | 82 million |
| Current release | v1.0, published 2025 Jan 27 |
| Languages and voices (v1.0) | 8 languages, 54 voices (see VOICES.md) |
| Architecture | StyleTTS 2 + ISTFTNet, decoder only: no diffusion, no encoder release |
| Output sample rate | 24,000 Hz (as used in the card’s example) |
| Training data | A few hundred hours of permissive/non-copyrighted audio plus IPA phoneme labels |
| Training cost | About $1,000 total (1,000 A100 80GB GPU hours across v0.19 and v1.0) |
| Hosted price | Under $1 per million characters as of April 2025; roughly 1,000 characters of input is about 1 minute of audio |
| License | Apache 2.0 |
The card links to an EVAL.md but doesn’t include benchmark numbers itself. Treat “comparable quality to larger models” as the authors’ claim. Listen to SAMPLES.md or try the demo Space before you commit.
Install
Kokoro ships as the kokoro Python package. According to the card, it uses the misaki G2P library under the hood. The card’s setup also installs espeak-ng as a system package. The card’s install commands are for a Debian or Ubuntu environment such as Colab:
pip install -q "kokoro>=0.9.2" soundfilesudo apt-get -qq -y install espeak-ngColab runs as root, so the card leaves out sudo. On a local Debian or Ubuntu machine, apt-get install needs root, so keep sudo in the command.
A general shell note that the card doesn’t mention: quote kokoro>=0.9.2. Without quotes, the shell reads > as a redirect, and pip installs an unpinned kokoro instead. The card’s own command is unquoted.
Run it
This is the card’s example rewritten as a plain script, without the notebook display code:
from kokoro import KPipelineimport soundfile as sf
pipeline = KPipeline(lang_code='a')
text = '''[Kokoro](/kˈOkəɹO/) is an open-weight TTS model with 82 million parameters.With Apache-licensed weights, [Kokoro](/kˈOkəɹO/) can be deployed anywherefrom production environments to personal projects.'''
generator = pipeline(text, voice='af_heart')for i, (gs, ps, audio) in enumerate(generator): print(i, gs, ps) sf.write(f'{i}.wav', audio, 24000)The pipeline returns a generator, not a single clip. Each item unpacks into (gs, ps, audio). The card doesn’t define these fields. Going by the names and how the card uses them, they appear to be:
gs: the text segment (graphemes)ps: its phonemesaudio: the waveform for that segment, which the card writes at 24,000 Hz
The card writes one file per item. Longer text appears to be split into several items, so you get several files unless you join them yourself.
Batch narration into a single file
A common job is turning a set of articles or docs pages into one audio file each. This script reads every .txt file in a folder and streams all the segments into a single WAV per input:
import sysfrom pathlib import Path
import soundfile as sffrom kokoro import KPipeline
SAMPLE_RATE = 24000
def narrate(pipeline, text, out_path, voice='af_heart'): with sf.SoundFile(out_path, 'w', samplerate=SAMPLE_RATE, channels=1) as f: for _, _, audio in pipeline(text, voice=voice): f.write(audio)
def main(in_dir, out_dir): pipeline = KPipeline(lang_code='a') # load once, reuse for every file out = Path(out_dir) out.mkdir(parents=True, exist_ok=True) for path in sorted(Path(in_dir).glob('*.txt')): text = path.read_text(encoding='utf-8') target = out / f'{path.stem}.wav' narrate(pipeline, text, target) print(f'{path.name} -> {target}')
if __name__ == '__main__': main(sys.argv[1], sys.argv[2])Run it with python batch_narrate.py articles/ audio/. Create the KPipeline once and reuse it, rather than building a new one per file. To estimate cost before you scale up, use the card’s rule of thumb: about 1,000 characters of input per minute of audio.
Fix pronunciation of product names
TTS models often mispronounce brand names, acronyms, and jargon. The card’s example text writes the model’s own name as [Kokoro](/kˈOkəɹO/). The card doesn’t explain this syntax. From the example, it looks like an inline phoneme override written as a Markdown link: [word](/phonemes/). The script below assumes the syntax works the way the example suggests, and applies a glossary of overrides before synthesis:
import re
import soundfile as sffrom kokoro import KPipeline
# Map words to phoneme strings. The Kokoro entry is from the model card;# add your own entries in the same /.../ format.GLOSSARY = { 'Kokoro': 'kˈOkəɹO',}
def apply_glossary(text, glossary): for word, phonemes in glossary.items(): # Skip words that are already wrapped as [word](/.../). pattern = rf'(?<!\[)\b{re.escape(word)}\b(?!\]\()' # A function replacement stops backslashes in phonemes from being # read as regex escapes. text = re.sub(pattern, lambda m: f'[{m.group(0)}](/{phonemes}/)', text) return text
pipeline = KPipeline(lang_code='a')raw = 'Kokoro is an open-weight TTS model. We use [Kokoro](/kˈOkəɹO/) for our audio docs.'text = apply_glossary(raw, GLOSSARY)print(text)
with sf.SoundFile('glossary.wav', 'w', samplerate=24000, channels=1) as f: for gs, ps, audio in pipeline(text, voice='af_heart'): print(ps) # confirm the override reached the phoneme string f.write(audio)Printing ps is the quickest way to check whether an override worked. If a word still sounds wrong, check the phoneme string before you blame the voice. The card doesn’t document which phoneme symbols are valid. Copy the style of the card’s example and test by ear.
Generate captions alongside the audio
Each segment includes its source text and its audio, so you can compute timestamps from the audio length and write an SRT subtitle file next to the WAV:
import soundfile as sffrom kokoro import KPipeline
SAMPLE_RATE = 24000
def fmt(seconds): ms = int(round(seconds * 1000)) h, ms = divmod(ms, 3_600_000) m, ms = divmod(ms, 60_000) s, ms = divmod(ms, 1000) return f'{h:02}:{m:02}:{s:02},{ms:03}'
text = '''Kokoro is an open-weight TTS model with 82 million parameters.It was trained exclusively on permissive and non-copyrighted audio data.The total training cost was about one thousand dollars.'''
pipeline = KPipeline(lang_code='a')cues = []t = 0.0
with sf.SoundFile('narration.wav', 'w', samplerate=SAMPLE_RATE, channels=1) as f: for i, (gs, ps, audio) in enumerate(pipeline(text, voice='af_heart'), start=1): f.write(audio) duration = len(audio) / SAMPLE_RATE cues.append(f'{i}\n{fmt(t)} --> {fmt(t + duration)}\n{gs.strip()}\n') t += duration
with open('narration.srt', 'w', encoding='utf-8') as srt: srt.write('\n'.join(cues))
print(f'Wrote {len(cues)} cues, {t:.1f}s of audio')This timing is per segment, not per word. That works for subtitles on narrated videos but not for karaoke-style word highlighting. The card doesn’t document word-level timing.
Gotchas
- System dependency. The card installs
espeak-ngbefore it usesKPipeline. It only shows theapt-getcommand, which needs root outside Colab. On other operating systems you’ll need to install it yourself. - Output is chunked.
pipeline(...)yields one(gs, ps, audio)tuple per segment. If you only grab the first item, you lose most of the audio on longer inputs. - Sample rate. The card writes and plays audio at 24,000 Hz. Use that same rate when saving, or playback speed and pitch will be off.
- Languages and voices. The card’s example only shows
lang_code='a'withvoice='af_heart'. For the other v1.0 languages, see the voice list in VOICES.md and the Advanced Usage section of the GitHub repo. The model card’s metadata lists only English. - No encoder release. The card says there’s no encoder release and doesn’t describe any voice cloning. The voices it documents are the 54 listed in VOICES.md.
- Training data attribution. The weights are Apache 2.0. Some of the v1.0 training data was CC BY audio (Koniwa
tncunder CC BY 3.0, SIWIS under CC BY 4.0). The card lists these under its attribution section. - Scam sites. The card warns that sites with “kokoro” in their root domain, such as kokorottsai_com and kokorotts_net, are not affiliated with the model. The official sources are the Hugging Face repo and
github.com/hexgrad/kokoro.
When to pick it
Pick Kokoro-82M when you need TTS you can self-host and use commercially. The weights are Apache 2.0, and the card says the model was trained on permissive data. The card makes no promise of legal clearance, though, so review the training-data notes yourself. It fits audio versions of articles, narration for docs or videos, and voice output for apps and agents. It’s small, and the setup is a pip install plus espeak-ng. If you’d rather not run it yourself, the card says hosted APIs charged under $1 per million characters as of April 2025.
Look elsewhere if you need to clone a specific person’s voice. The card documents no cloning and says there’s no encoder release. If you need word-level timestamps, check the kokoro package docs first, because the card doesn’t document them. Kokoro is also a poor fit if your language isn’t among the 8 covered in VOICES.md. The card makes quality claims but publishes no benchmark numbers on the page itself. Listen to the samples and check EVAL.md against your own content before you replace an existing TTS provider.
Related
AuK: Clone, Edit, and Clean Up Speech Through One Instruction Interface
Tencent's MIT-licensed AuK does zero-shot TTS, speech editing, enhancement, and separation from natural-language instructions.
tencent/AuK
YuE2-3B: Generate and Edit Full Songs on a 24GB GPU
YuE2-3B turns lyrics and a style prompt into a full 48 kHz stereo song. You can edit the melody and chords as an ABC score and render it again.
m-a-p/YuE2-3B
Run Audio8 ASR Infinite for 24/7 Chinese and English Transcription
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite