Model deep dive · Oct 8, 2026
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite
Model deep dive · Oct 8, 2026
Tencent's MIT-licensed AuK does zero-shot TTS, speech editing, enhancement, and separation from natural-language instructions.
tencent/AuK
Model deep dive · Oct 8, 2026
Kokoro-82M is an 82M-parameter open-weight TTS model. How to install it, generate speech, and use it for narration, pronunciation fixes, and captions.
hexgrad/Kokoro-82M
Model deep dive · Oct 8, 2026
YuE2-3B turns lyrics and a style prompt into a full 48 kHz stereo song. You can edit the melody and chords as an ABC score and render it again.
m-a-p/YuE2-3B