AuK: Clone, Edit, and Clean Up Speech Through One Instruction Interface
Tencent's MIT-licensed AuK does zero-shot TTS, speech editing, enhancement, and separation from natural-language instructions.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
AuK is a 1.5B-parameter foundation model for speech generation and editing from Tencent’s Hunyuan team, released under the MIT license. It puts zero-shot TTS, voice-description TTS, content editing, acoustic and paralinguistic editing, speech enhancement, and source separation behind one natural-language instruction interface. One set of AuK weights covers all of these tasks. Running it also needs a separate MLLM encoder (Qwen2.5-Omni-3B) and a VAE, so it isn’t a single download.
Key specs
| Parameters | 1.5B |
| Training data | “Millions of hours of diverse audio data” |
| Architecture | Diffusion transformer + layer-fusion weights; MLLM encoder and VAE loaded separately at runtime |
| MLLM encoder | Qwen2.5-Omni-3B (separate download) |
| Variants | AuK (base, quality) and AuK-Flash (distilled, 4-step inference) |
| Interface | Natural-language instructions for every task |
| Serving | Day-0 support in SGLang-Omni |
| License | MIT |
Supported tasks, grouped the way the card groups them:
- Speech generation: zero-shot TTS (speak text in a reference voice), instruct TTS (generate a voice from a description, no reference audio)
- Content editing: replace, insert, or remove words in a recording; rewrite lyrics in a singing recording while keeping the melody and voice
- Acoustic editing: pitch (semitones), speed (output length scales with the factor), volume (decibels)
- Paralinguistic editing: emotion, timbre, de-accent, add or remove nonverbal sounds (breaths, laughs, coughs), normal/whisper conversion
- Enhancement and separation: denoise and dereverberate, keep one speaker by talking order, extract singing voice from music, keep a target speaker identified by what they say
The card shows benchmark results only as a chart image (assets/performance.png), with no numbers in the text. This post doesn’t quote scores. Look at the chart and the arXiv paper (2609.08936) before you rely on quality claims.
Install
The card lists three downloads: the AuK weights, the Qwen2.5-Omni-3B encoder, and optionally the Flash variant.
pip install -U "huggingface_hub[cli]"
# AuK-Basehf download tencent/AuK --local-dir ./ckpts/AuK
# AuK-Flash (4-step distilled)hf download tencent/AuK-Flash --local-dir ./ckpts/AuK-Flash
# MLLM Encoderhf download Qwen/Qwen2.5-Omni-3B --local-dir ./ckpts/Qwen2.5-Omni-3BIf you’re in a region where ModelScope is faster:
pip install -U modelscope
modelscope download --model Tencent-Hunyuan/AuK --local_dir ./ckpts/AuKmodelscope download --model Tencent-Hunyuan/AuK-Flash --local_dir ./ckpts/AuK-Flashmodelscope download --model Qwen/Qwen2.5-Omni-3B --local_dir ./ckpts/Qwen2.5-Omni-3BThe card gives this expected layout:
ckpts/├── AuK/├── AuK-Flash/ # optional└── Qwen2.5-Omni-3B/For the inference package, Gradio app, ComfyUI nodes, and fine-tuning, the card points to the GitHub README. It doesn’t include the install commands for those itself.
Run it
The only launch command the card gives is for SGLang-Omni, which had Day-0 support:
python -m sglang_omni.cli serve --model-path tencent/AuKThis command points at the base model. The card doesn’t say whether SGLang-Omni also serves AuK-Flash, so don’t assume the Flash variant runs there without checking.
Request formats and configuration options are in the SGLang Omni AuK Cookbook. Optimization work is tracked in a GitHub issue, the AuK optimization roadmap, so the integration may still be changing. The model card doesn’t show a request payload or a Python API, so this post doesn’t make one up. Take the exact call shape from the SGLang Omni AuK Cookbook, or from the project’s own AuK Cookbook, which has CLI and Python examples.
Voice cloning and voice design
Zero-shot TTS is the obvious first use. You give it a reference clip and target text, and AuK speaks the text in that voice. If you don’t have a reference recording, instruct TTS builds a voice from a text description. One use we’d suggest is prototyping characters, narrators, or IVR prompts before you hire voice talent. The card doesn’t show instruct-TTS output quality, so judge that yourself on the demo Space first.
Both tasks use the same natural-language instruction interface. The card doesn’t say which tasks the SGLang-Omni integration supports, so check the SGLang Omni AuK Cookbook before you assume a given task works there.
The instruction templates for both modes are in the AuK Cookbook under 1.1 Zero-shot TTS and 1.2 Instruct TTS, with CLI and Python examples. If you’re producing a lot of audio, look at AuK-Flash. The card describes it as distilled for fast 4-step inference, which makes it the option to try when you need throughput. As noted above, the card doesn’t confirm SGLang-Omni support for Flash, so you may need the project’s own code to run it.
Fixing recordings without re-recording
The editing tasks set AuK apart from a plain TTS model. Say a podcast or course recording has a wrong product name, a cough mid-sentence, or a flat delivery. The card lists tasks for each of those:
- Speech content editing rewrites what is said: it replaces, inserts, or removes text.
- Nonverbal editing removes (or adds) breaths, laughs, and coughs.
- Emotion editing changes the emotion while keeping the content and voice.
- Speed, pitch, and volume editing take numeric amounts (speed factor, semitones, decibels).
- De-accent and whisper conversion cover more specialized cleanup.
All of these use the same weights and the same instruction interface as TTS. A post-production tool could expose them as separate operations without loading a different model for each one. The card doesn’t say whether SGLang-Omni serves every editing task, so check its cookbook or use the project’s own code. The templates are in the AuK Cookbook under sections 2, 3, and 4.
Cleaning and separating audio before ASR
AuK can also work as a preprocessing step in front of a speech recognition or voice-cloning pipeline:
- Speech enhancement denoises, dereverberates, or restores clear speech.
- Speech separation keeps one speaker by talking order (for example, the first speaker) and removes the others.
- Target speaker extraction keeps the speaker identified by what they say. This helps when you know a phrase from the transcript but don’t have a clean voice sample.
- Music separation extracts the singing voice from a mix, or keeps all human voices.
An idea we haven’t tested: run enhancement on a noisy reference clip, then use the cleaned clip as the reference for zero-shot TTS. Neither the card nor this post confirms that this improves results. Templates are in AuK Cookbook section 5.
Gotchas
- The AuK weights aren’t enough on their own. The card says the model checkpoint contains the diffusion transformer and layer-fusion weights. The Qwen2.5-Omni-3B MLLM encoder is a separate download. The card also says the VAE is loaded from separate files at runtime, but it doesn’t say where those files come from. Check the GitHub README for that.
- Missing-key warnings are expected. The card says missing
text_encoder.*keys during checkpoint loading are normal, because the encoder is loaded separately. Don’t spend time debugging them. - Memory planning is on you. The card gives no VRAM, GPU, or dtype guidance. It calls AuK a 1.5B model but doesn’t say whether that count includes the Qwen2.5-Omni-3B encoder. Plan for both the AuK checkpoint and the encoder, and test on your own hardware before you commit to a deployment size.
- There’s no standard pipeline. The repo has no library tag and no
transformerssnippet. Plan on the project’s own code (GitHub), SGLang-Omni, Gradio, or ComfyUI. - Separation means specific things. Speech separation selects a speaker by talking order, and target speaker extraction selects by spoken content. Neither one is a general “isolate this voice from a sample” feature as described on the card.
- Speed edits change duration. Output length scales with the speed factor, so adjust any downstream timing (subtitles, video sync) to match.
- Licensing covers AuK only. AuK is MIT. Qwen2.5-Omni-3B is a separate model with its own license, which the card doesn’t state, so check it yourself before shipping commercially.
- Voice cloning needs consent. The license allows a lot. Your users’ expectations and local law may not, so get permission before cloning someone’s voice.
When to pick it
Pick AuK when you need several speech operations, like cloning, word-level edits, emotion changes, denoising, and separation, and want one set of weights and one instruction interface instead of a chain of task-specific models. The MIT license and SGLang-Omni support make it a reasonable candidate for self-hosting. AuK-Flash is available when you need faster inference, though the card only shows the SGLang-Omni command for the base model.
Skip it or wait if you want a drop-in pip install with a short Python snippet. The card sends you to external cookbooks for every actual call. If you need published benchmark numbers to justify the choice, read the paper first, because the card only shows a chart. And if all you need is plain TTS on a small GPU, carrying the separate Qwen2.5-Omni-3B encoder alongside AuK may cost more resources than you need.
Related
Run Breeze TTS 2 for Voice Cloning, Design and Streaming
Breeze TTS 2 is an open-weight English and Chinese TTS model with voice cloning, text-described voices and a streaming API. Weights are non-commercial.
BreezeBlue/Breeze-TTS-2
Run Kokoro-82M: Apache-Licensed Text-to-Speech in a Few Lines
Kokoro-82M is an 82M-parameter open-weight TTS model. How to install it, generate speech, and use it for narration, pronunciation fixes, and captions.
hexgrad/Kokoro-82M
Run Audio8 ASR Infinite for 24/7 Chinese and English Transcription
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
Edge0/Audio8-ASR-Infinite