Tencent Releases AuK: Open-Source Speech Generation and Editing Model
Key Info
Tencent has officially released AuK, an open-source foundation model for unified speech generation and editing. It combines natural-language instructions with reference audio through a single interface.
Highlights
- AuK supports zero-shot TTS, instruction-controlled generation, content editing, whisper conversion, de-accenting, timbre/style/emotion editing, and speed/pitch control.
- It also provides enhancement, denoising, multi-speaker, and music separation capabilities.
- AuK-Flash, a faster variant, delivers 4-step inference and runs ~4.5× faster under matched conditions; code, weights, and related resources are being released.