tencent/AuK-Flash · Hugging Face
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| AuK-Flash: Fast 4-Step Speech Generation and Editing
IntroductionAuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface. AuK has two variants:
This repository contains the official weights for AuK-Flash, the distilled variant with fast 4-step inference. Supported TasksAuK exposes every task through the same natural-language instruction interface. The table below groups the supported tasks by category, with a short description and a link to its section in the Cookbook, which provides instruction templates plus CLI and Python examples.
[link] [comments] |
More from r/LocalLLaMA
-
GGUFs in transformers natively!
Sep 23
-
Pirate Face - pirate bay for LLMs
Sep 23
-
DeepSeek and Moonshot AI face Beijing's probe over potential data leaks to Anthropic
Sep 23
-
Nathan Lambert's written Congressional testimony on the state of open models - Chinese open-weight downloads now 2x America's, >80% of OpenRouter open-model usage
Sep 23
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.