We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context.</p>\n","updatedAt":"2026-09-09T05:54:36.184Z","author":{"_id":"630388b0d14428368d1616c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg","fullname":"Ziyang Ma","name":"BoJack","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8693998456001282},"editors":["BoJack"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg"],"reactions":[{"reaction":"👍","users":["tutu0604","worstchan","yfyeung"],"count":3},{"reaction":"🚀","users":["tutu0604","yfyeung"],"count":2}],"isReport":false}},{"id":"6aa1306a432a5d5823cf0522","author":{"_id":"6aa12fd2917300912a1ce0da","avatarUrl":"/avatars/9a41b165a21deb73f6ffa32d555b892a.svg","fullname":"3 patti best","name":"3pattibest","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-09T10:09:46.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"The AuK Technical Report presents an interesting approach to improving speech generation and editing through an open-source foundational model. It is great to see technology becoming more accessible across different digital fields. In gaming, 3 Patti Best also focuses on an accessible and engaging card-game experience for players. Those interested in 3 Patti Best can visit https://3patti-best.pk/ to explore the game.\n","html":"<p>The AuK Technical Report presents an interesting approach to improving speech generation and editing through an open-source foundational model. It is great to see technology becoming more accessible across different digital fields. In gaming, 3 Patti Best also focuses on an accessible and engaging card-game experience for players. Those interested in 3 Patti Best can visit <a href=\"https://3patti-best.pk/\" rel=\"nofollow\">https://3patti-best.pk/</a> to explore the game.</p>\n","updatedAt":"2026-09-09T10:09:46.214Z","author":{"_id":"6aa12fd2917300912a1ce0da","avatarUrl":"/avatars/9a41b165a21deb73f6ffa32d555b892a.svg","fullname":"3 patti best","name":"3pattibest","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9326133728027344},"editors":["3pattibest"],"editorAvatarUrls":["/avatars/9a41b165a21deb73f6ffa32d555b892a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08936","authors":[{"_id":"6aa0e4ccd0174964227bedd0","name":"Ziyang Ma","hidden":false},{"_id":"6aa0e4ccd0174964227bedd1","name":"Zhikang Niu","hidden":false},{"_id":"6aa0e4ccd0174964227bedd2","user":{"_id":"6440656b757aa3c2ad86a67a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6440656b757aa3c2ad86a67a/oVSwa0QfQYJjTBUPwXhO-.jpeg","isPro":false,"fullname":"Wenming Tu","user":"tutu0604","type":"user","name":"tutu0604"},"name":"Wenming Tu","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.708Z","hidden":false},{"_id":"6aa0e4ccd0174964227bedd3","name":"Tianrui Wang","hidden":false},{"_id":"6aa0e4ccd0174964227bedd4","name":"Ruiqi Yan","hidden":false},{"_id":"6aa0e4ccd0174964227bedd5","name":"Junxi Liu","hidden":false},{"_id":"6aa0e4ccd0174964227bedd6","name":"Yanru Huo","hidden":false},{"_id":"6aa0e4ccd0174964227bedd7","name":"Nickk Huang","hidden":false},{"_id":"6aa0e4ccd0174964227bedd8","name":"Yang Liu","hidden":false},{"_id":"6aa0e4ccd0174964227bedd9","name":"Qicong Xie","hidden":false},{"_id":"6aa0e4ccd0174964227bedda","name":"Zeyu Xie","hidden":false},{"_id":"6aa0e4ccd0174964227beddb","name":"Hui Wang","hidden":false},{"_id":"6aa0e4ccd0174964227beddc","name":"Haitao Li","hidden":false},{"_id":"6aa0e4ccd0174964227beddd","name":"Zixuan Jiang","hidden":false},{"_id":"6aa0e4ccd0174964227bedde","name":"Yalin Li","hidden":false},{"_id":"6aa0e4ccd0174964227beddf","name":"Jie Fang","hidden":false},{"_id":"6aa0e4ccd0174964227bede0","name":"Yifan Duan","hidden":false},{"_id":"6aa0e4ccd0174964227bede1","name":"Zeyue Tian","hidden":false},{"_id":"6aa0e4ccd0174964227bede2","name":"Guangzheng Li","hidden":false},{"_id":"6aa0e4ccd0174964227bede3","name":"Haina Zhu","hidden":false},{"_id":"6aa0e4ccd0174964227bede4","name":"Shuyi Wang","hidden":false},{"_id":"6aa0e4ccd0174964227bede5","name":"Jinwen Wang","hidden":false},{"_id":"6aa0e4ccd0174964227bede6","name":"Mingyu Cui","hidden":false},{"_id":"6aa0e4ccd0174964227bede7","name":"Tian Tan","hidden":false},{"_id":"6aa0e4ccd0174964227bede8","name":"Auden","hidden":false},{"_id":"6aa0e4ccd0174964227bede9","name":"Sen Liang","hidden":false},{"_id":"6aa0e4ccd0174964227bedea","name":"Steve Yves","hidden":false},{"_id":"6aa0e4ccd0174964227bedeb","name":"Shan Yang","hidden":false},{"_id":"6aa0e4ccd0174964227bedec","name":"Liefeng Bo","hidden":false},{"_id":"6aa0e4ccd0174964227beded","name":"Zilong Zheng","hidden":false},{"_id":"6aa0e4ccd0174964227bedee","name":"Kai Yu","hidden":false},{"_id":"6aa0e4ccd0174964227bedef","name":"Eng-Siong Chng","hidden":false},{"_id":"6aa0e4ccd0174964227bedf0","name":"Xie Chen","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/630388b0d14428368d1616c5/PJUD7031IjCZPiNdac6aw.mp4"],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing","submittedOnDailyBy":{"_id":"630388b0d14428368d1616c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg","isPro":false,"fullname":"Ziyang Ma","user":"BoJack","type":"user","name":"BoJack"},"summary":"We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.","upvotes":132,"discussionId":"6aa0e4ccd0174964227bedf1","projectPage":"https://auk-project.github.io/","githubRepo":"https://github.com/Tencent-Hunyuan/AuK","githubRepoAddedBy":"user","ai_summary":"AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.","ai_keywords":["multimodal large language model","VAE","rectified-flow Transformer","MMDiT","DiT","preference optimization","reinforcement learning","consistency initialization","Decoupled DMD","classifier-free guidance"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":27,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"630388b0d14428368d1616c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg","isPro":false,"fullname":"Ziyang Ma","user":"BoJack","type":"user"},{"_id":"6440656b757aa3c2ad86a67a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6440656b757aa3c2ad86a67a/oVSwa0QfQYJjTBUPwXhO-.jpeg","isPro":false,"fullname":"Wenming Tu","user":"tutu0604","type":"user"},{"_id":"6447d332ab5c7251886d6fd1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6447d332ab5c7251886d6fd1/bm5nwIp5CA_HosO8wXFvI.jpeg","isPro":false,"fullname":"ZhikangNiu-SII","user":"zkniu","type":"user"},{"_id":"6786007928014d20f37fb228","avatarUrl":"/avatars/08fa31dc953362bb263127c00aae922d.svg","isPro":false,"fullname":"Yanru Huo","user":"iHateTheWorld555","type":"user"},{"_id":"662554dc79d897d7dd1ca7a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662554dc79d897d7dd1ca7a4/ybZsbGZv7aekFV5fPOPT2.jpeg","isPro":false,"fullname":"xcczach","user":"xcczach","type":"user"},{"_id":"63e65c742d2c508de9fbc7ab","avatarUrl":"/avatars/67f2c37bcafd15bb9991385c13d550fe.svg","isPro":false,"fullname":"yangguanrou","user":"yhaha","type":"user"},{"_id":"6864e44244dc36f3a7b896eb","avatarUrl":"/avatars/21f5dfd9858e3a9042c3a0cbc9fbcc69.svg","isPro":false,"fullname":"Li","user":"yayll","type":"user"},{"_id":"67b3f529d21021f9eb29fa36","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67b3f529d21021f9eb29fa36/2bJpoGozgLCUs6VmhZqXv.jpeg","isPro":false,"fullname":"Zixuan Jiang","user":"Andrew0425","type":"user"},{"_id":"67852e5a3d49517c54365945","avatarUrl":"/avatars/3ab8df5628e91bf89ac39f6fd6955c3e.svg","isPro":false,"fullname":"Zezhong Qian","user":"XuWuLingYu","type":"user"},{"_id":"67320e4996d5da4801a69199","avatarUrl":"/avatars/bf85472384886c19121a5b32bb4dbeea.svg","isPro":false,"fullname":"AlexTYJ","user":"AlexTYJ","type":"user"},{"_id":"670cc406a48acae9350394b3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/v6_kCriBnU2qXFuyrYVI4.png","isPro":false,"fullname":"PengchaoFeng","user":"the-bird-F","type":"user"},{"_id":"6946651d4c20c7f3d0f671e1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6946651d4c20c7f3d0f671e1/0KZpnxCydGniukHJAMyU-.png","isPro":false,"fullname":"Hengtao Wu","user":"HengtaoWu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08936.md","query":{}}">
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
Abstract
AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.
Community
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context.
The AuK Technical Report presents an interesting approach to improving speech generation and editing through an open-source foundational model. It is great to see technology becoming more accessible across different digital fields. In gaming, 3 Patti Best also focuses on an accessible and engaging card-game experience for players. Those interested in 3 Patti Best can visit https://3patti-best.pk/ to explore the game.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.08936 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.