Hugging Face Daily Papers · · 5 min read

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context.</p>\n","updatedAt":"2026-09-09T05:54:36.184Z","author":{"_id":"630388b0d14428368d1616c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg","fullname":"Ziyang Ma","name":"BoJack","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8693998456001282},"editors":["BoJack"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg"],"reactions":[{"reaction":"👍","users":["tutu0604","worstchan","yfyeung"],"count":3},{"reaction":"🚀","users":["tutu0604","yfyeung"],"count":2}],"isReport":false}},{"id":"6aa1306a432a5d5823cf0522","author":{"_id":"6aa12fd2917300912a1ce0da","avatarUrl":"/avatars/9a41b165a21deb73f6ffa32d555b892a.svg","fullname":"3 patti best","name":"3pattibest","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-09T10:09:46.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"The AuK Technical Report presents an interesting approach to improving speech generation and editing through an open-source foundational model. It is great to see technology becoming more accessible across different digital fields. In gaming, 3 Patti Best also focuses on an accessible and engaging card-game experience for players. Those interested in 3 Patti Best can visit https://3patti-best.pk/ to explore the game.\n","html":"<p>The AuK Technical Report presents an interesting approach to improving speech generation and editing through an open-source foundational model. It is great to see technology becoming more accessible across different digital fields. In gaming, 3 Patti Best also focuses on an accessible and engaging card-game experience for players. Those interested in 3 Patti Best can visit <a href=\"https://3patti-best.pk/\" rel=\"nofollow\">https://3patti-best.pk/</a> to explore the game.</p>\n","updatedAt":"2026-09-09T10:09:46.214Z","author":{"_id":"6aa12fd2917300912a1ce0da","avatarUrl":"/avatars/9a41b165a21deb73f6ffa32d555b892a.svg","fullname":"3 patti best","name":"3pattibest","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9326133728027344},"editors":["3pattibest"],"editorAvatarUrls":["/avatars/9a41b165a21deb73f6ffa32d555b892a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08936","authors":[{"_id":"6aa0e4ccd0174964227bedd0","name":"Ziyang Ma","hidden":false},{"_id":"6aa0e4ccd0174964227bedd1","name":"Zhikang Niu","hidden":false},{"_id":"6aa0e4ccd0174964227bedd2","user":{"_id":"6440656b757aa3c2ad86a67a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6440656b757aa3c2ad86a67a/oVSwa0QfQYJjTBUPwXhO-.jpeg","isPro":false,"fullname":"Wenming Tu","user":"tutu0604","type":"user","name":"tutu0604"},"name":"Wenming Tu","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.708Z","hidden":false},{"_id":"6aa0e4ccd0174964227bedd3","name":"Tianrui Wang","hidden":false},{"_id":"6aa0e4ccd0174964227bedd4","name":"Ruiqi Yan","hidden":false},{"_id":"6aa0e4ccd0174964227bedd5","name":"Junxi Liu","hidden":false},{"_id":"6aa0e4ccd0174964227bedd6","name":"Yanru Huo","hidden":false},{"_id":"6aa0e4ccd0174964227bedd7","name":"Nickk Huang","hidden":false},{"_id":"6aa0e4ccd0174964227bedd8","name":"Yang Liu","hidden":false},{"_id":"6aa0e4ccd0174964227bedd9","name":"Qicong Xie","hidden":false},{"_id":"6aa0e4ccd0174964227bedda","name":"Zeyu Xie","hidden":false},{"_id":"6aa0e4ccd0174964227beddb","name":"Hui Wang","hidden":false},{"_id":"6aa0e4ccd0174964227beddc","name":"Haitao Li","hidden":false},{"_id":"6aa0e4ccd0174964227beddd","name":"Zixuan Jiang","hidden":false},{"_id":"6aa0e4ccd0174964227bedde","name":"Yalin Li","hidden":false},{"_id":"6aa0e4ccd0174964227beddf","name":"Jie Fang","hidden":false},{"_id":"6aa0e4ccd0174964227bede0","name":"Yifan Duan","hidden":false},{"_id":"6aa0e4ccd0174964227bede1","name":"Zeyue Tian","hidden":false},{"_id":"6aa0e4ccd0174964227bede2","name":"Guangzheng Li","hidden":false},{"_id":"6aa0e4ccd0174964227bede3","name":"Haina Zhu","hidden":false},{"_id":"6aa0e4ccd0174964227bede4","name":"Shuyi Wang","hidden":false},{"_id":"6aa0e4ccd0174964227bede5","name":"Jinwen Wang","hidden":false},{"_id":"6aa0e4ccd0174964227bede6","name":"Mingyu Cui","hidden":false},{"_id":"6aa0e4ccd0174964227bede7","name":"Tian Tan","hidden":false},{"_id":"6aa0e4ccd0174964227bede8","name":"Auden","hidden":false},{"_id":"6aa0e4ccd0174964227bede9","name":"Sen Liang","hidden":false},{"_id":"6aa0e4ccd0174964227bedea","name":"Steve Yves","hidden":false},{"_id":"6aa0e4ccd0174964227bedeb","name":"Shan Yang","hidden":false},{"_id":"6aa0e4ccd0174964227bedec","name":"Liefeng Bo","hidden":false},{"_id":"6aa0e4ccd0174964227beded","name":"Zilong Zheng","hidden":false},{"_id":"6aa0e4ccd0174964227bedee","name":"Kai Yu","hidden":false},{"_id":"6aa0e4ccd0174964227bedef","name":"Eng-Siong Chng","hidden":false},{"_id":"6aa0e4ccd0174964227bedf0","name":"Xie Chen","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/630388b0d14428368d1616c5/PJUD7031IjCZPiNdac6aw.mp4"],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing","submittedOnDailyBy":{"_id":"630388b0d14428368d1616c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg","isPro":false,"fullname":"Ziyang Ma","user":"BoJack","type":"user","name":"BoJack"},"summary":"We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.","upvotes":132,"discussionId":"6aa0e4ccd0174964227bedf1","projectPage":"https://auk-project.github.io/","githubRepo":"https://github.com/Tencent-Hunyuan/AuK","githubRepoAddedBy":"user","ai_summary":"AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.","ai_keywords":["multimodal large language model","VAE","rectified-flow Transformer","MMDiT","DiT","preference optimization","reinforcement learning","consistency initialization","Decoupled DMD","classifier-free guidance"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":27,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"630388b0d14428368d1616c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630388b0d14428368d1616c5/Z8O82fDVlB5qkM8Jmq65_.jpeg","isPro":false,"fullname":"Ziyang Ma","user":"BoJack","type":"user"},{"_id":"6440656b757aa3c2ad86a67a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6440656b757aa3c2ad86a67a/oVSwa0QfQYJjTBUPwXhO-.jpeg","isPro":false,"fullname":"Wenming Tu","user":"tutu0604","type":"user"},{"_id":"6447d332ab5c7251886d6fd1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6447d332ab5c7251886d6fd1/bm5nwIp5CA_HosO8wXFvI.jpeg","isPro":false,"fullname":"ZhikangNiu-SII","user":"zkniu","type":"user"},{"_id":"6786007928014d20f37fb228","avatarUrl":"/avatars/08fa31dc953362bb263127c00aae922d.svg","isPro":false,"fullname":"Yanru Huo","user":"iHateTheWorld555","type":"user"},{"_id":"662554dc79d897d7dd1ca7a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662554dc79d897d7dd1ca7a4/ybZsbGZv7aekFV5fPOPT2.jpeg","isPro":false,"fullname":"xcczach","user":"xcczach","type":"user"},{"_id":"63e65c742d2c508de9fbc7ab","avatarUrl":"/avatars/67f2c37bcafd15bb9991385c13d550fe.svg","isPro":false,"fullname":"yangguanrou","user":"yhaha","type":"user"},{"_id":"6864e44244dc36f3a7b896eb","avatarUrl":"/avatars/21f5dfd9858e3a9042c3a0cbc9fbcc69.svg","isPro":false,"fullname":"Li","user":"yayll","type":"user"},{"_id":"67b3f529d21021f9eb29fa36","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67b3f529d21021f9eb29fa36/2bJpoGozgLCUs6VmhZqXv.jpeg","isPro":false,"fullname":"Zixuan Jiang","user":"Andrew0425","type":"user"},{"_id":"67852e5a3d49517c54365945","avatarUrl":"/avatars/3ab8df5628e91bf89ac39f6fd6955c3e.svg","isPro":false,"fullname":"Zezhong Qian","user":"XuWuLingYu","type":"user"},{"_id":"67320e4996d5da4801a69199","avatarUrl":"/avatars/bf85472384886c19121a5b32bb4dbeea.svg","isPro":false,"fullname":"AlexTYJ","user":"AlexTYJ","type":"user"},{"_id":"670cc406a48acae9350394b3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/v6_kCriBnU2qXFuyrYVI4.png","isPro":false,"fullname":"PengchaoFeng","user":"the-bird-F","type":"user"},{"_id":"6946651d4c20c7f3d0f671e1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6946651d4c20c7f3d0f671e1/0KZpnxCydGniukHJAMyU-.png","isPro":false,"fullname":"Hengtao Wu","user":"HengtaoWu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08936.md","query":{}}">
Papers
arxiv:2609.08936

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

Published on Sep 8
· Submitted by
Ziyang Ma
on Sep 9
#2 Paper of the day
Authors:
,

Abstract

AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.

Community

Paper submitter about 8 hours ago

We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context.

The AuK Technical Report presents an interesting approach to improving speech generation and editing through an open-source foundational model. It is great to see technology becoming more accessible across different digital fields. In gaming, 3 Patti Best also focuses on an accessible and engaging card-game experience for players. Those interested in 3 Patti Best can visit https://3patti-best.pk/ to explore the game.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.08936
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.08936 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers