Hugging Face Daily Papers · · 6 min read

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio.</p>\n","updatedAt":"2026-08-04T03:02:56.011Z","author":{"_id":"66569729ea21cfae5f5797c4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66569729ea21cfae5f5797c4/IguwJzljFN3QiEd1bn5BP.jpeg","fullname":"Yu Zhang","name":"AaronZ345","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9171043634414673},"editors":["AaronZ345"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66569729ea21cfae5f5797c4/IguwJzljFN3QiEd1bn5BP.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.02023","authors":[{"_id":"6a714ee6ec5082b9f872ccf1","name":"Yu Zhang","hidden":false},{"_id":"6a714ee6ec5082b9f872ccf2","name":"Ruiqi Li","hidden":false},{"_id":"6a714ee6ec5082b9f872ccf3","name":"Changhao Pan","hidden":false},{"_id":"6a714ee6ec5082b9f872ccf4","name":"Ke Lei","hidden":false},{"_id":"6a714ee6ec5082b9f872ccf5","name":"Xiang Yin","hidden":false},{"_id":"6a714ee6ec5082b9f872ccf6","name":"Cheng Yang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66569729ea21cfae5f5797c4/j7RwwE6myoJNrUuucGA5h.mp4","https://cdn-uploads.huggingface.co/production/uploads/66569729ea21cfae5f5797c4/5B4IeMfkjWuYuVjziauLr.mp4","https://cdn-uploads.huggingface.co/production/uploads/66569729ea21cfae5f5797c4/Eqwhxgn29By86jrX5tWyF.mp4"],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks","submittedOnDailyBy":{"_id":"66569729ea21cfae5f5797c4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66569729ea21cfae5f5797c4/IguwJzljFN3QiEd1bn5BP.jpeg","isPro":false,"fullname":"Yu Zhang","user":"AaronZ345","type":"user","name":"AaronZ345"},"summary":"Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\\#swantale.","upvotes":61,"discussionId":"6a714ee6ec5082b9f872ccf7","projectPage":"https://swanaigc.github.io/#swantale","organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66569729ea21cfae5f5797c4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66569729ea21cfae5f5797c4/IguwJzljFN3QiEd1bn5BP.jpeg","isPro":false,"fullname":"Yu Zhang","user":"AaronZ345","type":"user"},{"_id":"66632347d71a4e1e6cbf891f","avatarUrl":"/avatars/b328a39d1a781866b90852ea44005718.svg","isPro":false,"fullname":"XiangYin","user":"StephenYX","type":"user"},{"_id":"67285bba520ec569b6a9f6ff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/TH5X9DTDrzYzah-5Fop94.png","isPro":true,"fullname":"salah","user":"Davidwang215","type":"user"},{"_id":"6645ea5638f0db40582bddcf","avatarUrl":"/avatars/216aeb4d365e28dff484cc275f9f90d7.svg","isPro":false,"fullname":"Yifu Chen","user":"1f","type":"user"},{"_id":"663a1a61197afc06304c7c32","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663a1a61197afc06304c7c32/d_q4Reb_hLaWHJ55S7ENf.jpeg","isPro":false,"fullname":"Lei Ke","user":"BrokenMoon","type":"user"},{"_id":"66568060c6a8cb4e884be331","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66568060c6a8cb4e884be331/jx8NsxV374oURta6JdTzU.jpeg","isPro":false,"fullname":"PanChanghao","user":"DavidPigeon","type":"user"},{"_id":"69dfc108d9321ac1d79a41ed","avatarUrl":"/avatars/a3346def3eae390e0b79ab52523d9ca3.svg","isPro":false,"fullname":"YiFei Fan","user":"Sixteennights","type":"user"},{"_id":"67bb195bb1077fc4c7ab4dbb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67bb195bb1077fc4c7ab4dbb/gGNMMf_zfR-7B3GL9znQj.jpeg","isPro":false,"fullname":"liangtianle","user":"liangtianle","type":"user"},{"_id":"6684a72f74af0ef94892a3fa","avatarUrl":"/avatars/69c8bb5696f55a83aab627316a629ba8.svg","isPro":false,"fullname":"XUMING HE","user":"hexmSeeU","type":"user"},{"_id":"68120a1375e6e2d3c078cc5b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Xh1AQCiYFggk-AjIT-gcB.png","isPro":false,"fullname":"yangrui","user":"yrainbow","type":"user"},{"_id":"69c674272f69bb1cc01cc986","avatarUrl":"/avatars/33b24348f0be413e9514bf55da0159bd.svg","isPro":false,"fullname":"Tianxiang Li","user":"char12345","type":"user"},{"_id":"667d4ae1144f0f683483f3cd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/6rZVS79wB8jTNUrZI9Ot2.jpeg","isPro":false,"fullname":"Ruiqi Li","user":"RL-2000","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"653b817d32c97d0655575872","name":"ByteDance","fullname":"ByteDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6535c9e88bde2fae19b6fb25/0clr54wj5Ly-RkYU9OXPp.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.02023.md","query":{}}">
Papers
arxiv:2608.02023

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Published on Aug 3
· Submitted by
Yu Zhang
on Aug 4
#1 Paper of the day
Authors:
,

Abstract

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/\#swantale.

Community

Paper submitter about 5 hours ago

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.02023
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.02023 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.02023 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.02023 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers