Hugging Face Daily Papers · · 2 min read

WanSong v1.0 Technical Report

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

this is open source or close?</p>\n","updatedAt":"2026-07-17T14:21:41.000Z","author":{"_id":"66d2e8d41ba71ac4c08e0307","avatarUrl":"/avatars/4cfc4b11d66739d072f7166bb9d43f08.svg","fullname":"SR.Suzume","name":"srsuzume","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9630875587463379},"editors":["srsuzume"],"editorAvatarUrls":["/avatars/4cfc4b11d66739d072f7166bb9d43f08.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.14749","authors":[{"_id":"6a59954d6c2e371e6ca380e0","name":"Binghui Chen","hidden":false},{"_id":"6a59954d6c2e371e6ca380e1","name":"Pandeng Li","hidden":false},{"_id":"6a59954d6c2e371e6ca380e2","name":"Yu Liu","hidden":false},{"_id":"6a59954d6c2e371e6ca380e3","name":"Jingren Zhou","hidden":false}],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"WanSong v1.0 Technical Report","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\\eg, AR followed by diffusion), WanSong is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.","upvotes":8,"discussionId":"6a59954d6c2e371e6ca380e4","organization":{"_id":"67bc7cd418dd753c02a82684","name":"Wan-AI","fullname":"Wan-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67b610677ea7952def8b29c6/N6jQbbeaa_FcUY-wI1dgG.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6270324ebecab9e2dcf245de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6270324ebecab9e2dcf245de/cMbtWSasyNlYc9hvsEEzt.jpeg","isPro":false,"fullname":"Kye Gomez","user":"kye","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"67bbade8a8c89b98ec377944","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67bbade8a8c89b98ec377944/HPtKDo8fnKr4OxpN1Z17D.png","isPro":false,"fullname":"Urodoc Oncall","user":"UDCAI","type":"user"},{"_id":"66d2e8d41ba71ac4c08e0307","avatarUrl":"/avatars/4cfc4b11d66739d072f7166bb9d43f08.svg","isPro":false,"fullname":"SR.Suzume","user":"srsuzume","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"69bd4f7b22b6ccba8310e4aa","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/YFn0ccaFDkkhr4tXJQ6BH.jpeg","isPro":false,"fullname":"佐藤蒼","user":"victoriajohnson","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"69ccf171be44414ace4df2db","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/tfSoQatYvAMQMu_sHVc1A.png","isPro":false,"fullname":"이민서","user":"jacksonmartin51","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67bc7cd418dd753c02a82684","name":"Wan-AI","fullname":"Wan-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67b610677ea7952def8b29c6/N6jQbbeaa_FcUY-wI1dgG.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.14749.md","query":{}}">
Papers
arxiv:2607.14749

WanSong v1.0 Technical Report

Published on Jul 16
· Submitted by
taesiri
on Jul 17
Authors:
,

Abstract

Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present WanSong, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), WanSong is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.

Community

this is open source or close?

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.14749
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.14749 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.14749 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.14749 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers