We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared 16×2048 residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at <a href=\"https://stepaudiollm.github.io/step-audio-3-gen/\" rel=\"nofollow\">https://stepaudiollm.github.io/step-audio-3-gen/</a>.</p>\n","updatedAt":"2026-09-14T07:01:48.954Z","author":{"_id":"66518fd07d8cb2629a514c18","avatarUrl":"/avatars/6280b33a6b1532ee938afd4aa303f709.svg","fullname":"Yang","name":"giantPanda0906","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8521384596824646},"editors":["giantPanda0906"],"editorAvatarUrls":["/avatars/6280b33a6b1532ee938afd4aa303f709.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.12945","authors":[{"_id":"6aa759c47ba345d44ad148db","name":"Bin Lin","hidden":false},{"_id":"6aa759c47ba345d44ad148dc","name":"Bo Zhao","hidden":false},{"_id":"6aa759c47ba345d44ad148dd","name":"Boyang Wang","hidden":false},{"_id":"6aa759c47ba345d44ad148de","name":"Boyang Zhang","hidden":false},{"_id":"6aa759c47ba345d44ad148df","user":{"_id":"642e7c3bbd7a26bb2cb3eef7","avatarUrl":"/avatars/0cf1f4bf805d0b84bc4c0de3f43446ae.svg","isPro":false,"fullname":"Boyong Wu","user":"BoyongWu","type":"user","name":"BoyongWu"},"name":"Boyong Wu","status":"claimed_verified","statusLastChangedAt":"2026-09-14T08:45:04.437Z","hidden":false},{"_id":"6aa759c47ba345d44ad148e0","name":"Chao Yan","hidden":false},{"_id":"6aa759c47ba345d44ad148e1","name":"Chen Geng","hidden":false},{"_id":"6aa759c47ba345d44ad148e2","name":"Chen Wu","hidden":false},{"_id":"6aa759c47ba345d44ad148e3","name":"Cheng Yi","hidden":false},{"_id":"6aa759c47ba345d44ad148e4","name":"Chengli Feng","hidden":false},{"_id":"6aa759c47ba345d44ad148e5","name":"Chenglin Zhu","hidden":false},{"_id":"6aa759c47ba345d44ad148e6","name":"DanNi Wan","hidden":false},{"_id":"6aa759c47ba345d44ad148e7","name":"Daxin Jiang","hidden":false},{"_id":"6aa759c47ba345d44ad148e8","name":"Dongqing Pang","hidden":false},{"_id":"6aa759c47ba345d44ad148e9","name":"Fei Tian","hidden":false},{"_id":"6aa759c47ba345d44ad148ea","name":"Feng Tian","hidden":false},{"_id":"6aa759c47ba345d44ad148eb","name":"Future Li","hidden":false},{"_id":"6aa759c47ba345d44ad148ec","name":"Gang Yu","hidden":false},{"_id":"6aa759c47ba345d44ad148ed","name":"Guanglong Yang","hidden":false},{"_id":"6aa759c47ba345d44ad148ee","name":"Jia Peng","hidden":false},{"_id":"6aa759c47ba345d44ad148ef","name":"Jiahao Song","hidden":false},{"_id":"6aa759c47ba345d44ad148f0","name":"Jiamin Fan","hidden":false},{"_id":"6aa759c47ba345d44ad148f1","name":"Jiangjie Zhen","hidden":false},{"_id":"6aa759c47ba345d44ad148f2","name":"Jianzheng Gao","hidden":false},{"_id":"6aa759c47ba345d44ad148f3","name":"Jun Chen","hidden":false},{"_id":"6aa759c47ba345d44ad148f4","name":"Li Xie","hidden":false},{"_id":"6aa759c47ba345d44ad148f5","name":"Lifang Zhang","hidden":false},{"_id":"6aa759c47ba345d44ad148f6","name":"Lingli Ji","hidden":false},{"_id":"6aa759c47ba345d44ad148f7","name":"Liying Shi","hidden":false},{"_id":"6aa759c47ba345d44ad148f8","name":"Lun Cai","hidden":false},{"_id":"6aa759c47ba345d44ad148f9","name":"Min Xu","hidden":false},{"_id":"6aa759c47ba345d44ad148fa","name":"Na Wang","hidden":false},{"_id":"6aa759c47ba345d44ad148fb","name":"Peilin Li","hidden":false},{"_id":"6aa759c47ba345d44ad148fc","name":"Peng Yang","hidden":false},{"_id":"6aa759c47ba345d44ad148fd","name":"Pengfei Tan","hidden":false},{"_id":"6aa759c47ba345d44ad148fe","name":"Qingjian Lin","hidden":false},{"_id":"6aa759c47ba345d44ad148ff","user":{"_id":"6548d151cf50edb69f3f6934","avatarUrl":"/avatars/f7ccbb303a8dd0439b63532e9b1d0329.svg","isPro":false,"fullname":"jerryxiong","user":"jerryxxx","type":"user","name":"jerryxxx"},"name":"Ruijie Xiong","status":"claimed_verified","statusLastChangedAt":"2026-09-14T09:17:55.610Z","hidden":false},{"_id":"6aa759c47ba345d44ad14900","name":"Runze Li","hidden":false},{"_id":"6aa759c47ba345d44ad14901","name":"Shenghua Hu","hidden":false},{"_id":"6aa759c47ba345d44ad14902","name":"Shi Qiu","hidden":false},{"_id":"6aa759c47ba345d44ad14903","name":"Siqi Tu","hidden":false},{"_id":"6aa759c47ba345d44ad14904","name":"Siyi Zhou","hidden":false},{"_id":"6aa759c47ba345d44ad14905","name":"Tianjiao Deng","hidden":false},{"_id":"6aa759c47ba345d44ad14906","name":"Wanying Lu","hidden":false},{"_id":"6aa759c47ba345d44ad14907","name":"Weiming Niu","hidden":false},{"_id":"6aa759c47ba345d44ad14908","name":"Wen Sun","hidden":false},{"_id":"6aa759c47ba345d44ad14909","name":"WenWen Qu","hidden":false},{"_id":"6aa759c47ba345d44ad1490a","name":"Xiangyu Zhang","hidden":false},{"_id":"6aa759c47ba345d44ad1490b","name":"Xianwei Zhang","hidden":false},{"_id":"6aa759c47ba345d44ad1490c","name":"XiaoSu Su","hidden":false},{"_id":"6aa759c47ba345d44ad1490d","name":"Xing Chen","hidden":false},{"_id":"6aa759c47ba345d44ad1490e","name":"Xinyu Liu","hidden":false},{"_id":"6aa759c47ba345d44ad1490f","name":"Xuerui Yang","hidden":false},{"_id":"6aa759c47ba345d44ad14910","name":"Yang Li","hidden":false},{"_id":"6aa759c47ba345d44ad14911","name":"Yang Yang","hidden":false},{"_id":"6aa759c47ba345d44ad14912","name":"Yechang Huang","hidden":false},{"_id":"6aa759c47ba345d44ad14913","name":"Yibo Zhu","hidden":false},{"_id":"6aa759c47ba345d44ad14914","name":"Yifan Zhang","hidden":false},{"_id":"6aa759c47ba345d44ad14915","name":"Yiyang Xu","hidden":false},{"_id":"6aa759c47ba345d44ad14916","name":"Yu Fu","hidden":false},{"_id":"6aa759c47ba345d44ad14917","name":"Yu Luo","hidden":false},{"_id":"6aa759c47ba345d44ad14918","name":"Yu Zhou","hidden":false},{"_id":"6aa759c47ba345d44ad14919","name":"Yumang Wang","hidden":false},{"_id":"6aa759c47ba345d44ad1491a","name":"Yunzhou Ju","hidden":false},{"_id":"6aa759c47ba345d44ad1491b","name":"Yuxiang Yang","hidden":false},{"_id":"6aa759c47ba345d44ad1491c","name":"Zekai Liu","hidden":false},{"_id":"6aa759c47ba345d44ad1491d","name":"Zengwei Yao","hidden":false},{"_id":"6aa759c47ba345d44ad1491e","name":"Zhenwei Mou","hidden":false},{"_id":"6aa759c47ba345d44ad1491f","name":"Zheqi Dai","hidden":false},{"_id":"6aa759c47ba345d44ad14920","name":"Zhiyue Wu","hidden":false},{"_id":"6aa759c47ba345d44ad14921","name":"Zichao Zhou","hidden":false}],"publishedAt":"2026-09-11T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"StepAudio 3 Gen Technical Report","submittedOnDailyBy":{"_id":"66518fd07d8cb2629a514c18","avatarUrl":"/avatars/6280b33a6b1532ee938afd4aa303f709.svg","isPro":false,"fullname":"Yang","user":"giantPanda0906","type":"user","name":"giantPanda0906"},"summary":"We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared 16 times 2048 residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at https://stepaudiollm.github.io/step-audio-3-gen/.","upvotes":23,"discussionId":"6aa759c57ba345d44ad14922","projectPage":"https://stepaudiollm.github.io/step-audio-3-gen/","ai_summary":"StepAudio 3 Gen is a discrete autoregressive audio generation model using residual vector quantization tokens to unify text-to-speech, voice design, sound effects, and music within a single framework.","ai_keywords":["residual vector quantization","RVQ tokens","autoregressive generator","discrete autoregressive modeling","causal Transformer","StepAudio Tokenizer","interference-aware progressive pretraining","RVQ Adaptor","multi-codebook acoustic representations"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66518fd07d8cb2629a514c18","avatarUrl":"/avatars/6280b33a6b1532ee938afd4aa303f709.svg","isPro":false,"fullname":"Yang","user":"giantPanda0906","type":"user"},{"_id":"67aeb3a4820d941aab225178","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/1BII4nBQiT-rG2kUwXM6V.png","isPro":false,"fullname":"chao yan","user":"yanchaomars","type":"user"},{"_id":"6925a4b3097da9cf89390456","avatarUrl":"/avatars/a926a11436d837bbe06f8d75ce207f17.svg","isPro":false,"fullname":"FeiTian","user":"FeiTia","type":"user"},{"_id":"642e7c3bbd7a26bb2cb3eef7","avatarUrl":"/avatars/0cf1f4bf805d0b84bc4c0de3f43446ae.svg","isPro":false,"fullname":"Boyong Wu","user":"BoyongWu","type":"user"},{"_id":"692807f75181f68b031a0d0c","avatarUrl":"/avatars/67229605aff54214cad13785fd8ff9f0.svg","isPro":false,"fullname":"Jun Chen","user":"RookieJune","type":"user"},{"_id":"64bf871fb9296eef6214ddaa","avatarUrl":"/avatars/889b675306f5302f8da32ce976ea5629.svg","isPro":false,"fullname":"Zhiyue Wu","user":"wzy11","type":"user"},{"_id":"66bf73d6399b14b08f74b59a","avatarUrl":"/avatars/ecb47fe892f2a96b28f9d9800c0f789d.svg","isPro":false,"fullname":"glowol","user":"glowol","type":"user"},{"_id":"68f0aba56a17cb5835e4a99e","avatarUrl":"/avatars/f8f7d2c7b8d387ffb19952814cdbfd8a.svg","isPro":false,"fullname":"Chen","user":"dacxas","type":"user"},{"_id":"6548d151cf50edb69f3f6934","avatarUrl":"/avatars/f7ccbb303a8dd0439b63532e9b1d0329.svg","isPro":false,"fullname":"jerryxiong","user":"jerryxxx","type":"user"},{"_id":"67ce5daeb54cdb852357e0f7","avatarUrl":"/avatars/02504589c103e1095d35cc436ce03c7e.svg","isPro":false,"fullname":"Jinghua Liang","user":"ljh222","type":"user"},{"_id":"665c2858aeeaa961877d5c03","avatarUrl":"/avatars/b5ae1e0989e1761d240a960ddc789e05.svg","isPro":false,"fullname":"咘噜咘噜的碰","user":"mysxs","type":"user"},{"_id":"6aa79ddb02c6b027646fe87d","avatarUrl":"/avatars/880a947f7f21386606fa25ea8935df16.svg","isPro":false,"fullname":"zekai liu","user":"cillian69","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.12945.md","query":{}}">
StepAudio 3 Gen Technical Report
Published on Sep 11
· Submitted by Yang on Sep 14 Abstract
StepAudio 3 Gen is a discrete autoregressive audio generation model using residual vector quantization tokens to unify text-to-speech, voice design, sound effects, and music within a single framework.
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared 16 times 2048 residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at https://stepaudiollm.github.io/step-audio-3-gen/.
Community
We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared 16×2048 residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi-codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio samples are available at https://stepaudiollm.github.io/step-audio-3-gen/.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.12945 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.12945 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.12945 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.