Motion-Omni replaces the speech-then-motion cascade with one end-to-end model that natively generates dialogue speech together with explicit facial expression, hand, upper-body, and lower-body motion, all from the hidden states that produce the speech. It stays within 2% of the teacher cascade on reference-free motion metrics, runs 5.4× faster at RTF 0.78 (faster than real time), and reaches a 2.62% WER, the lowest among the omni-modal systems compared.</p>\n","updatedAt":"2026-09-07T02:47:13.123Z","author":{"_id":"6355473d525beaee688b7ba1","avatarUrl":"/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg","fullname":"Wei Tao","name":"itaowe","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8925673961639404},"editors":["itaowe"],"editorAvatarUrls":["/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg"],"reactions":[{"reaction":"🔥","users":["ChengqianMa"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04250","authors":[{"_id":"6a9e226fde5ea82090db6591","name":"Chengqian Ma","hidden":false},{"_id":"6a9e226fde5ea82090db6592","name":"Wei Tao","hidden":false},{"_id":"6a9e226fde5ea82090db6593","name":"Haoyu Zhang","hidden":false},{"_id":"6a9e226fde5ea82090db6594","name":"Yiwen Guo","hidden":false}],"publishedAt":"2026-08-28T00:00:00.000Z","submittedOnDailyAt":"2026-09-07T00:00:00.000Z","title":"Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue","submittedOnDailyBy":{"_id":"6355473d525beaee688b7ba1","avatarUrl":"/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg","isPro":false,"fullname":"Wei Tao","user":"itaowe","type":"user","name":"itaowe"},"summary":"An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. We present Motion-Omni, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech. Joint training is not optional here: with the speech pathway frozen, motion remains misaligned with the audio, and co-adapting the LLM, Speech Generator and Motion Generator under both objectives is what recovers alignment while retaining spoken-dialogue ability. Supervision comes from a scalable, model-agnostic pipeline that pseudo-labels consistent-voice speech responses with a replaceable motion teacher, yielding 422,856 quality-ranked pairs (1,402 hours). We further release SwDA-500 and, to our knowledge, the first public evaluation protocol for stochastic open-ended full-body spoken dialogue, matching audio across motion systems while unifying rendering, automatic metrics, human evaluation, and latency measurement. Instantiated with a Qwen2.5-7B-Instruct backbone, Motion-Omni-Q7 matches the same-audio teacher cascade to within 2% on reference-free motion metrics while responding 5.4 x faster (RTF=0.78, faster than real time), surpasses all non-teacher cascades on beat correlation and diversity, and reaches a 2.62% word error rate, the lowest among the omni-modal systems compared.","upvotes":19,"discussionId":"6a9e226fde5ea82090db6595","projectPage":"https://step-out.github.io/Motion-Omni-Page/","githubRepo":"https://github.com/step-out/Motion-Omni","githubRepoAddedBy":"user","ai_summary":"Motion-Omni is an end-to-end framework that jointly generates spoken dialogue and full-body co-speech motion from shared hidden states, using scalable pseudo-labeling and a unified evaluation protocol to achieve real-time, aligned responses.","ai_keywords":["spoken dialogue model","co-speech motion","Motion-Omni","LLM","Speech Generator","Motion Generator","pseudo-labeling","full-body motion","beat correlation","real-time factor"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"63fd9ec5ed9eead590ff216b","name":"PKU1898","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677565514018-61f8e5934a8e5a275b2b3e5a.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6355473d525beaee688b7ba1","avatarUrl":"/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg","isPro":false,"fullname":"Wei Tao","user":"itaowe","type":"user"},{"_id":"660383b2527470e0164533a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660383b2527470e0164533a9/CXIpr6_vtoxPFXW5EKh8n.jpeg","isPro":false,"fullname":"Chengqian Ma","user":"ChengqianMa","type":"user"},{"_id":"645307e8c895e6437dc8010b","avatarUrl":"/avatars/c8b18e1eea21b3618c488aaeb1a59377.svg","isPro":false,"fullname":"郝家诚","user":"Haojiacheng","type":"user"},{"_id":"67d27fb95785e2093c553182","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/P0JWb5xIVV4Az-1mvXNcA.png","isPro":false,"fullname":"ZianHuang","user":"ZianHuang","type":"user"},{"_id":"66b064b1e48856bb71a1feff","avatarUrl":"/avatars/5f9b321267ea7353948b650c0e9d61a8.svg","isPro":false,"fullname":"zhangansen","user":"zhangansen","type":"user"},{"_id":"654bcc8686cad6da97acbb48","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/654bcc8686cad6da97acbb48/sG1q7dcOS1MrrVUup_Fb6.jpeg","isPro":false,"fullname":"Xiangyu Zhao","user":"xyzhaocs","type":"user"},{"_id":"656c3503e0ff1cebe95e5730","avatarUrl":"/avatars/e9e6ecde1ca79833bfe23df51ee0dc68.svg","isPro":false,"fullname":"Tsuki","user":"Tsukihjy","type":"user"},{"_id":"65dd503d56cdc976b31a6ea2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/0Y-04s1sH2aRF-qrP1x6f.png","isPro":false,"fullname":"Reina","user":"reinaqwq","type":"user"},{"_id":"67825a33f0590f8a4671c869","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/etgxuroBooQTJhaGthcOR.png","isPro":false,"fullname":"jason","user":"hubertjason","type":"user"},{"_id":"663b8bc4eb73b1a397f056ae","avatarUrl":"/avatars/7dc79fe553ec5e4b2659f3c0fd09ce29.svg","isPro":false,"fullname":"tiez","user":"tiez","type":"user"},{"_id":"68051cfedb92e32b7068bee7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/vxO-1hFj_y0fRO4X1_rc0.png","isPro":false,"fullname":"Liu","user":"phenanthra","type":"user"},{"_id":"6314518e5f47a1896274d080","avatarUrl":"/avatars/d69ad2d8b14a87961603b29d1ef2eba6.svg","isPro":true,"fullname":"Zhiyuan Peng","user":"pzy2000","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"63fd9ec5ed9eead590ff216b","name":"PKU1898","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677565514018-61f8e5934a8e5a275b2b3e5a.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04250.md","query":{}}">
Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
Abstract
Motion-Omni is an end-to-end framework that jointly generates spoken dialogue and full-body co-speech motion from shared hidden states, using scalable pseudo-labeling and a unified evaluation protocol to achieve real-time, aligned responses.
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. We present Motion-Omni, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech. Joint training is not optional here: with the speech pathway frozen, motion remains misaligned with the audio, and co-adapting the LLM, Speech Generator and Motion Generator under both objectives is what recovers alignment while retaining spoken-dialogue ability. Supervision comes from a scalable, model-agnostic pipeline that pseudo-labels consistent-voice speech responses with a replaceable motion teacher, yielding 422,856 quality-ranked pairs (1,402 hours). We further release SwDA-500 and, to our knowledge, the first public evaluation protocol for stochastic open-ended full-body spoken dialogue, matching audio across motion systems while unifying rendering, automatic metrics, human evaluation, and latency measurement. Instantiated with a Qwen2.5-7B-Instruct backbone, Motion-Omni-Q7 matches the same-audio teacher cascade to within 2% on reference-free motion metrics while responding 5.4 x faster (RTF=0.78, faster than real time), surpasses all non-teacher cascades on beat correlation and diversity, and reaches a 2.62% word error rate, the lowest among the omni-modal systems compared.
Community
Motion-Omni replaces the speech-then-motion cascade with one end-to-end model that natively generates dialogue speech together with explicit facial expression, hand, upper-body, and lower-body motion, all from the hidden states that produce the speech. It stays within 2% of the teacher cascade on reference-free motion metrics, runs 5.4× faster at RTF 0.78 (faster than real time), and reaches a 2.62% WER, the lowest among the omni-modal systems compared.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.04250 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.04250 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.04250 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.