What if an omni-modal dialogue model could not only listen and speak, but also appear? Ex-Omni-2D generates coordinated text, personalized speech, and expressive avatar video within a unified dialogue framework. We would love to hear your thoughts on visual presence, streaming generation, and the future of embodied dialogue systems.</p>\n","updatedAt":"2026-08-12T03:12:42.165Z","author":{"_id":"665d72007bef1cfc313a92dd","avatarUrl":"/avatars/6d56671153bbf1ffff072472678819da.svg","fullname":"Haoyu Zhang","name":"lemonade666","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png","fullname":"Tencent","name":"tencent","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9222422242164612},"editors":["lemonade666"],"editorAvatarUrls":["/avatars/6d56671153bbf1ffff072472678819da.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.10720","authors":[{"_id":"6a7be3a21653ef87c6af1c2f","name":"Haoyu Zhang","hidden":false},{"_id":"6a7be3a21653ef87c6af1c30","name":"Zhipeng Li","hidden":false},{"_id":"6a7be3a21653ef87c6af1c31","name":"Xiaoying Tang","hidden":false},{"_id":"6a7be3a21653ef87c6af1c32","name":"Tianshu Yu","hidden":false},{"_id":"6a7be3a21653ef87c6af1c33","name":"Yiwen Guo","hidden":false}],"publishedAt":"2026-08-11T00:00:00.000Z","submittedOnDailyAt":"2026-08-12T00:00:00.000Z","title":"Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence","submittedOnDailyBy":{"_id":"665d72007bef1cfc313a92dd","avatarUrl":"/avatars/6d56671153bbf1ffff072472678819da.svg","isPro":true,"fullname":"Haoyu Zhang","user":"lemonade666","type":"user","name":"lemonade666"},"summary":"Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.","upvotes":9,"discussionId":"6a7be3a21653ef87c6af1c34","projectPage":"https://logo-cuhksz.github.io/Ex-Omni-2D","githubRepo":"https://github.com/LOGO-CUHKSZ/Ex-Omni-2D-Code","githubRepoAddedBy":"user","ai_summary":"Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator.","ai_keywords":["omni-modal dialogue","Visual Thought Plan","multi-codebook speech units","block-causal Streaming Student","Prefix Streaming","Video Generator","RTF"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":16},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65164444bc0631719873af81","avatarUrl":"/avatars/0e68ea5b5369273a07e5889480ca9421.svg","isPro":false,"fullname":"Wei Pang","user":"weipang142857","type":"user"},{"_id":"665d72007bef1cfc313a92dd","avatarUrl":"/avatars/6d56671153bbf1ffff072472678819da.svg","isPro":true,"fullname":"Haoyu Zhang","user":"lemonade666","type":"user"},{"_id":"673d927a3af47d1d2b99b090","avatarUrl":"/avatars/18bf08b95d09fc3c12704619540b8bbb.svg","isPro":false,"fullname":"Daoyuan Zheng","user":"zdyzdyzdy","type":"user"},{"_id":"68380f4f231cf484dd4e87f4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/xp34hfiSLf-DiE1DVVhHk.png","isPro":false,"fullname":"Xinjian Zhao","user":"Xinjiansz","type":"user"},{"_id":"6a1595088e1e41b41f23e647","avatarUrl":"/avatars/df6c69d98af331f835ed5c76282af586.svg","isPro":false,"fullname":"가은 오","user":"olivertaylor8","type":"user"},{"_id":"6355473d525beaee688b7ba1","avatarUrl":"/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg","isPro":false,"fullname":"Wei Tao","user":"itaowe","type":"user"},{"_id":"660383b2527470e0164533a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660383b2527470e0164533a9/CXIpr6_vtoxPFXW5EKh8n.jpeg","isPro":false,"fullname":"Chengqian Ma","user":"ChengqianMa","type":"user"},{"_id":"67f38ee467c6c8a4ae8445b2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JFpCVX_uC2GtvqNO-h_Tm.png","isPro":false,"fullname":"lalala","user":"adlalala","type":"user"},{"_id":"65672c089450460026f602eb","avatarUrl":"/avatars/5064d4849c6af7677da937343e37cd11.svg","isPro":false,"fullname":"ruiying LIU","user":"sholyu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Abstract
Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator.
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.
Community
What if an omni-modal dialogue model could not only listen and speak, but also appear? Ex-Omni-2D generates coordinated text, personalized speech, and expressive avatar video within a unified dialogue framework. We would love to hear your thoughts on visual presence, streaming generation, and the future of embodied dialogue systems.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.10720 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.10720 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.10720 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.