Hugging Face Daily Papers · · 3 min read

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

What if an omni-modal dialogue model could not only listen and speak, but also appear? Ex-Omni-2D generates coordinated text, personalized speech, and expressive avatar video within a unified dialogue framework. We would love to hear your thoughts on visual presence, streaming generation, and the future of embodied dialogue systems.</p>\n","updatedAt":"2026-08-12T03:12:42.165Z","author":{"_id":"665d72007bef1cfc313a92dd","avatarUrl":"/avatars/6d56671153bbf1ffff072472678819da.svg","fullname":"Haoyu Zhang","name":"lemonade666","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5dd96eb166059660ed1ee413/Lp3m-XLpjQGwBItlvn69q.png","fullname":"Tencent","name":"tencent","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9222422242164612},"editors":["lemonade666"],"editorAvatarUrls":["/avatars/6d56671153bbf1ffff072472678819da.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.10720","authors":[{"_id":"6a7be3a21653ef87c6af1c2f","name":"Haoyu Zhang","hidden":false},{"_id":"6a7be3a21653ef87c6af1c30","name":"Zhipeng Li","hidden":false},{"_id":"6a7be3a21653ef87c6af1c31","name":"Xiaoying Tang","hidden":false},{"_id":"6a7be3a21653ef87c6af1c32","name":"Tianshu Yu","hidden":false},{"_id":"6a7be3a21653ef87c6af1c33","name":"Yiwen Guo","hidden":false}],"publishedAt":"2026-08-11T00:00:00.000Z","submittedOnDailyAt":"2026-08-12T00:00:00.000Z","title":"Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence","submittedOnDailyBy":{"_id":"665d72007bef1cfc313a92dd","avatarUrl":"/avatars/6d56671153bbf1ffff072472678819da.svg","isPro":true,"fullname":"Haoyu Zhang","user":"lemonade666","type":"user","name":"lemonade666"},"summary":"Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.","upvotes":9,"discussionId":"6a7be3a21653ef87c6af1c34","projectPage":"https://logo-cuhksz.github.io/Ex-Omni-2D","githubRepo":"https://github.com/LOGO-CUHKSZ/Ex-Omni-2D-Code","githubRepoAddedBy":"user","ai_summary":"Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator.","ai_keywords":["omni-modal dialogue","Visual Thought Plan","multi-codebook speech units","block-causal Streaming Student","Prefix Streaming","Video Generator","RTF"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":16},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65164444bc0631719873af81","avatarUrl":"/avatars/0e68ea5b5369273a07e5889480ca9421.svg","isPro":false,"fullname":"Wei Pang","user":"weipang142857","type":"user"},{"_id":"665d72007bef1cfc313a92dd","avatarUrl":"/avatars/6d56671153bbf1ffff072472678819da.svg","isPro":true,"fullname":"Haoyu Zhang","user":"lemonade666","type":"user"},{"_id":"673d927a3af47d1d2b99b090","avatarUrl":"/avatars/18bf08b95d09fc3c12704619540b8bbb.svg","isPro":false,"fullname":"Daoyuan Zheng","user":"zdyzdyzdy","type":"user"},{"_id":"68380f4f231cf484dd4e87f4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/xp34hfiSLf-DiE1DVVhHk.png","isPro":false,"fullname":"Xinjian Zhao","user":"Xinjiansz","type":"user"},{"_id":"6a1595088e1e41b41f23e647","avatarUrl":"/avatars/df6c69d98af331f835ed5c76282af586.svg","isPro":false,"fullname":"가은 오","user":"olivertaylor8","type":"user"},{"_id":"6355473d525beaee688b7ba1","avatarUrl":"/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg","isPro":false,"fullname":"Wei Tao","user":"itaowe","type":"user"},{"_id":"660383b2527470e0164533a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660383b2527470e0164533a9/CXIpr6_vtoxPFXW5EKh8n.jpeg","isPro":false,"fullname":"Chengqian Ma","user":"ChengqianMa","type":"user"},{"_id":"67f38ee467c6c8a4ae8445b2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JFpCVX_uC2GtvqNO-h_Tm.png","isPro":false,"fullname":"lalala","user":"adlalala","type":"user"},{"_id":"65672c089450460026f602eb","avatarUrl":"/avatars/5064d4849c6af7677da937343e37cd11.svg","isPro":false,"fullname":"ruiying LIU","user":"sholyu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Papers
arxiv:2608.10720

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Published on Aug 11
· Submitted by
Haoyu Zhang
on Aug 12
Authors:
,

Abstract

Ex-Omni-2D is an omni-modal dialogue framework that produces coordinated text, speech, and video responses via a visual thought plan and a distilled streaming video generator.

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. We introduce Ex-Omni-2D, an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query--text--speech--video supervision. A full-sequence Video Generator serves as the primary Teacher. For efficient incremental generation, we further distill it into a few-step block-causal Streaming Student whose Prefix Streaming mechanism carries a clean latent across consecutive chunks to reduce cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400times720/720times400, providing a practical quality--efficiency operating point.

Community

What if an omni-modal dialogue model could not only listen and speak, but also appear? Ex-Omni-2D generates coordinated text, personalized speech, and expressive avatar video within a unified dialogue framework. We would love to hear your thoughts on visual presence, streaming generation, and the future of embodied dialogue systems.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.10720 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.10720 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.10720 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers