full-duplex interaction system</p>\n","updatedAt":"2026-09-15T04:37:28.604Z","author":{"_id":"678f850bec882f210c1b59f2","avatarUrl":"/avatars/f25bab4afb86de92f75bf9a90e02f59f.svg","fullname":"XIE ZHIFEI","name":"zhifeixie","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":28,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7556428909301758},"editors":["zhifeixie"],"editorAvatarUrls":["/avatars/f25bab4afb86de92f75bf9a90e02f59f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.13814","authors":[{"_id":"6aa8c1eb5dd4cb9b4cc028b5","name":"Ruixiang Zhao","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028b6","name":"Hualei Wang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028b7","name":"Renhe Sun","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028b8","name":"Enzhi Zhou","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028b9","name":"Jincenzi Wu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028ba","name":"Xujie Song","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028bb","name":"Kexin Shi","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028bc","name":"Zihang Liu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028bd","name":"Pengcheng Zhu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028be","name":"Jiayi Zhou","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028bf","name":"Baoyue Zhang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c0","name":"Changhao Zhang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c1","name":"Zitong Wang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c2","name":"Jinhong Wang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c3","name":"Tong Niu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c4","name":"Jingjing Liu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c5","name":"Junan Lin","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c6","name":"Haolin He","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c7","name":"Hengshuo Chu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c8","name":"Yuhui Chen","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028c9","name":"Jian Liu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028ca","name":"Yuge Huang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028cb","name":"Junliang Xing","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028cc","name":"Yuntao Wang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028cd","name":"Weiqiang Wang","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028ce","name":"Chun Yu","hidden":false},{"_id":"6aa8c1eb5dd4cb9b4cc028cf","name":"Yuanchun Shi","hidden":false}],"publishedAt":"2026-09-12T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"Realtime-Venus: A full-duplex interaction system with asynchronous delegation","submittedOnDailyBy":{"_id":"678f850bec882f210c1b59f2","avatarUrl":"/avatars/f25bab4afb86de92f75bf9a90e02f59f.svg","isPro":false,"fullname":"XIE ZHIFEI","user":"zhifeixie","type":"user","name":"zhifeixie"},"summary":"Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events.\n A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue.\n Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows.\n Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.","upvotes":5,"discussionId":"6aa8c1ec5dd4cb9b4cc028d0","ai_summary":"Realtime-Venus is a proactive full-duplex system with separate audio-visual and audio models that integrate continuous perception, conversational control, and native speech generation via a shared causal timeline and dual-loop runtime.","ai_keywords":["full-duplex","causal timeline","dual-loop runtime","background reasoning","tool execution","post-training","proactive full-duplex trajectories","delegation workflows","StreamingBench","OVO-Bench","Daily-Omni","MMAU","MMAU-Pro","Speech CMMLU","VoiceBench AlpacaEval","Full-Duplex-Bench"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"67c1d682826160b28f778510","name":"antgroup","fullname":"Ant Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/7VcPHdLSGlged3ixK1dys.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64c377b85b7862ce873a929e","avatarUrl":"/avatars/2797d477c780b2f42f278aec9195aa2d.svg","isPro":false,"fullname":"ruixiang zhao","user":"RuixiangZhao","type":"user"},{"_id":"67c7d44419b236e0565358f4","avatarUrl":"/avatars/36f6dfd8c8dcdd69d8a9af5e58d978a4.svg","isPro":false,"fullname":"Zihang Liu","user":"zh-liu799","type":"user"},{"_id":"640f27dba92fedb0e84fe578","avatarUrl":"/avatars/7c81cadb2a6570e19feeaf2cab8d222c.svg","isPro":false,"fullname":"Wu","user":"Jincenzi","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6641dba91a96431427a006ce","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6641dba91a96431427a006ce/uKOGbZaxi0AYNJpCznLsn.jpeg","isPro":false,"fullname":"Zhiqiu Zhang","user":"ZZQ987","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67c1d682826160b28f778510","name":"antgroup","fullname":"Ant Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/7VcPHdLSGlged3ixK1dys.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.13814.md","query":{}}">
Realtime-Venus: A full-duplex interaction system with asynchronous delegation
Abstract
Realtime-Venus is a proactive full-duplex system with separate audio-visual and audio models that integrate continuous perception, conversational control, and native speech generation via a shared causal timeline and dual-loop runtime.
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events.
A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue.
Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows.
Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.
Community
full-duplex interaction system
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.13814 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.13814 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.13814 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.