🧠 VideoCoCo uses executable Blender code as process-level CoT✨, generating deterministic spatiotemporal drafts that guide a video editor toward photorealistic and physically consistent results. 🎬</p>\n","updatedAt":"2026-07-31T02:07:39.734Z","author":{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","fullname":"Lin Chen","name":"Lin-Chen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":99,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7859587073326111},"editors":["Lin-Chen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27380","authors":[{"_id":"6a6c03567bd25d8874c070d9","name":"Haodong Li","hidden":false},{"_id":"6a6c03567bd25d8874c070da","user":{"_id":"69a04978be4e0dfbcf999758","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69a04978be4e0dfbcf999758/uFAYXLPMj21t4IVnF_exk.jpeg","isPro":false,"fullname":"Tianfei Ren","user":"rentianfei122","type":"user","name":"rentianfei122"},"name":"Tianfei Ren","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.449Z","hidden":false},{"_id":"6a6c03567bd25d8874c070db","name":"Xiaoxiao Ma","hidden":false},{"_id":"6a6c03567bd25d8874c070dc","name":"Chunmei Qing","hidden":false},{"_id":"6a6c03567bd25d8874c070dd","name":"Zhen Fang","hidden":false},{"_id":"6a6c03567bd25d8874c070de","name":"Sipeng He","hidden":false},{"_id":"6a6c03567bd25d8874c070df","name":"Ziyu Guo","hidden":false},{"_id":"6a6c03567bd25d8874c070e0","name":"Haoyu Wu","hidden":false},{"_id":"6a6c03567bd25d8874c070e1","user":{"_id":"670880950e79a8b46f7ff9dd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670880950e79a8b46f7ff9dd/hA1TLhwlQblkFsq8wLrkB.jpeg","isPro":false,"fullname":"Juanxi Tian","user":"Juanxi","type":"user","name":"Juanxi"},"name":"Juanxi Tian","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.457Z","hidden":false},{"_id":"6a6c03567bd25d8874c070e2","name":"Yihang Zou","hidden":false},{"_id":"6a6c03567bd25d8874c070e3","name":"Ruichuan An","hidden":false},{"_id":"6a6c03567bd25d8874c070e4","name":"Dongzhi Jiang","hidden":false},{"_id":"6a6c03567bd25d8874c070e5","name":"Boxue Yang","hidden":false},{"_id":"6a6c03567bd25d8874c070e6","name":"Ji Xie","hidden":false},{"_id":"6a6c03567bd25d8874c070e7","name":"Xu Huang","hidden":false},{"_id":"6a6c03567bd25d8874c070e8","name":"Wenhao Yan","hidden":false},{"_id":"6a6c03567bd25d8874c070e9","name":"Jialv Zou","hidden":false},{"_id":"6a6c03567bd25d8874c070ea","name":"Zhengrong Yue","hidden":false},{"_id":"6a6c03567bd25d8874c070eb","name":"Yaxin Luo","hidden":false},{"_id":"6a6c03567bd25d8874c070ec","name":"Xiaotong Li","hidden":false},{"_id":"6a6c03567bd25d8874c070ed","name":"Yuzhu Wang","hidden":false},{"_id":"6a6c03567bd25d8874c070ee","name":"Junyan Ye","hidden":false},{"_id":"6a6c03567bd25d8874c070ef","name":"Jinjing Zhao","hidden":false},{"_id":"6a6c03567bd25d8874c070f0","name":"Zehui Chen","hidden":false},{"_id":"6a6c03567bd25d8874c070f1","user":{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","isPro":false,"fullname":"Lin Chen","user":"Lin-Chen","type":"user","name":"Lin-Chen"},"name":"Lin Chen","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.440Z","hidden":false},{"_id":"6a6c03567bd25d8874c070f2","name":"Renye Yan","hidden":false},{"_id":"6a6c03567bd25d8874c070f3","name":"Feng Zhao","hidden":false},{"_id":"6a6c03567bd25d8874c070f4","name":"Pheng-Ann Heng","hidden":false}],"publishedAt":"2026-07-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System","submittedOnDailyBy":{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","isPro":false,"fullname":"Lin Chen","user":"Lin-Chen","type":"user","name":"Lin-Chen"},"summary":"Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.","upvotes":60,"discussionId":"6a6c03567bd25d8874c070f5","githubRepo":"https://github.com/micky-li-hd/VideoCoCo","githubRepoAddedBy":"user","githubStars":6},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","isPro":false,"fullname":"Lin Chen","user":"Lin-Chen","type":"user"},{"_id":"69a04978be4e0dfbcf999758","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69a04978be4e0dfbcf999758/uFAYXLPMj21t4IVnF_exk.jpeg","isPro":false,"fullname":"Tianfei Ren","user":"rentianfei122","type":"user"},{"_id":"64b0a5037a475fba70a7260d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b0a5037a475fba70a7260d/MauBbb6raMA23yrR1Zq21.jpeg","isPro":false,"fullname":"Zhen Fang","user":"CostaliyA","type":"user"},{"_id":"64892d31cbda0d1cdb956897","avatarUrl":"/avatars/3cdafe03a8295124636347d15a099aaf.svg","isPro":false,"fullname":"Zehui Chen","user":"lovesnowbest","type":"user"},{"_id":"681c90c5866cdb5f8216078d","avatarUrl":"/avatars/ebca826c6b622c8f4912ff936afb762b.svg","isPro":false,"fullname":"hartleychen","user":"rayrayhartley","type":"user"},{"_id":"68106c88b924dd6c328889c2","avatarUrl":"/avatars/8accf835b711bffa2ea307158950ab33.svg","isPro":false,"fullname":"Hongbo Peng","user":"M1chaelPeng","type":"user"},{"_id":"67543820c3af453d7b3e1d5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67543820c3af453d7b3e1d5e/RbAZ9AQlxpy5E5Is-QN8b.jpeg","isPro":false,"fullname":"Dingming Li","user":"lidingm","type":"user"},{"_id":"665d785d2080b38e5d74a742","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/665d785d2080b38e5d74a742/hC38hs5EEHXyqHHm8VXhg.png","isPro":false,"fullname":"Junxian Mu","user":"xxz9","type":"user"},{"_id":"670880950e79a8b46f7ff9dd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670880950e79a8b46f7ff9dd/hA1TLhwlQblkFsq8wLrkB.jpeg","isPro":false,"fullname":"Juanxi Tian","user":"Juanxi","type":"user"},{"_id":"64f8962bce75bb0fb50bdbdb","avatarUrl":"/avatars/c85537df848bda7ec92565f56cd32eed.svg","isPro":false,"fullname":"Jinjing Zhao","user":"Jinjing713","type":"user"},{"_id":"6349214f8146350b3a4c5cdf","avatarUrl":"/avatars/cfd24caac9a87efb528d0f4c375932bc.svg","isPro":false,"fullname":"Dongzhi Jiang","user":"CaraJ","type":"user"},{"_id":"6590f7880c993129053a2344","avatarUrl":"/avatars/5cf5c5185d10feeaecd5e8add7c4330d.svg","isPro":false,"fullname":"Haoyu wu","user":"Haoyuwu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27380.md","query":{}}">
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Abstract
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
Community
🧠 VideoCoCo uses executable Blender code as process-level CoT✨, generating deterministic spatiotemporal drafts that guide a video editor toward photorealistic and physically consistent results. 🎬
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.27380 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.27380 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.27380 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.