Hugging Face Daily Papers · · 4 min read

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🧠 VideoCoCo uses executable Blender code as process-level CoT✨, generating deterministic spatiotemporal drafts that guide a video editor toward photorealistic and physically consistent results. 🎬</p>\n","updatedAt":"2026-07-31T02:07:39.734Z","author":{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","fullname":"Lin Chen","name":"Lin-Chen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":99,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7859587073326111},"editors":["Lin-Chen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.27380","authors":[{"_id":"6a6c03567bd25d8874c070d9","name":"Haodong Li","hidden":false},{"_id":"6a6c03567bd25d8874c070da","user":{"_id":"69a04978be4e0dfbcf999758","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69a04978be4e0dfbcf999758/uFAYXLPMj21t4IVnF_exk.jpeg","isPro":false,"fullname":"Tianfei Ren","user":"rentianfei122","type":"user","name":"rentianfei122"},"name":"Tianfei Ren","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.449Z","hidden":false},{"_id":"6a6c03567bd25d8874c070db","name":"Xiaoxiao Ma","hidden":false},{"_id":"6a6c03567bd25d8874c070dc","name":"Chunmei Qing","hidden":false},{"_id":"6a6c03567bd25d8874c070dd","name":"Zhen Fang","hidden":false},{"_id":"6a6c03567bd25d8874c070de","name":"Sipeng He","hidden":false},{"_id":"6a6c03567bd25d8874c070df","name":"Ziyu Guo","hidden":false},{"_id":"6a6c03567bd25d8874c070e0","name":"Haoyu Wu","hidden":false},{"_id":"6a6c03567bd25d8874c070e1","user":{"_id":"670880950e79a8b46f7ff9dd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670880950e79a8b46f7ff9dd/hA1TLhwlQblkFsq8wLrkB.jpeg","isPro":false,"fullname":"Juanxi Tian","user":"Juanxi","type":"user","name":"Juanxi"},"name":"Juanxi Tian","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.457Z","hidden":false},{"_id":"6a6c03567bd25d8874c070e2","name":"Yihang Zou","hidden":false},{"_id":"6a6c03567bd25d8874c070e3","name":"Ruichuan An","hidden":false},{"_id":"6a6c03567bd25d8874c070e4","name":"Dongzhi Jiang","hidden":false},{"_id":"6a6c03567bd25d8874c070e5","name":"Boxue Yang","hidden":false},{"_id":"6a6c03567bd25d8874c070e6","name":"Ji Xie","hidden":false},{"_id":"6a6c03567bd25d8874c070e7","name":"Xu Huang","hidden":false},{"_id":"6a6c03567bd25d8874c070e8","name":"Wenhao Yan","hidden":false},{"_id":"6a6c03567bd25d8874c070e9","name":"Jialv Zou","hidden":false},{"_id":"6a6c03567bd25d8874c070ea","name":"Zhengrong Yue","hidden":false},{"_id":"6a6c03567bd25d8874c070eb","name":"Yaxin Luo","hidden":false},{"_id":"6a6c03567bd25d8874c070ec","name":"Xiaotong Li","hidden":false},{"_id":"6a6c03567bd25d8874c070ed","name":"Yuzhu Wang","hidden":false},{"_id":"6a6c03567bd25d8874c070ee","name":"Junyan Ye","hidden":false},{"_id":"6a6c03567bd25d8874c070ef","name":"Jinjing Zhao","hidden":false},{"_id":"6a6c03567bd25d8874c070f0","name":"Zehui Chen","hidden":false},{"_id":"6a6c03567bd25d8874c070f1","user":{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","isPro":false,"fullname":"Lin Chen","user":"Lin-Chen","type":"user","name":"Lin-Chen"},"name":"Lin Chen","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.440Z","hidden":false},{"_id":"6a6c03567bd25d8874c070f2","name":"Renye Yan","hidden":false},{"_id":"6a6c03567bd25d8874c070f3","name":"Feng Zhao","hidden":false},{"_id":"6a6c03567bd25d8874c070f4","name":"Pheng-Ann Heng","hidden":false}],"publishedAt":"2026-07-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System","submittedOnDailyBy":{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","isPro":false,"fullname":"Lin Chen","user":"Lin-Chen","type":"user","name":"Lin-Chen"},"summary":"Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.","upvotes":60,"discussionId":"6a6c03567bd25d8874c070f5","githubRepo":"https://github.com/micky-li-hd/VideoCoCo","githubRepoAddedBy":"user","githubStars":6},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64b02ec0e5000ae8a572ced5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b02ec0e5000ae8a572ced5/6ifLntBU2ICQK7SW8WxKU.png","isPro":false,"fullname":"Lin Chen","user":"Lin-Chen","type":"user"},{"_id":"69a04978be4e0dfbcf999758","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69a04978be4e0dfbcf999758/uFAYXLPMj21t4IVnF_exk.jpeg","isPro":false,"fullname":"Tianfei Ren","user":"rentianfei122","type":"user"},{"_id":"64b0a5037a475fba70a7260d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b0a5037a475fba70a7260d/MauBbb6raMA23yrR1Zq21.jpeg","isPro":false,"fullname":"Zhen Fang","user":"CostaliyA","type":"user"},{"_id":"64892d31cbda0d1cdb956897","avatarUrl":"/avatars/3cdafe03a8295124636347d15a099aaf.svg","isPro":false,"fullname":"Zehui Chen","user":"lovesnowbest","type":"user"},{"_id":"681c90c5866cdb5f8216078d","avatarUrl":"/avatars/ebca826c6b622c8f4912ff936afb762b.svg","isPro":false,"fullname":"hartleychen","user":"rayrayhartley","type":"user"},{"_id":"68106c88b924dd6c328889c2","avatarUrl":"/avatars/8accf835b711bffa2ea307158950ab33.svg","isPro":false,"fullname":"Hongbo Peng","user":"M1chaelPeng","type":"user"},{"_id":"67543820c3af453d7b3e1d5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67543820c3af453d7b3e1d5e/RbAZ9AQlxpy5E5Is-QN8b.jpeg","isPro":false,"fullname":"Dingming Li","user":"lidingm","type":"user"},{"_id":"665d785d2080b38e5d74a742","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/665d785d2080b38e5d74a742/hC38hs5EEHXyqHHm8VXhg.png","isPro":false,"fullname":"Junxian Mu","user":"xxz9","type":"user"},{"_id":"670880950e79a8b46f7ff9dd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670880950e79a8b46f7ff9dd/hA1TLhwlQblkFsq8wLrkB.jpeg","isPro":false,"fullname":"Juanxi Tian","user":"Juanxi","type":"user"},{"_id":"64f8962bce75bb0fb50bdbdb","avatarUrl":"/avatars/c85537df848bda7ec92565f56cd32eed.svg","isPro":false,"fullname":"Jinjing Zhao","user":"Jinjing713","type":"user"},{"_id":"6349214f8146350b3a4c5cdf","avatarUrl":"/avatars/cfd24caac9a87efb528d0f4c375932bc.svg","isPro":false,"fullname":"Dongzhi Jiang","user":"CaraJ","type":"user"},{"_id":"6590f7880c993129053a2344","avatarUrl":"/avatars/5cf5c5185d10feeaecd5e8add7c4330d.svg","isPro":false,"fullname":"Haoyu wu","user":"Haoyuwu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.27380.md","query":{}}">
Papers
arxiv:2607.27380

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Published on Jul 29
· Submitted by
Lin Chen
on Jul 31
Authors:
,

Abstract

Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.

Community

Paper author Paper submitter about 8 hours ago

🧠 VideoCoCo uses executable Blender code as process-level CoT✨, generating deterministic spatiotemporal drafts that guide a video editor toward photorealistic and physically consistent results. 🎬

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.27380
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.27380 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.27380 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.27380 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers