VisualPatchWorld (VPW) learns programmatic dynamics from LeWM expert trajectories on four environments (Two-room, Reacher, PushT, Cube), then evaluates them with fair frozen CEM / library-shoot planners.</p>\n","updatedAt":"2026-07-29T09:23:35.029Z","author":{"_id":"64670b0e0ed2f7a8cba98aa3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64670b0e0ed2f7a8cba98aa3/mfZlx51fLVRoeC0RyjOvd.jpeg","fullname":"Jiaxin","name":"jbai0318","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8231328129768372},"editors":["jbai0318"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64670b0e0ed2f7a8cba98aa3/mfZlx51fLVRoeC0RyjOvd.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.25236","authors":[{"_id":"6a69c29e9d3a1231d492b9fb","name":"Jiaxin Bai","hidden":false},{"_id":"6a69c29e9d3a1231d492b9fc","name":"Jiaxuan Xiong","hidden":false}],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-29T00:00:00.000Z","title":"VisualPatchWorld: Code World Models as Latent Structured Representations for Planning","submittedOnDailyBy":{"_id":"64670b0e0ed2f7a8cba98aa3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64670b0e0ed2f7a8cba98aa3/mfZlx51fLVRoeC0RyjOvd.jpeg","isPro":false,"fullname":"Jiaxin","user":"jbai0318","type":"user","name":"jbai0318"},"summary":"Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception, simulation, and planning. Two prominent realizations are neural predictors that learn dynamics in continuous vector spaces, and hand-built physics engines that expose explicit state and physical laws. Neural predictors scale from data but leave the form of the dynamics implicit; physics engines are inspectable and editable but difficult to construct at scale. We introduce VisualPatchWorld (VPW), which represents world dynamics as code. VPW first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error. The resulting programs can be rolled forward like a simulator, inspected in source form, and used inside model-predictive control; image-derived scene graphs can supply the live state at replan time. Across comparisons with prior code-based world models, VPW attains 69.0% mean planning success and exceeds the strongest code baseline by 23.5 points. The largest gains arise when choosing the correct qualitative dynamics is essential. Under the same planner, the induced models approach ground-truth engine success on navigation and grasp-rich control; a residual gap remains for contact-rich pushing, and checking a shortlist of promising plans in the engine closes most of that gap. These results establish a practical route toward automatically constructed code world models that are useful for planning. Code is available at https://github.com/HKBU-KnowComp/VisualPatchWorld/.","upvotes":1,"discussionId":"6a69c29e9d3a1231d492b9fd","githubRepo":"https://github.com/HKBU-KnowComp/VisualPatchWorld","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"6a17e7fd5cefe89a1409a23c","name":"HKBU-KnowComp","fullname":"HKBU Knowledge Computation Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64670b0e0ed2f7a8cba98aa3/yFMrUnXOVCS8Xi9ru0mdg.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64670b0e0ed2f7a8cba98aa3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64670b0e0ed2f7a8cba98aa3/mfZlx51fLVRoeC0RyjOvd.jpeg","isPro":false,"fullname":"Jiaxin","user":"jbai0318","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a17e7fd5cefe89a1409a23c","name":"HKBU-KnowComp","fullname":"HKBU Knowledge Computation Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64670b0e0ed2f7a8cba98aa3/yFMrUnXOVCS8Xi9ru0mdg.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.25236.md","query":{}}">
VisualPatchWorld: Code World Models as Latent Structured Representations for Planning
Published on Jul 28
· Submitted by Jiaxin on Jul 29 Abstract
Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception, simulation, and planning. Two prominent realizations are neural predictors that learn dynamics in continuous vector spaces, and hand-built physics engines that expose explicit state and physical laws. Neural predictors scale from data but leave the form of the dynamics implicit; physics engines are inspectable and editable but difficult to construct at scale. We introduce VisualPatchWorld (VPW), which represents world dynamics as code. VPW first selects a qualitative dynamical form with short active probes, then fits that form's free parameters from recorded state-action traces by minimizing multi-step prediction error. The resulting programs can be rolled forward like a simulator, inspected in source form, and used inside model-predictive control; image-derived scene graphs can supply the live state at replan time. Across comparisons with prior code-based world models, VPW attains 69.0% mean planning success and exceeds the strongest code baseline by 23.5 points. The largest gains arise when choosing the correct qualitative dynamics is essential. Under the same planner, the induced models approach ground-truth engine success on navigation and grasp-rich control; a residual gap remains for contact-rich pushing, and checking a shortlist of promising plans in the engine closes most of that gap. These results establish a practical route toward automatically constructed code world models that are useful for planning. Code is available at https://github.com/HKBU-KnowComp/VisualPatchWorld/.
Community
VisualPatchWorld (VPW) learns programmatic dynamics from LeWM expert trajectories on four environments (Two-room, Reacher, PushT, Cube), then evaluates them with fair frozen CEM / library-shoot planners.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.25236 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.25236 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.25236 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.