Hugging Face Daily Papers · · 4 min read

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

\n<li>Paper: <a href=\"https://arxiv.org/abs/2608.09926\" rel=\"nofollow\">https://arxiv.org/abs/2608.09926</a></li>\n<li>Page: <a href=\"https://lat-dyn-reason.github.io/\" rel=\"nofollow\">https://lat-dyn-reason.github.io/</a></li>\n<li>Code: <a href=\"https://github.com/Lat-Dyn-Reason/Lat-Dyn-Reason\" rel=\"nofollow\">https://github.com/Lat-Dyn-Reason/Lat-Dyn-Reason</a></li>\n<li>Model: <a href=\"https://huggingface.co/haodongli/LDR\">https://huggingface.co/haodongli/LDR</a></li>\n<li>Data: <a href=\"https://huggingface.co/datasets/haodongli/LDR\">https://huggingface.co/datasets/haodongli/LDR</a></li>\n</ul>\n","updatedAt":"2026-08-13T18:40:40.527Z","author":{"_id":"641d211e353524fe41f16387","avatarUrl":"/avatars/01d8fec857faac05f15b772f65127565.svg","fullname":"Haodong Li","name":"haodongli","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":38,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5088064670562744},"editors":["haodongli"],"editorAvatarUrls":["/avatars/01d8fec857faac05f15b772f65127565.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09926","authors":[{"_id":"6a7bf01a1653ef87c6af1cd8","name":"Haodong Li","hidden":false},{"_id":"6a7bf01a1653ef87c6af1cd9","name":"Shaoteng Liu","hidden":false},{"_id":"6a7bf01a1653ef87c6af1cda","name":"Tianyu Wang","hidden":false},{"_id":"6a7bf01a1653ef87c6af1cdb","name":"Chongjian Ge","hidden":false},{"_id":"6a7bf01a1653ef87c6af1cdc","name":"Sihui Ji","hidden":false},{"_id":"6a7bf01a1653ef87c6af1cdd","name":"Jiahan Zhang","hidden":false},{"_id":"6a7bf01a1653ef87c6af1cde","name":"Xin Lin","hidden":false},{"_id":"6a7bf01a1653ef87c6af1cdf","name":"Haolin Lu","hidden":false},{"_id":"6a7bf01a1653ef87c6af1ce0","name":"Zhe Lin","hidden":false},{"_id":"6a7bf01a1653ef87c6af1ce1","name":"Manmohan Chandraker","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/641d211e353524fe41f16387/aynXWfCWS7Ouu9qLUnC0X.mp4"],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning","submittedOnDailyBy":{"_id":"641d211e353524fe41f16387","avatarUrl":"/avatars/01d8fec857faac05f15b772f65127565.svg","isPro":false,"fullname":"Haodong Li","user":"haodongli","type":"user","name":"haodongli"},"summary":"The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from pixels, we introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features. Following PhyWorld, we validate LDR on a controlled white-box physics benchmark spanning five tasks (uniform motion, parabola, collision, bouncing, looming), focusing on out-of-distribution scenarios that reveal whether a model has truly learned the underlying dynamics. LDR extrapolates the learned dynamics far better: the gap between its in- and out-of-distribution error is over 20times smaller than the video diffusion baseline's, under both single- and joint-task training at 256^2 resolution, while using 26times fewer parameters and running 143times faster. LDR can even generalize under severe shift: for example, trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left. To our knowledge, this is the first video world model that extrapolates learned dynamics beyond its training distribution. Project page: https://lat-dyn-reason.github.io/","upvotes":3,"discussionId":"6a7bf01a1653ef87c6af1ce2","projectPage":"https://lat-dyn-reason.github.io/","githubRepo":"https://github.com/Lat-Dyn-Reason/Lat-Dyn-Reason","githubRepoAddedBy":"user","ai_summary":"Latent Dynamics Reasoning integrates kinematic dynamics in structured latent space to enable video world models that extrapolate physical laws far beyond training distributions with far fewer parameters and faster inference.","ai_keywords":["Latent Dynamics Reasoning","kinematic integration","structured latent","video diffusion","out-of-distribution extrapolation","physics benchmark","world model"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":28,"organization":{"_id":"61e5d14f77496de0a6d95c6b","name":"adobe","fullname":"Adobe","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1645217431826-61e35e517ac6b6d06cfa8081.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"641d211e353524fe41f16387","avatarUrl":"/avatars/01d8fec857faac05f15b772f65127565.svg","isPro":false,"fullname":"Haodong Li","user":"haodongli","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"6644b9853f0604b318b7bf80","avatarUrl":"/avatars/0c3d89117489dca7fb203c334c3d6d7c.svg","isPro":false,"fullname":" X","user":"AnonymityX","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61e5d14f77496de0a6d95c6b","name":"adobe","fullname":"Adobe","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1645217431826-61e35e517ac6b6d06cfa8081.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09926.md","query":{}}">
Papers
arxiv:2608.09926

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Published on Aug 10
· Submitted by
Haodong Li
on Aug 13
Authors:
,

Abstract

Latent Dynamics Reasoning integrates kinematic dynamics in structured latent space to enable video world models that extrapolate physical laws far beyond training distributions with far fewer parameters and faster inference.

The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from pixels, we introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features. Following PhyWorld, we validate LDR on a controlled white-box physics benchmark spanning five tasks (uniform motion, parabola, collision, bouncing, looming), focusing on out-of-distribution scenarios that reveal whether a model has truly learned the underlying dynamics. LDR extrapolates the learned dynamics far better: the gap between its in- and out-of-distribution error is over 20times smaller than the video diffusion baseline's, under both single- and joint-task training at 256^2 resolution, while using 26times fewer parameters and running 143times faster. LDR can even generalize under severe shift: for example, trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left. To our knowledge, this is the first video world model that extrapolates learned dynamics beyond its training distribution. Project page: https://lat-dyn-reason.github.io/

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.09926
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers