Hugging Face Daily Papers · · 3 min read

Self-Supervised Learning of Structured Dynamics from Videos

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Check out our project page: <a href=\"https://lukasknobel.github.io/projects/StructuredDynamics\" rel=\"nofollow\">https://lukasknobel.github.io/projects/StructuredDynamics</a></p>\n","updatedAt":"2026-07-24T10:52:04.442Z","author":{"_id":"62f4cafb31ee3f3670f597a4","avatarUrl":"/avatars/9a214de17a8e6ab8ceb2b2ee8727368b.svg","fullname":"Lukas Knobel","name":"Lukas431","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.9014908075332642},"editors":["Lukas431"],"editorAvatarUrls":["/avatars/9a214de17a8e6ab8ceb2b2ee8727368b.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.21576","authors":[{"_id":"6a6337463f6711f3a0ba9395","user":{"_id":"62f4cafb31ee3f3670f597a4","avatarUrl":"/avatars/9a214de17a8e6ab8ceb2b2ee8727368b.svg","isPro":false,"fullname":"Lukas Knobel","user":"Lukas431","type":"user","name":"Lukas431"},"name":"Lukas Knobel","status":"claimed_verified","statusLastChangedAt":"2026-07-24T16:45:04.144Z","hidden":false},{"_id":"6a6337463f6711f3a0ba9396","name":"Andrew Zisserman","hidden":false},{"_id":"6a6337463f6711f3a0ba9397","name":"Yuki M. Asano","hidden":false}],"publishedAt":"2026-07-23T00:00:00.000Z","submittedOnDailyAt":"2026-07-24T00:00:00.000Z","title":"Self-Supervised Learning of Structured Dynamics from Videos","submittedOnDailyBy":{"_id":"62f4cafb31ee3f3670f597a4","avatarUrl":"/avatars/9a214de17a8e6ab8ceb2b2ee8727368b.svg","isPro":false,"fullname":"Lukas Knobel","user":"Lukas431","type":"user","name":"Lukas431"},"summary":"Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.","upvotes":13,"discussionId":"6a6337473f6711f3a0ba9398","projectPage":"https://lukasknobel.github.io/projects/StructuredDynamics","githubRepo":"https://github.com/lukasknobel/StructuredDynamics","githubRepoAddedBy":"user","githubStars":2,"organization":{"_id":"670635814d8c8b01343c35c5","name":"FunAILab","fullname":"Fundamental AI Lab at UTN","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637d21239a5217b88b7549c3/5rDBCg9OOamKK_UqCof_z.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69bcdac24df1e2c004bc1778","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/NZw9N0T0ZME6j5qBHA-RD.png","isPro":false,"fullname":"허 다은","user":"averyqg","type":"user"},{"_id":"660544b83d1c41c482d0cb05","avatarUrl":"/avatars/5258bd94730b8e7fcc5e95467fbcac7b.svg","isPro":false,"fullname":"Cameron Braunstein","user":"CameronBraunstein","type":"user"},{"_id":"651e947837233162ed764542","avatarUrl":"/avatars/32156abf958138e16b688e90148c26be.svg","isPro":false,"fullname":"Imanol G. Estepa","user":"St3p","type":"user"},{"_id":"66b1f7361b87a6bc9ddc4e5f","avatarUrl":"/avatars/766a8ebd999c31c9a99a182c29bd35a7.svg","isPro":false,"fullname":"Noor Ahmed","user":"noorahmedds","type":"user"},{"_id":"6690e4b4c7d72a1ae0de19cb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6690e4b4c7d72a1ae0de19cb/ixHTWIo9kQwSAHbT63eG2.jpeg","isPro":false,"fullname":"Dehghanighobadi ","user":"ZahraDehghanighobadi","type":"user"},{"_id":"637de0fa1342ba1762422495","avatarUrl":"/avatars/b1cce386a8b33007fc381fcbfc5cdc9a.svg","isPro":false,"fullname":"Dawid","user":"dakopi","type":"user"},{"_id":"649eaa2dfbdfd3c1612ee842","avatarUrl":"/avatars/64f255a15b0d39a81c4cfafa4de29fea.svg","isPro":false,"fullname":"Walter Simoncini","user":"walterdev","type":"user"},{"_id":"6870fa24e42263da14ea0bf2","avatarUrl":"/avatars/012eeb2a499f34a05be17ea4de286db5.svg","isPro":false,"fullname":"Raza Yunus","user":"razayunus","type":"user"},{"_id":"65ca04ab68e1c1a48e20f745","avatarUrl":"/avatars/e2376adaa309e30a03ae1ddda6cea773.svg","isPro":false,"fullname":"Danilo de Goede","user":"ddgoede","type":"user"},{"_id":"67b85f84d01134f8987b8ba1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/iJrVlo0vvtITPGhV1q9fe.png","isPro":false,"fullname":"Marc","user":"glasbruch","type":"user"},{"_id":"669132e03202b01a4307088f","avatarUrl":"/avatars/a0c59a67c80ea021c3793f1cb3676c4b.svg","isPro":false,"fullname":"Filipe Laitenberger","user":"flaitenberger","type":"user"},{"_id":"619f93fd6d79d41e14199f4e","avatarUrl":"/avatars/7a799632ef67767dd733562764d48a20.svg","isPro":false,"fullname":"Alexey Bezgin","user":"elderberry17","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"670635814d8c8b01343c35c5","name":"FunAILab","fullname":"Fundamental AI Lab at UTN","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637d21239a5217b88b7549c3/5rDBCg9OOamKK_UqCof_z.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.21576.md","query":{}}">
Papers
arxiv:2607.21576

Self-Supervised Learning of Structured Dynamics from Videos

Published on Jul 23
· Submitted by
Lukas Knobel
on Jul 24
Authors:

Abstract

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.21576
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.21576 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.21576 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.21576 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers