Hugging Face Daily Papers · · 4 min read

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

ShadowDancer learns to control video world models by watching each dynamics twice — a video and its “shadow” (same motion, resampled appearance). This cross-shadow paradigm yields one unified action interface for any demonstrable behavior: first/third-person gameplay, open worlds, human motion, camera, and robot manipulation — no labels, no fine-tuning.</p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/6753fe5ef49782992ebc00b6/Ubgpeu3Cfz9TMR5M8yOVM.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-07-31T06:37:23.236Z","author":{"_id":"6753fe5ef49782992ebc00b6","avatarUrl":"/avatars/899954f5582237520ee9491ac85bc979.svg","fullname":"Jin Cao","name":"TmaKiss","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7875763773918152},"editors":["TmaKiss"],"editorAvatarUrls":["/avatars/899954f5582237520ee9491ac85bc979.svg"],"reactions":[],"isReport":false}},{"id":"6a6c471a03a9a0556db5790a","author":{"_id":"65f1713552c38a91e0a445e8","avatarUrl":"/avatars/47ab3ada51c9b9976ac1cd0c4301c373.svg","fullname":"kaipeng","name":"kpzhang996","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false},"createdAt":"2026-07-31T06:56:26.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"ShadowDancer gives interactive video world models an any-action, frame-level control.\n![teaser](https://cdn-uploads.huggingface.co/production/uploads/65f1713552c38a91e0a445e8/OK8LhoMsy_wVlmmFX77ki.png)","html":"<p>ShadowDancer gives interactive video world models an any-action, frame-level control.<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65f1713552c38a91e0a445e8/OK8LhoMsy_wVlmmFX77ki.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65f1713552c38a91e0a445e8/OK8LhoMsy_wVlmmFX77ki.png\" alt=\"teaser\"></a></p>\n","updatedAt":"2026-07-31T06:56:26.516Z","author":{"_id":"65f1713552c38a91e0a445e8","avatarUrl":"/avatars/47ab3ada51c9b9976ac1cd0c4301c373.svg","fullname":"kaipeng","name":"kpzhang996","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.38963860273361206},"editors":["kpzhang996"],"editorAvatarUrls":["/avatars/47ab3ada51c9b9976ac1cd0c4301c373.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28362","authors":[{"_id":"6a6c4023202e2d9e3ffdb845","user":{"_id":"6753fe5ef49782992ebc00b6","avatarUrl":"/avatars/899954f5582237520ee9491ac85bc979.svg","isPro":false,"fullname":"Jin Cao","user":"TmaKiss","type":"user","name":"TmaKiss"},"name":"Jin Cao","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.650Z","hidden":false},{"_id":"6a6c4023202e2d9e3ffdb846","name":"Zian Meng","hidden":false},{"_id":"6a6c4023202e2d9e3ffdb847","name":"Kaipeng Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65f1713552c38a91e0a445e8/eLIiEORJuq4aA8DuoZb8W.png"],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow","submittedOnDailyBy":{"_id":"65f1713552c38a91e0a445e8","avatarUrl":"/avatars/47ab3ada51c9b9976ac1cd0c4301c373.svg","isPro":false,"fullname":"kaipeng","user":"kpzhang996","type":"user","name":"kpzhang996"},"summary":"We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io","upvotes":15,"discussionId":"6a6c4024202e2d9e3ffdb848","projectPage":"https://ShadowDancer-1.github.io","githubRepo":"https://github.com/AlayaLab/ShadowDancer","githubRepoAddedBy":"user","githubStars":6,"organization":{"_id":"689f08c50df4fcf7fddc0b08","name":"AlayaLab","fullname":"Alaya Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63342778d92c5842ae728aef/dNCvNz9MMshksG2xspIbM.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6753fe5ef49782992ebc00b6","avatarUrl":"/avatars/899954f5582237520ee9491ac85bc979.svg","isPro":false,"fullname":"Jin Cao","user":"TmaKiss","type":"user"},{"_id":"65f1713552c38a91e0a445e8","avatarUrl":"/avatars/47ab3ada51c9b9976ac1cd0c4301c373.svg","isPro":false,"fullname":"kaipeng","user":"kpzhang996","type":"user"},{"_id":"6a3b5265c2005a67a66ea05d","avatarUrl":"/avatars/bc23812644fd1457a764552d59ca263e.svg","isPro":false,"fullname":"Aurora","user":"AuroraRyan2","type":"user"},{"_id":"64266cc885f26ab94af48b82","avatarUrl":"/avatars/d4cfa427f641f301a376c4ef68491aa1.svg","isPro":false,"fullname":"DougeFu","user":"DougeFu","type":"user"},{"_id":"63342778d92c5842ae728aef","avatarUrl":"/avatars/888eb265643633c5fdd7048be9bfe98f.svg","isPro":false,"fullname":"Fengbo Lan","user":"fblan","type":"user"},{"_id":"68bb92636e97e5a7f85e18f4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FOLs8rV-CoJpYxx_MbxpW.png","isPro":false,"fullname":"Andrew Kane","user":"Andrew1129","type":"user"},{"_id":"68323f961e5e5c17eb1f0de4","avatarUrl":"/avatars/e0b56c721c2aec1daf52b67f05093a2c.svg","isPro":false,"fullname":"sue","user":"jzf0634","type":"user"},{"_id":"674ea59a8f2e7614a6c72f26","avatarUrl":"/avatars/86fd4c6d7d33435de49b35659bf65265.svg","isPro":false,"fullname":"Chuanhao","user":"ChuanhaoLi","type":"user"},{"_id":"64a6f3defd819e42d2d28402","avatarUrl":"/avatars/079d6c1ad19f38c25ccfdfbe56a778e6.svg","isPro":false,"fullname":"xjxu","user":"xjxu21","type":"user"},{"_id":"6896ba75214e7b510999f0ca","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ZQ10kysVT3L8bQxzDIqNT.png","isPro":false,"fullname":"VSTP(SII)","user":"LECVSTP","type":"user"},{"_id":"67d91432fcb289e3f21c881b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/EIyAq-ta-GLP6QkN7kf1n.png","isPro":false,"fullname":"Jiaming Tan","user":"oneandonly211","type":"user"},{"_id":"676bc71e490f3664721e81eb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/DRAdleyZ_Zy5h4IldZ3gb.png","isPro":false,"fullname":"Sakura Sato","user":"SakuraSato","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"689f08c50df4fcf7fddc0b08","name":"AlayaLab","fullname":"Alaya Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63342778d92c5842ae728aef/dNCvNz9MMshksG2xspIbM.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.28362.md","query":{}}">
Papers
arxiv:2607.28362

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

Published on Jul 30
· Submitted by
kaipeng
on Jul 31
Authors:

Abstract

We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io

Community

Paper author about 3 hours ago

ShadowDancer learns to control video world models by watching each dynamics twice — a video and its “shadow” (same motion, resampled appearance). This cross-shadow paradigm yields one unified action interface for any demonstrable behavior: first/third-person gameplay, open worlds, human motion, camera, and robot manipulation — no labels, no fine-tuning.

Paper submitter about 3 hours ago

ShadowDancer gives interactive video world models an any-action, frame-level control.
teaser

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.28362
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.28362 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.28362 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.28362 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers