BadWAM models World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. It instantiates two complementary objectives. The action-only adversarial attack prioritizes disruption by driving the model toward task-failing actions. The imagination-preserving adversarial attack additionally keeps the predicted future close to the clean imagination, producing a stealthier failure mode.</p>\n","updatedAt":"2026-07-17T06:52:25.813Z","author":{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","fullname":"LIQIIIII","name":"LIQIIIII","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":2,"identifiedLanguage":{"language":"en","probability":0.883745014667511},"editors":["LIQIIIII"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png"],"reactions":[],"isReport":false}},{"id":"6a59d1dee23a838753696a81","author":{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","fullname":"LIQIIIII","name":"LIQIIIII","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false},"createdAt":"2026-07-17T06:55:26.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\nhttps://cdn-uploads.huggingface.co/production/uploads/6706ab1168e9971e91bad6f7/YBnJj_XTeJlkN19lGnl0O.qt\n","html":"<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/6706ab1168e9971e91bad6f7/YBnJj_XTeJlkN19lGnl0O.qt\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-07-17T06:55:26.921Z","author":{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","fullname":"LIQIIIII","name":"LIQIIIII","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3740061819553375},"editors":["LIQIIIII"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.15207","authors":[{"_id":"6a5993cd6c2e371e6ca380b7","name":"Qi Li","hidden":false},{"_id":"6a5993cd6c2e371e6ca380b8","name":"Xingyi Yang","hidden":false},{"_id":"6a5993cd6c2e371e6ca380b9","name":"Xinchao Wang","hidden":false}],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"BadWAM: When World-Action Models Dream Right but Act Wrong","submittedOnDailyBy":{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","isPro":false,"fullname":"LIQIIIII","user":"LIQIIIII","type":"user","name":"LIQIIIII"},"summary":"World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.","upvotes":35,"discussionId":"6a5993cd6c2e371e6ca380ba","projectPage":"https://liqiiiii.github.io/BadWAM/","githubRepo":"https://github.com/LiQiiiii/BadWAM","githubRepoAddedBy":"user","githubStars":32},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6706ab1168e9971e91bad6f7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tWSXpBEAm0d8gTDWFRxTS.png","isPro":false,"fullname":"LIQIIIII","user":"LIQIIIII","type":"user"},{"_id":"634cfebc350bcee9bed20a4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634cfebc350bcee9bed20a4d/fN47nN5rhw-HJaFLBZWQy.png","isPro":false,"fullname":"Xingyi Yang","user":"adamdad","type":"user"},{"_id":"677fbbf5f2e19477cb809830","avatarUrl":"/avatars/51af04f28038870f3ec418cc4909ecd0.svg","isPro":false,"fullname":"Tianbo Pan","user":"pan7386","type":"user"},{"_id":"5df833bdda6d0311fd3d5403","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/5df833bdda6d0311fd3d5403/62OtGJEQXdOuhV9yCd4HS.png","isPro":false,"fullname":"Weihao Yu","user":"whyu","type":"user"},{"_id":"6860fe55a1ab4d5c885c3edf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/HNSwVGoTd33mgbyjxA-fc.jpeg","isPro":false,"fullname":"QIN ZHIBIN","user":"tuantuan0321","type":"user"},{"_id":"66eee2b9ce6e5db9b30cccc0","avatarUrl":"/avatars/667b44f36405c1f60b2d42bb1a4d1cc1.svg","isPro":false,"fullname":"刘昊","user":"lejjej","type":"user"},{"_id":"686e948128daebed525bd6ea","avatarUrl":"/avatars/6a918a1e2c1d27ce803e9eb1aeb43bb8.svg","isPro":false,"fullname":"tauzhao","user":"tauzhao","type":"user"},{"_id":"69048472836d624e6d1b27f9","avatarUrl":"/avatars/37fd5d7ac5b56d6c428e67ef4f5eba79.svg","isPro":false,"fullname":"bojun zou","user":"nrbzd","type":"user"},{"_id":"6624f53748e016b5ea587d40","avatarUrl":"/avatars/f8c16f45de0c3e32437f6e960a5b0959.svg","isPro":false,"fullname":"Shihua Zhang","user":"SuhZhang","type":"user"},{"_id":"647dd8f9a49bffab5d6fe46e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/UkYcsNfnvKOotfKfTNcEk.png","isPro":false,"fullname":"Yin Bo","user":"YINBO0927","type":"user"},{"_id":"6503706736bc3431217f5935","avatarUrl":"/avatars/42ca8111d3aba9db7ac269835dacd8c9.svg","isPro":false,"fullname":"Qingyuan Wang","user":"QingyuanWang","type":"user"},{"_id":"69f05abccbbd5b9b39fdebd9","avatarUrl":"/avatars/844bb8b47c961985609221d0d41bd84f.svg","isPro":false,"fullname":"HWQ","user":"FDZUO","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.15207.md","query":{}}">
BadWAM: When World-Action Models Dream Right but Act Wrong
Abstract
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.
Community
BadWAM models World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. It instantiates two complementary objectives. The action-only adversarial attack prioritizes disruption by driving the model toward task-failing actions. The imagination-preserving adversarial attack additionally keeps the predicted future close to the clean imagination, producing a stealthier failure mode.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.15207 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.15207 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.