World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving.</p>\n","updatedAt":"2026-08-10T03:06:29.864Z","author":{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","fullname":"Dingkang Liang","name":"dkliang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8884835839271545},"editors":["dkliang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07468","authors":[{"_id":"6a793f888e9301703eaa5e8c","name":"Zongchuang Zhao","hidden":false},{"_id":"6a793f888e9301703eaa5e8d","name":"Xin Zhou","hidden":false},{"_id":"6a793f888e9301703eaa5e8e","name":"Tianyang Xu","hidden":false},{"_id":"6a793f888e9301703eaa5e8f","name":"Zhengyang Sun","hidden":false},{"_id":"6a793f888e9301703eaa5e90","name":"Kaixuan Zhou","hidden":false},{"_id":"6a793f888e9301703eaa5e91","name":"Honglin Li","hidden":false},{"_id":"6a793f888e9301703eaa5e92","name":"Dingkang Liang","hidden":false},{"_id":"6a793f888e9301703eaa5e93","name":"Xiang Bai","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"SimWAM: A Simple World Action Model for End-to-End Autonomous Driving","submittedOnDailyBy":{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user","name":"dkliang"},"summary":"World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/","upvotes":18,"discussionId":"6a793f888e9301703eaa5e94","githubRepo":"https://github.com/H-EmbodVis/SimWAM","githubRepoAddedBy":"user","githubStars":15,"organization":{"_id":"687cebf73858638f66e59f56","name":"H-EmbodVis","fullname":"H-EmbodVis","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/XkebfHCVngT9o82kTAVAV.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user"},{"_id":"6505a83e38b7f6bcfa72274a","avatarUrl":"/avatars/51180e15108460173ca0eea8e492e89b.svg","isPro":false,"fullname":"ZongchuangZhao","user":"zczhao","type":"user"},{"_id":"68711a2ad7fb53afed1a90ff","avatarUrl":"/avatars/2e47648ad22fec6df3e9a9bece11aba1.svg","isPro":false,"fullname":"xuxiaopeng","user":"xuxp","type":"user"},{"_id":"668cb84b410a13fa3d3d1297","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/668cb84b410a13fa3d3d1297/65g-71-QYm4DrJAdUa4O2.png","isPro":false,"fullname":"Xianjin-Wu","user":"HyperbolicCurve","type":"user"},{"_id":"6717c5c36bc2876059ed23ab","avatarUrl":"/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg","isPro":false,"fullname":"Xin Zhou","user":"LMD0311","type":"user"},{"_id":"68d8da2b9541f51bd687b1e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d8da2b9541f51bd687b1e2/3vO1pqV1tnMLbWhxfjnMZ.png","isPro":false,"fullname":"Hengyi Xie","user":"yoloiwig","type":"user"},{"_id":"66744b07ab975c85911ed26e","avatarUrl":"/avatars/913221a99619e05ca4fcd178e625a098.svg","isPro":false,"fullname":"xu","user":"wxu2023","type":"user"},{"_id":"69e6032d501979cfa288e338","avatarUrl":"/avatars/21c5a63188a4d90f4b3eb47dd236ba18.svg","isPro":false,"fullname":"Zhengyang Sun","user":"ZhengyangSun","type":"user"},{"_id":"63ad3de96ee60ca58a409280","avatarUrl":"/avatars/7461f4fda3692f042e556d2a7c339bc0.svg","isPro":false,"fullname":"Qi Liu","user":"QiLiuHKU","type":"user"},{"_id":"69606e67f52664bec09f31ec","avatarUrl":"/avatars/58b176cc04347c095984652c655b4c02.svg","isPro":false,"fullname":"覃品然","user":"Dkuriax","type":"user"},{"_id":"68d5f4340abfe8b8120a56ca","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/zBxGi50dBL0o4OpudiFno.png","isPro":false,"fullname":"yangli_leo_00","user":"YangLi00","type":"user"},{"_id":"67bb4345489cb4dc98b873cb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mRIkBXm-SfdvMUvX7xqdl.png","isPro":false,"fullname":"Wang","user":"Ruzhuo","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"687cebf73858638f66e59f56","name":"H-EmbodVis","fullname":"H-EmbodVis","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/XkebfHCVngT9o82kTAVAV.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07468.md","query":{}}">
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
Abstract
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/
Community
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.07468 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.07468 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.