Hugging Face Daily Papers · · 5 min read

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving.</p>\n","updatedAt":"2026-08-10T03:06:29.864Z","author":{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","fullname":"Dingkang Liang","name":"dkliang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8884835839271545},"editors":["dkliang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07468","authors":[{"_id":"6a793f888e9301703eaa5e8c","name":"Zongchuang Zhao","hidden":false},{"_id":"6a793f888e9301703eaa5e8d","name":"Xin Zhou","hidden":false},{"_id":"6a793f888e9301703eaa5e8e","name":"Tianyang Xu","hidden":false},{"_id":"6a793f888e9301703eaa5e8f","name":"Zhengyang Sun","hidden":false},{"_id":"6a793f888e9301703eaa5e90","name":"Kaixuan Zhou","hidden":false},{"_id":"6a793f888e9301703eaa5e91","name":"Honglin Li","hidden":false},{"_id":"6a793f888e9301703eaa5e92","name":"Dingkang Liang","hidden":false},{"_id":"6a793f888e9301703eaa5e93","name":"Xiang Bai","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"SimWAM: A Simple World Action Model for End-to-End Autonomous Driving","submittedOnDailyBy":{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user","name":"dkliang"},"summary":"World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/","upvotes":18,"discussionId":"6a793f888e9301703eaa5e94","githubRepo":"https://github.com/H-EmbodVis/SimWAM","githubRepoAddedBy":"user","githubStars":15,"organization":{"_id":"687cebf73858638f66e59f56","name":"H-EmbodVis","fullname":"H-EmbodVis","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/XkebfHCVngT9o82kTAVAV.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user"},{"_id":"6505a83e38b7f6bcfa72274a","avatarUrl":"/avatars/51180e15108460173ca0eea8e492e89b.svg","isPro":false,"fullname":"ZongchuangZhao","user":"zczhao","type":"user"},{"_id":"68711a2ad7fb53afed1a90ff","avatarUrl":"/avatars/2e47648ad22fec6df3e9a9bece11aba1.svg","isPro":false,"fullname":"xuxiaopeng","user":"xuxp","type":"user"},{"_id":"668cb84b410a13fa3d3d1297","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/668cb84b410a13fa3d3d1297/65g-71-QYm4DrJAdUa4O2.png","isPro":false,"fullname":"Xianjin-Wu","user":"HyperbolicCurve","type":"user"},{"_id":"6717c5c36bc2876059ed23ab","avatarUrl":"/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg","isPro":false,"fullname":"Xin Zhou","user":"LMD0311","type":"user"},{"_id":"68d8da2b9541f51bd687b1e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d8da2b9541f51bd687b1e2/3vO1pqV1tnMLbWhxfjnMZ.png","isPro":false,"fullname":"Hengyi Xie","user":"yoloiwig","type":"user"},{"_id":"66744b07ab975c85911ed26e","avatarUrl":"/avatars/913221a99619e05ca4fcd178e625a098.svg","isPro":false,"fullname":"xu","user":"wxu2023","type":"user"},{"_id":"69e6032d501979cfa288e338","avatarUrl":"/avatars/21c5a63188a4d90f4b3eb47dd236ba18.svg","isPro":false,"fullname":"Zhengyang Sun","user":"ZhengyangSun","type":"user"},{"_id":"63ad3de96ee60ca58a409280","avatarUrl":"/avatars/7461f4fda3692f042e556d2a7c339bc0.svg","isPro":false,"fullname":"Qi Liu","user":"QiLiuHKU","type":"user"},{"_id":"69606e67f52664bec09f31ec","avatarUrl":"/avatars/58b176cc04347c095984652c655b4c02.svg","isPro":false,"fullname":"覃品然","user":"Dkuriax","type":"user"},{"_id":"68d5f4340abfe8b8120a56ca","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/zBxGi50dBL0o4OpudiFno.png","isPro":false,"fullname":"yangli_leo_00","user":"YangLi00","type":"user"},{"_id":"67bb4345489cb4dc98b873cb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mRIkBXm-SfdvMUvX7xqdl.png","isPro":false,"fullname":"Wang","user":"Ruzhuo","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"687cebf73858638f66e59f56","name":"H-EmbodVis","fullname":"H-EmbodVis","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/XkebfHCVngT9o82kTAVAV.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07468.md","query":{}}">
Papers
arxiv:2608.07468

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

Published on Aug 7
· Submitted by
Dingkang Liang
on Aug 10
#1 Paper of the day
Authors:
,

Abstract

World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/

Community

Paper submitter about 3 hours ago

World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing the video branch to be discarded after training and leaving a self-contained planner that directly predicts trajectories. Since the two experts share no parameters and interact only through a unified attention interface, the video backbone could be replaced and the action expert scaled independently without modifying the learning objective or inference pipeline. We further apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Our SimWAM achieves 91.5 PDMS on NAVSIM, surpasses state-of-the-art WAM-based planners with substantially lower latency, and transfers zero-shot to nuScenes. These results position SimWAM as a simple yet solid baseline that could readily benefit from advances in video generation for efficient autonomous driving.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.07468
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.07468 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.07468 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers