We identify an anchoring bias in autonomous-driving VLMs caused by GT-trajectory-conditioned CoT supervision. To address this issue, we introduce AD-MCQ and DEFT-RLVR, a candidate-grounded reinforcement learning framework that reformulates AD planning as verifiable trajectory selection. This formulation encourages more causally faithful reasoning while preserving the model’s general visual capabilities. Moreover, because AD-MCQ operates entirely within the VLM and allows task difficulty to be flexibly controlled through candidate construction, it provides a practical, scalable, and readily deployable foundation for future research.</p>\n","updatedAt":"2026-08-04T07:36:46.902Z","author":{"_id":"6593d2329e16fa7510e0876a","avatarUrl":"/avatars/c200676dd9dc1db8a3e27388251aea49.svg","fullname":"hzx","name":"hzxllll","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8895012140274048},"editors":["hzxllll"],"editorAvatarUrls":["/avatars/c200676dd9dc1db8a3e27388251aea49.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.01755","authors":[{"_id":"6a714b44ec5082b9f872ccd6","name":"Zixuan Huang","hidden":false},{"_id":"6a714b44ec5082b9f872ccd7","name":"Yang Zhou","hidden":false},{"_id":"6a714b44ec5082b9f872ccd8","name":"Kaixuan Wang","hidden":false},{"_id":"6a714b44ec5082b9f872ccd9","name":"Guli Zhang","hidden":false},{"_id":"6a714b44ec5082b9f872ccda","name":"Hongyan Xie","hidden":false},{"_id":"6a714b44ec5082b9f872ccdb","name":"Yakun Zhu","hidden":false},{"_id":"6a714b44ec5082b9f872ccdc","name":"Hao Geng","hidden":false},{"_id":"6a714b44ec5082b9f872ccdd","name":"Yikun Ban","hidden":false},{"_id":"6a714b44ec5082b9f872ccde","name":"Deqing Wang","hidden":false}],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs","submittedOnDailyBy":{"_id":"6593d2329e16fa7510e0876a","avatarUrl":"/avatars/c200676dd9dc1db8a3e27388251aea49.svg","isPro":false,"fullname":"hzx","user":"hzxllll","type":"user","name":"hzxllll"},"summary":"Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.","upvotes":1,"discussionId":"6a714b44ec5082b9f872ccdf","githubRepo":"https://github.com/hzx122/DEFT-RLVR","githubRepoAddedBy":"user","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6593d2329e16fa7510e0876a","avatarUrl":"/avatars/c200676dd9dc1db8a3e27388251aea49.svg","isPro":false,"fullname":"hzx","user":"hzxllll","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.01755.md","query":{}}">
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Published on Aug 3
· Submitted by hzx on Aug 4 Abstract
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, producing less causally faithful CoTs and substantially more severe hallucinations, especially in causally challenging scenes. Removing the GT trajectory eliminates this shortcut, but open-ended trajectory generation entangles high-level decision-making with precise geometric synthesis and low-level dynamics. To make trajectory-level driving decisions verifiable without requiring open-ended trajectory synthesis, we introduce Autonomous-Driving Multiple-Choice Question (AD-MCQ), which casts planning as selection among explicit trajectory candidates. Taking this a step further, we propose Deferred Exposure of Future Trajectories for RLVR (DEFT-RLVR) to transform future trajectories from pre-decision anchors into post-decision verification targets. Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. With VLM-only inference and controllable difficulty through candidate construction, AD-MCQ provides a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.
Community
We identify an anchoring bias in autonomous-driving VLMs caused by GT-trajectory-conditioned CoT supervision. To address this issue, we introduce AD-MCQ and DEFT-RLVR, a candidate-grounded reinforcement learning framework that reformulates AD planning as verifiable trajectory selection. This formulation encourages more causally faithful reasoning while preserving the model’s general visual capabilities. Moreover, because AD-MCQ operates entirely within the VLM and allows task difficulty to be flexibly controlled through candidate construction, it provides a practical, scalable, and readily deployable foundation for future research.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.01755 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.