We introduce Qwen-Drive-1.0, the first vision-language foundation model for autonomous driving that unifies 3D perception and visual question answering at the pretraining stage and further extends to motion planning, while keeping the pretrained VLM architecture entirely untouched. Built on the natively multimodal Qwen3.5-4B, it attaches two external modules. A BEV perception head serves as an explicit, inspectable 3D probe, jointly performing 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert generates future ego trajectories through flow matching. Through staged training, we substantially boost the autonomous driving capability of a general-purpose VLM and validate it on 3D perception, driving visual question answering, and motion planning tasks, forming a unified driving vision-language model that offers a new-generation VLM base for driving-scenario adaptation.</p>\n","updatedAt":"2026-09-02T06:49:48.606Z","author":{"_id":"6717c5c36bc2876059ed23ab","avatarUrl":"/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg","fullname":"Xin Zhou","name":"LMD0311","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8759468197822571},"editors":["LMD0311"],"editorAvatarUrls":["/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00111","authors":[{"_id":"6a97859efe3c2f89286c3912","name":"Xin Zhou","hidden":false},{"_id":"6a97859efe3c2f89286c3913","name":"Zongchuang Zhao","hidden":false},{"_id":"6a97859efe3c2f89286c3914","name":"Zhibo Yang","hidden":false},{"_id":"6a97859efe3c2f89286c3915","name":"Mingsheng Li","hidden":false},{"_id":"6a97859efe3c2f89286c3916","name":"Humen Zhong","hidden":false},{"_id":"6a97859efe3c2f89286c3917","name":"Shuai Bai","hidden":false},{"_id":"6a97859efe3c2f89286c3918","name":"Du Chu","hidden":false},{"_id":"6a97859efe3c2f89286c3919","name":"Ruizhe Chen","hidden":false},{"_id":"6a97859efe3c2f89286c391a","name":"Zhaohai Li","hidden":false},{"_id":"6a97859efe3c2f89286c391b","name":"Jun Tang","hidden":false},{"_id":"6a97859efe3c2f89286c391c","name":"Qiuyue Wang","hidden":false},{"_id":"6a97859efe3c2f89286c391d","name":"Mingkun Yang","hidden":false},{"_id":"6a97859efe3c2f89286c391e","name":"Jiazhao Zhang","hidden":false},{"_id":"6a97859efe3c2f89286c391f","name":"Dayiheng Liu","hidden":false},{"_id":"6a97859efe3c2f89286c3920","name":"Dingkang Liang","hidden":false},{"_id":"6a97859efe3c2f89286c3921","name":"Xiang Bai","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-02T00:00:00.000Z","title":"Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.","upvotes":56,"discussionId":"6a97859efe3c2f89286c3922","ai_summary":"Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.","ai_keywords":["vision-language foundation model","3D perception","bird's-eye-view","BEV perception head","3D object detection","semantic occupancy prediction","BEV map segmentation","Planning Expert","motion planning","staged training"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"64c8b5837fe12ecd0a7e92eb","name":"Qwen","fullname":"Qwen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6215ca5692c0ecfba9186921/hrRM50-6XcdWgg2AKpENG.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"64b8a72952b7353d8c669086","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b8a72952b7353d8c669086/3PUTNmx9kd17gtZ9-Yviw.jpeg","isPro":false,"fullname":"Qi Fan","user":"fanqiNO1","type":"user"},{"_id":"6234a8105e7398c64ac62199","avatarUrl":"/avatars/1ae2fc3910a64bd91d20aadb267c0bc3.svg","isPro":false,"fullname":"Maozhou Ge","user":"Gmc2","type":"user"},{"_id":"6717c5c36bc2876059ed23ab","avatarUrl":"/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg","isPro":false,"fullname":"Xin Zhou","user":"LMD0311","type":"user"},{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user"},{"_id":"6363a1fa123a5d5cd4a800e2","avatarUrl":"/avatars/a0961ca5463aae05de0b1574c0064fae.svg","isPro":false,"fullname":"gbz","user":"greeky","type":"user"},{"_id":"6777a782cb3b36883e4c99d7","avatarUrl":"/avatars/b203a55dd12e95621ffabef78d54d33a.svg","isPro":false,"fullname":"Gangwei Xu","user":"gangweix","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6a6c970d0f04a66c65e4f768","avatarUrl":"/avatars/4fe9d3587f118e29697c4bf1295d7341.svg","isPro":false,"fullname":"jilangqu","user":"jdbhh","type":"user"},{"_id":"668cb84b410a13fa3d3d1297","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/668cb84b410a13fa3d3d1297/65g-71-QYm4DrJAdUa4O2.png","isPro":false,"fullname":"Xianjin-Wu","user":"HyperbolicCurve","type":"user"},{"_id":"68d8da2b9541f51bd687b1e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d8da2b9541f51bd687b1e2/3vO1pqV1tnMLbWhxfjnMZ.png","isPro":false,"fullname":"Hengyi Xie","user":"yoloiwig","type":"user"},{"_id":"69d4e377984dc690fa30a0b9","avatarUrl":"/avatars/be597d31f10da27621447c2b56f3f24b.svg","isPro":false,"fullname":"Chaoqun Zheng","user":"ThisYQ","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"64c8b5837fe12ecd0a7e92eb","name":"Qwen","fullname":"Qwen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6215ca5692c0ecfba9186921/hrRM50-6XcdWgg2AKpENG.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00111.md","query":{}}">
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
Abstract
Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.
We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
Community
We introduce Qwen-Drive-1.0, the first vision-language foundation model for autonomous driving that unifies 3D perception and visual question answering at the pretraining stage and further extends to motion planning, while keeping the pretrained VLM architecture entirely untouched. Built on the natively multimodal Qwen3.5-4B, it attaches two external modules. A BEV perception head serves as an explicit, inspectable 3D probe, jointly performing 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert generates future ego trajectories through flow matching. Through staged training, we substantially boost the autonomous driving capability of a general-purpose VLM and validate it on 3D perception, driving visual question answering, and motion planning tasks, forming a unified driving vision-language model that offers a new-generation VLM base for driving-scenario adaptation.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.00111 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.00111 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.00111 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.