Hugging Face Daily Papers · · 4 min read

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We introduce Qwen-Drive-1.0, the first vision-language foundation model for autonomous driving that unifies 3D perception and visual question answering at the pretraining stage and further extends to motion planning, while keeping the pretrained VLM architecture entirely untouched. Built on the natively multimodal Qwen3.5-4B, it attaches two external modules. A BEV perception head serves as an explicit, inspectable 3D probe, jointly performing 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert generates future ego trajectories through flow matching. Through staged training, we substantially boost the autonomous driving capability of a general-purpose VLM and validate it on 3D perception, driving visual question answering, and motion planning tasks, forming a unified driving vision-language model that offers a new-generation VLM base for driving-scenario adaptation.</p>\n","updatedAt":"2026-09-02T06:49:48.606Z","author":{"_id":"6717c5c36bc2876059ed23ab","avatarUrl":"/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg","fullname":"Xin Zhou","name":"LMD0311","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8759468197822571},"editors":["LMD0311"],"editorAvatarUrls":["/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00111","authors":[{"_id":"6a97859efe3c2f89286c3912","name":"Xin Zhou","hidden":false},{"_id":"6a97859efe3c2f89286c3913","name":"Zongchuang Zhao","hidden":false},{"_id":"6a97859efe3c2f89286c3914","name":"Zhibo Yang","hidden":false},{"_id":"6a97859efe3c2f89286c3915","name":"Mingsheng Li","hidden":false},{"_id":"6a97859efe3c2f89286c3916","name":"Humen Zhong","hidden":false},{"_id":"6a97859efe3c2f89286c3917","name":"Shuai Bai","hidden":false},{"_id":"6a97859efe3c2f89286c3918","name":"Du Chu","hidden":false},{"_id":"6a97859efe3c2f89286c3919","name":"Ruizhe Chen","hidden":false},{"_id":"6a97859efe3c2f89286c391a","name":"Zhaohai Li","hidden":false},{"_id":"6a97859efe3c2f89286c391b","name":"Jun Tang","hidden":false},{"_id":"6a97859efe3c2f89286c391c","name":"Qiuyue Wang","hidden":false},{"_id":"6a97859efe3c2f89286c391d","name":"Mingkun Yang","hidden":false},{"_id":"6a97859efe3c2f89286c391e","name":"Jiazhao Zhang","hidden":false},{"_id":"6a97859efe3c2f89286c391f","name":"Dayiheng Liu","hidden":false},{"_id":"6a97859efe3c2f89286c3920","name":"Dingkang Liang","hidden":false},{"_id":"6a97859efe3c2f89286c3921","name":"Xiang Bai","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-02T00:00:00.000Z","title":"Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.","upvotes":56,"discussionId":"6a97859efe3c2f89286c3922","ai_summary":"Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.","ai_keywords":["vision-language foundation model","3D perception","bird's-eye-view","BEV perception head","3D object detection","semantic occupancy prediction","BEV map segmentation","Planning Expert","motion planning","staged training"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"64c8b5837fe12ecd0a7e92eb","name":"Qwen","fullname":"Qwen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6215ca5692c0ecfba9186921/hrRM50-6XcdWgg2AKpENG.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"64b8a72952b7353d8c669086","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b8a72952b7353d8c669086/3PUTNmx9kd17gtZ9-Yviw.jpeg","isPro":false,"fullname":"Qi Fan","user":"fanqiNO1","type":"user"},{"_id":"6234a8105e7398c64ac62199","avatarUrl":"/avatars/1ae2fc3910a64bd91d20aadb267c0bc3.svg","isPro":false,"fullname":"Maozhou Ge","user":"Gmc2","type":"user"},{"_id":"6717c5c36bc2876059ed23ab","avatarUrl":"/avatars/52c68fb315760df5ef9323cd8ada5a3c.svg","isPro":false,"fullname":"Xin Zhou","user":"LMD0311","type":"user"},{"_id":"67467b5979406f42a14517e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67467b5979406f42a14517e9/wgnUxTd8vWOo0Cr2gaoyG.jpeg","isPro":false,"fullname":"Dingkang Liang","user":"dkliang","type":"user"},{"_id":"6363a1fa123a5d5cd4a800e2","avatarUrl":"/avatars/a0961ca5463aae05de0b1574c0064fae.svg","isPro":false,"fullname":"gbz","user":"greeky","type":"user"},{"_id":"6777a782cb3b36883e4c99d7","avatarUrl":"/avatars/b203a55dd12e95621ffabef78d54d33a.svg","isPro":false,"fullname":"Gangwei Xu","user":"gangweix","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6a6c970d0f04a66c65e4f768","avatarUrl":"/avatars/4fe9d3587f118e29697c4bf1295d7341.svg","isPro":false,"fullname":"jilangqu","user":"jdbhh","type":"user"},{"_id":"668cb84b410a13fa3d3d1297","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/668cb84b410a13fa3d3d1297/65g-71-QYm4DrJAdUa4O2.png","isPro":false,"fullname":"Xianjin-Wu","user":"HyperbolicCurve","type":"user"},{"_id":"68d8da2b9541f51bd687b1e2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d8da2b9541f51bd687b1e2/3vO1pqV1tnMLbWhxfjnMZ.png","isPro":false,"fullname":"Hengyi Xie","user":"yoloiwig","type":"user"},{"_id":"69d4e377984dc690fa30a0b9","avatarUrl":"/avatars/be597d31f10da27621447c2b56f3f24b.svg","isPro":false,"fullname":"Chaoqun Zheng","user":"ThisYQ","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"64c8b5837fe12ecd0a7e92eb","name":"Qwen","fullname":"Qwen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6215ca5692c0ecfba9186921/hrRM50-6XcdWgg2AKpENG.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00111.md","query":{}}">
Papers
arxiv:2609.00111

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Published on Aug 31
· Submitted by
taesiri
on Sep 2
#2 Paper of the day
Authors:
,

Abstract

Qwen-Drive-1.0 is a vision-language foundation model for autonomous driving that unifies 3D perception, visual question answering, and motion planning via shared representations and staged training.

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained vision-language model (VLM) and integrates 3D perception, visual question answering, and motion planning within a unified framework. An external bird's-eye-view (BEV) perception head jointly performs 3D object detection, semantic occupancy prediction, and BEV map segmentation. It serves as a probe of the 3D information accessible from the shared representations and provides an explicit, inspectable interface to 3D scene structure. A Planning Expert conditions on shared VLM representations to generate future ego trajectories. A staged training recipe combines driving supervision with general-purpose vision-language data to acquire driving-specific competence while helping preserve broad visual understanding and instruction-following capabilities. Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability. Comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

Community

We introduce Qwen-Drive-1.0, the first vision-language foundation model for autonomous driving that unifies 3D perception and visual question answering at the pretraining stage and further extends to motion planning, while keeping the pretrained VLM architecture entirely untouched. Built on the natively multimodal Qwen3.5-4B, it attaches two external modules. A BEV perception head serves as an explicit, inspectable 3D probe, jointly performing 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert generates future ego trajectories through flow matching. Through staged training, we substantially boost the autonomous driving capability of a general-purpose VLM and validate it on 3D perception, driving visual question answering, and motion planning tasks, forming a unified driving vision-language model that offers a new-generation VLM base for driving-scenario adaptation.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.00111
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.00111 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.00111 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.00111 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers