Hugging Face Daily Papers · · 6 min read

DriveZero: End-to-End Driving Beyond Human Demonstrations

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/6448baa3e780dbfc89058bc3/hMFGq7hdLd-qlLaYcnfpv.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6448baa3e780dbfc89058bc3/hMFGq7hdLd-qlLaYcnfpv.png\" alt=\"teaser\"></a></p>\n<p>DriveZero decomposes driving into an action model and a perception model, pretrains each in the regime best suited to it, and unifies them by distillation.</p>\n<ul>\n<li><strong>DriveRL</strong>, the action model, learns to drive from scratch with closed-loop RL. It converts real nuPlan logs into mixed-agent interactive worlds, where each background actor receives its own behavior provider, so that log replay, rule-based behaviors, and learned policies coexist within one scene. In these worlds, a 5.7M-parameter privileged policy is trained with PPO.</li>\n<li><strong>DriveVFM</strong>, the perception model, learns from massive raw images by distilling frozen vision foundation models. It consolidates DINOv3, SigLIP2, SAM, and Depth Anything V2 into a single driving backbone, with no task labels.</li>\n<li><strong>DriveZero</strong> unifies the two. It encodes multi-view images with DriveVFM, decodes trajectory proposals, and learns them from DriveRL rollouts instead of human trajectories, including rollouts under augmented goals that the human log never contained.</li>\n</ul>\n<p>🏆 <strong>DriveRL exceeds the log-replay expert on all six nuPlan closed-loop settings. DriveZero, trained only on DriveRL rollouts, surpasses the human driver on NAVSIMv1 and leads NAVSIMv2 and the true closed-loop HUGSIM benchmark.</strong></p>\n","updatedAt":"2026-09-09T01:58:33.190Z","author":{"_id":"6448baa3e780dbfc89058bc3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448baa3e780dbfc89058bc3/vVwmu6aFPa7YTa8qEYexB.jpeg","fullname":"Haochen Tian","name":"StarBurger","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.860975980758667},"editors":["StarBurger"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6448baa3e780dbfc89058bc3/vVwmu6aFPa7YTa8qEYexB.jpeg"],"reactions":[],"isReport":false}},{"id":"6aa0e55d8e8152f9fee584cc","author":{"_id":"647ee09e454af0237bd23c06","avatarUrl":"/avatars/3f639ebbec1693e9eb0cf37ab05c9968.svg","fullname":"Zhizhao Duan","name":"deepdream-dzz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-09-09T04:49:33.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"小米这版智驾1.17太拉了呀","html":"<p>小米这版智驾1.17太拉了呀</p>\n","updatedAt":"2026-09-09T04:49:33.064Z","author":{"_id":"647ee09e454af0237bd23c06","avatarUrl":"/avatars/3f639ebbec1693e9eb0cf37ab05c9968.svg","fullname":"Zhizhao Duan","name":"deepdream-dzz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"zh","probability":0.34823399782180786},"editors":["deepdream-dzz"],"editorAvatarUrls":["/avatars/3f639ebbec1693e9eb0cf37ab05c9968.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06055","authors":[{"_id":"6aa0bbd4d0174964227bec05","name":"Hao He","hidden":false},{"_id":"6aa0bbd4d0174964227bec06","name":"Chengcheng Hu","hidden":false},{"_id":"6aa0bbd4d0174964227bec07","name":"Zirun Su","hidden":false},{"_id":"6aa0bbd4d0174964227bec08","user":{"_id":"6aa0c6d20cb4390655a73862","avatarUrl":"/avatars/6d54c7d6ab5400e8ad26be5f728faa40.svg","isPro":false,"fullname":"zhangheng19931123","user":"zhangheng1123","type":"user","name":"zhangheng1123"},"name":"Heng Zhang","status":"claimed_verified","statusLastChangedAt":"2026-09-09T09:22:59.055Z","hidden":false},{"_id":"6aa0bbd4d0174964227bec09","name":"Haisong Liu","hidden":false},{"_id":"6aa0bbd4d0174964227bec0a","name":"Jinke Li","hidden":false},{"_id":"6aa0bbd4d0174964227bec0b","user":{"_id":"6448baa3e780dbfc89058bc3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448baa3e780dbfc89058bc3/vVwmu6aFPa7YTa8qEYexB.jpeg","isPro":false,"fullname":"Haochen Tian","user":"StarBurger","type":"user","name":"StarBurger"},"name":"Haochen Tian","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.560Z","hidden":false},{"_id":"6aa0bbd4d0174964227bec0c","name":"Zhenwei Shen","hidden":false},{"_id":"6aa0bbd4d0174964227bec0d","name":"Hongyang Li","hidden":false},{"_id":"6aa0bbd4d0174964227bec0e","name":"Zhichao Li","hidden":false},{"_id":"6aa0bbd4d0174964227bec0f","name":"Yunchen Yang","hidden":false},{"_id":"6aa0bbd4d0174964227bec10","name":"Bochao Huang","hidden":false},{"_id":"6aa0bbd4d0174964227bec11","name":"Siyu Zhang","hidden":false},{"_id":"6aa0bbd4d0174964227bec12","name":"Kuangye Chen","hidden":false},{"_id":"6aa0bbd4d0174964227bec13","name":"Xiongjie Zhang","hidden":false},{"_id":"6aa0bbd4d0174964227bec14","name":"Wentao Dai","hidden":false},{"_id":"6aa0bbd4d0174964227bec15","name":"Hengchen Dai","hidden":false},{"_id":"6aa0bbd4d0174964227bec16","name":"Siyuan Liu","hidden":false},{"_id":"6aa0bbd4d0174964227bec17","name":"Zehao Huang","hidden":false},{"_id":"6aa0bbd4d0174964227bec18","name":"Naiyan Wang","hidden":false}],"publishedAt":"2026-09-05T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"DriveZero: End-to-End Driving Beyond Human Demonstrations","submittedOnDailyBy":{"_id":"6448baa3e780dbfc89058bc3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448baa3e780dbfc89058bc3/vVwmu6aFPa7YTa8qEYexB.jpeg","isPro":false,"fullname":"Haochen Tian","user":"StarBurger","type":"user","name":"StarBurger"},"summary":"Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded trajectories. This report presents DriveZero, an end-to-end system that learns driving behavior beyond human demonstrations. It decomposes driving into a perception model and an action model, pretrains each in the regime best suited to it, and combines them into one end-to-end planner. The two models call for different learning recipes: perception must understand the world, and benefits from massive and diverse visual data; action must interact with it, and requires closed-loop feedback. On the action side, we introduce DriveRL, a mixed-agent closed-loop reinforcement-learning framework. It converts real driving logs into interactive worlds, where a privileged teacher policy is trained with PPO through closed-loop rollouts. For the perception model, DriveVFM consolidates multiple frozen vision foundation models, including DINOv3, SigLIP2, SAM and Depth Anything V2, into a single backbone from raw images alone, requiring no task-specific annotations. DriveZero then unifies the two: a camera-only planner that distills the frozen DriveRL teacher through its rolled-out trajectories. The goal-conditioned teacher can moreover be queried under augmented driving intents, yielding diverse, goal-consistent supervision that logged data cannot provide. On nuPlan, DriveRL with value-guided test-time action search achieves a mean score of 93.57 across the Val14, Test14-hard, and Test14-random community splits in both non-reactive and reactive modes, exceeding the Log-Replay expert on all three splits. DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2 and the closed-loop HUGSIM benchmark without any human trajectory supervision.","upvotes":50,"discussionId":"6aa0bbd4d0174964227bec19","projectPage":"https://xiaomiautol3.github.io/DriveZero/","githubRepo":"https://github.com/XiaomiAutoL3/DriveZero","githubRepoAddedBy":"user","ai_summary":"DriveZero is an end-to-end autonomous driving system that combines a vision foundation model for perception with a closed-loop reinforcement learning action model to learn driving behaviors beyond human demonstrations.","ai_keywords":["DriveZero","DriveRL","PPO","closed-loop reinforcement learning","vision foundation models","DINOv3","SigLIP2","SAM","Depth Anything V2","goal-conditioned teacher","value-guided test-time action search"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":64},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6448baa3e780dbfc89058bc3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6448baa3e780dbfc89058bc3/vVwmu6aFPa7YTa8qEYexB.jpeg","isPro":false,"fullname":"Haochen Tian","user":"StarBurger","type":"user"},{"_id":"6a9bd62f0edd4c549dd8bfca","avatarUrl":"/avatars/c072abb2d9b439a64134b25f4f6bb189.svg","isPro":false,"fullname":"thc","user":"haochentian","type":"user"},{"_id":"64e57309b78bc92221ce3b70","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e57309b78bc92221ce3b70/jAc2mnn_rWcqH7epCCwT2.png","isPro":true,"fullname":"OpenDriveLab","user":"OpenDriveLab-org","type":"user"},{"_id":"658e8954a6567cb93c186888","avatarUrl":"/avatars/0d4aee3af5909f931f33e841f42b4e28.svg","isPro":false,"fullname":"zhang zhifang","user":"zhangzhifang","type":"user"},{"_id":"65102cfc469c325dc4c919ab","avatarUrl":"/avatars/cfcdfa83190e80d58c25f5551036bb27.svg","isPro":false,"fullname":"Haisong Liu","user":"afterthat97","type":"user"},{"_id":"687a03096005d067e1a99ce8","avatarUrl":"/avatars/e497966170ded000260f5f683308e8eb.svg","isPro":false,"fullname":"winstywang","user":"winsty","type":"user"},{"_id":"695e647d1a521ad0a5c46e89","avatarUrl":"/avatars/fc5633585f040b9a586d79952c95a916.svg","isPro":false,"fullname":"suzirun","user":"suzirun","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"692fe17e27917f8ddeeff8d8","avatarUrl":"/avatars/6415ade540e0752900111b4c82cdabf7.svg","isPro":false,"fullname":"Yuan","user":"Lime7","type":"user"},{"_id":"6aa0c6d20cb4390655a73862","avatarUrl":"/avatars/6d54c7d6ab5400e8ad26be5f728faa40.svg","isPro":false,"fullname":"zhangheng19931123","user":"zhangheng1123","type":"user"},{"_id":"6aa0c9f46c5c4d72c1ce69dd","avatarUrl":"/avatars/c2b8734bc82bc511fb473b81ad3d094f.svg","isPro":false,"fullname":"Hao He","user":"Hugohhh","type":"user"},{"_id":"66a27f8cd3449709d69216ce","avatarUrl":"/avatars/71cd4df83a9f086073768c2fc481fc7c.svg","isPro":false,"fullname":"fenfenda","user":"fenfenda","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06055.md","query":{}}">
Papers
arxiv:2609.06055

DriveZero: End-to-End Driving Beyond Human Demonstrations

Published on Sep 5
· Submitted by
Haochen Tian
on Sep 9
Authors:
,

Abstract

DriveZero is an end-to-end autonomous driving system that combines a vision foundation model for perception with a closed-loop reinforcement learning action model to learn driving behaviors beyond human demonstrations.

Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded trajectories. This report presents DriveZero, an end-to-end system that learns driving behavior beyond human demonstrations. It decomposes driving into a perception model and an action model, pretrains each in the regime best suited to it, and combines them into one end-to-end planner. The two models call for different learning recipes: perception must understand the world, and benefits from massive and diverse visual data; action must interact with it, and requires closed-loop feedback. On the action side, we introduce DriveRL, a mixed-agent closed-loop reinforcement-learning framework. It converts real driving logs into interactive worlds, where a privileged teacher policy is trained with PPO through closed-loop rollouts. For the perception model, DriveVFM consolidates multiple frozen vision foundation models, including DINOv3, SigLIP2, SAM and Depth Anything V2, into a single backbone from raw images alone, requiring no task-specific annotations. DriveZero then unifies the two: a camera-only planner that distills the frozen DriveRL teacher through its rolled-out trajectories. The goal-conditioned teacher can moreover be queried under augmented driving intents, yielding diverse, goal-consistent supervision that logged data cannot provide. On nuPlan, DriveRL with value-guided test-time action search achieves a mean score of 93.57 across the Val14, Test14-hard, and Test14-random community splits in both non-reactive and reactive modes, exceeding the Log-Replay expert on all three splits. DriveZero achieves state-of-the-art performance on NAVSIMv1, NAVSIMv2 and the closed-loop HUGSIM benchmark without any human trajectory supervision.

Community

Paper author Paper submitter about 12 hours ago

teaser

DriveZero decomposes driving into an action model and a perception model, pretrains each in the regime best suited to it, and unifies them by distillation.

  • DriveRL, the action model, learns to drive from scratch with closed-loop RL. It converts real nuPlan logs into mixed-agent interactive worlds, where each background actor receives its own behavior provider, so that log replay, rule-based behaviors, and learned policies coexist within one scene. In these worlds, a 5.7M-parameter privileged policy is trained with PPO.
  • DriveVFM, the perception model, learns from massive raw images by distilling frozen vision foundation models. It consolidates DINOv3, SigLIP2, SAM, and Depth Anything V2 into a single driving backbone, with no task labels.
  • DriveZero unifies the two. It encodes multi-view images with DriveVFM, decodes trajectory proposals, and learns them from DriveRL rollouts instead of human trajectories, including rollouts under augmented goals that the human log never contained.

🏆 DriveRL exceeds the log-replay expert on all six nuPlan closed-loop settings. DriveZero, trained only on DriveRL rollouts, surpasses the human driver on NAVSIMv1 and leads NAVSIMv2 and the true closed-loop HUGSIM benchmark.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.06055
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.06055 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.06055 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.06055 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers