Hugging Face Daily Papers · · 4 min read

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🤖 Why do robots still struggle to act on what they see in 3D?</p>\n<p>Most Vision-Language-Action planners reason in 2D — they predict pixel waypoints and just inherit whatever depth lies beneath them. The plan looks right on the image, but it isn’t grounded in the geometry the robot actually has to move through.</p>\n<p>Our new work 3D HAMSTER tackles this head-on: it predicts metrically grounded 3D end-effector trajectories from a single RGB-D observation + a language instruction, keeping planning and control geometrically consistent end to end.</p>\n<p>Proud to share it’s been accepted to IROS 2026 🎉 </p>\n<p>📄 Paper: <a href=\"https://arxiv.org/abs/2606.31329\" rel=\"nofollow\">https://arxiv.org/abs/2606.31329</a><br>🌐 Project: <a href=\"https://davian-robotics.github.io/3D_HAMSTER/\" rel=\"nofollow\">https://davian-robotics.github.io/3D_HAMSTER/</a><br>💻 Code: <a href=\"https://github.com/DAVIAN-Robotics/3D_HAMSTER\" rel=\"nofollow\">https://github.com/DAVIAN-Robotics/3D_HAMSTER</a><br>🤗 Model: <a href=\"https://huggingface.co/DAVIAN-Robotics/3D_HAMSTER\">https://huggingface.co/DAVIAN-Robotics/3D_HAMSTER</a></p>\n","updatedAt":"2026-07-08T05:01:36.546Z","author":{"_id":"67061bf00cc840f4462a0705","avatarUrl":"/avatars/0fab53fa981f166483aa8ad61172d32d.svg","fullname":"Dongyoon","name":"godnpeter","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8481305241584778},"editors":["godnpeter"],"editorAvatarUrls":["/avatars/0fab53fa981f166483aa8ad61172d32d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.31329","authors":[{"_id":"6a450ee64f1dd35e48fb8c9e","name":"Dongyoon Hwang","hidden":false},{"_id":"6a450ee64f1dd35e48fb8c9f","name":"Byungkun Lee","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca0","name":"Dongjin Kim","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca1","name":"Hyojin Jang","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca2","name":"Hoiyeong Jin","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca3","name":"Jueun Mun","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca4","name":"Minho Park","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca5","name":"Hojoon Lee","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca6","name":"Hyunseung Kim","hidden":false},{"_id":"6a450ee64f1dd35e48fb8ca7","name":"Jaegul Choo","hidden":false}],"publishedAt":"2026-06-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-08T00:00:00.000Z","title":"3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance","submittedOnDailyBy":{"_id":"67061bf00cc840f4462a0705","avatarUrl":"/avatars/0fab53fa981f166483aa8ad61172d32d.svg","isPro":false,"fullname":"Dongyoon","user":"godnpeter","type":"user","name":"godnpeter"},"summary":"Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.","upvotes":4,"discussionId":"6a450ee74f1dd35e48fb8ca8","projectPage":"https://davian-robotics.github.io/3D_HAMSTER/","githubRepo":"https://github.com/DAVIAN-Robotics/3D_HAMSTER","githubRepoAddedBy":"user","ai_summary":"3D HAMSTER framework enhances robot manipulation by integrating a vision-language model with depth encoding to generate metrically accurate 3D trajectories for point cloud-based control policies.","ai_keywords":["Vision-Language Model","point clouds","3D trajectory prediction","depth encoder","dense depth reconstruction","hierarchical framework","low-level policy","metric space"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":7,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67061bf00cc840f4462a0705","avatarUrl":"/avatars/0fab53fa981f166483aa8ad61172d32d.svg","isPro":false,"fullname":"Dongyoon","user":"godnpeter","type":"user"},{"_id":"631c386bc73939ffc0716a37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662793811119-noauth.jpeg","isPro":false,"fullname":"SeongWan Kim","user":"idgmatrix","type":"user"},{"_id":"687363d49a81c7dcbcfa2d84","avatarUrl":"/avatars/5d943a5c811ed931c3fdcfee19253049.svg","isPro":false,"fullname":"jj","user":"realman123","type":"user"},{"_id":"65275bb0d82f71e8fca12a70","avatarUrl":"/avatars/69c1acebd02642d2ff5c0dd445822405.svg","isPro":false,"fullname":"Byungjun Yoon","user":"happyhappy-jun","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.31329.md","query":{}}">
Papers
arxiv:2606.31329

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance

Published on Jun 30
· Submitted by
Dongyoon
on Jul 8
Authors:
,

Abstract

3D HAMSTER framework enhances robot manipulation by integrating a vision-language model with depth encoding to generate metrically accurate 3D trajectories for point cloud-based control policies.

Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.

Community

Paper submitter about 12 hours ago

🤖 Why do robots still struggle to act on what they see in 3D?

Most Vision-Language-Action planners reason in 2D — they predict pixel waypoints and just inherit whatever depth lies beneath them. The plan looks right on the image, but it isn’t grounded in the geometry the robot actually has to move through.

Our new work 3D HAMSTER tackles this head-on: it predicts metrically grounded 3D end-effector trajectories from a single RGB-D observation + a language instruction, keeping planning and control geometrically consistent end to end.

Proud to share it’s been accepted to IROS 2026 🎉

📄 Paper: https://arxiv.org/abs/2606.31329
🌐 Project: https://davian-robotics.github.io/3D_HAMSTER/
💻 Code: https://github.com/DAVIAN-Robotics/3D_HAMSTER
🤗 Model: https://huggingface.co/DAVIAN-Robotics/3D_HAMSTER

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.31329
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2606.31329 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2606.31329 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers