Hugging Face Daily Papers · · 4 min read

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Attention-DP3 enhances 3D diffusion policies with object-aware geometric attention, improving target localization and robustness in cluttered scenes while keeping the DP3 backbone unchanged. It consistently outperforms DP3 across simulation and real-world benchmarks, with gains of up to 31% under heavy clutter. Attention-DP3 is accepted by ECCV 2026. Code is avaliable at <a href=\"https://github.com/zhangzhongbo2213/Attention-DP3\" rel=\"nofollow\">https://github.com/zhangzhongbo2213/Attention-DP3</a>.</p>\n","updatedAt":"2026-09-15T02:29:18.333Z","author":{"_id":"6575702b15b1ca184b0b2700","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg","fullname":"Zaibin Zhang","name":"MrBean2024","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8857409358024597},"editors":["MrBean2024"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.13318","authors":[{"_id":"6aa8aca25dd4cb9b4cc024fa","name":"Changbo Yan","hidden":false},{"_id":"6aa8aca25dd4cb9b4cc024fb","name":"Zhongbo Zhang","hidden":false},{"_id":"6aa8aca25dd4cb9b4cc024fc","name":"Zaibin Zhang","hidden":false},{"_id":"6aa8aca25dd4cb9b4cc024fd","name":"Yifan Wang","hidden":false},{"_id":"6aa8aca25dd4cb9b4cc024fe","name":"Lijun Wang","hidden":false},{"_id":"6aa8aca25dd4cb9b4cc024ff","name":"Huchuan Lu","hidden":false}],"publishedAt":"2026-09-10T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning","submittedOnDailyBy":{"_id":"6575702b15b1ca184b0b2700","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg","isPro":false,"fullname":"Zaibin Zhang","user":"MrBean2024","type":"user","name":"MrBean2024"},"summary":"3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may be partially occluded or tightly intermingled with visually similar distractors. As a result, standard 3D diffusion policies often struggle to localize and exploit task-relevant geometry as scene complexity grows. We propose Attention-DP3, a spatially object-aware 3D diffusion policy that injects object-level geometric cues via attention while keeping the DP3 diffusion backbone unchanged. Our pipeline performs open-vocabulary 2D segmentation on RGB images, then lifts predicted target masks into 3D using calibrated camera geometry to obtain object-centric geometric priors. We incorporate these cues through Tri-field Attentional Conditioning, which constructs three complementary fields: (i) a targetness field to anchor the target object, (ii) an intra-target saliency field to emphasize task-relevant geometry within the target, and (iii) a backgroundness field to suppress distractors and clutter. Experiments on Adroit, DexArt, MetaWorld, and the real-world SO101 platform show consistent improvements over DP3, achieving state-of-the-art performance across benchmarks. Notably, as distractor objects increase, DP3 drops sharply, whereas Attention-DP3 remains stable and outperforms DP3 by up to 31\\% under heavy clutter. The code is publicly available at https://github.com/zhangzhongbo2213/Attention-DP3.","upvotes":1,"discussionId":"6aa8aca25dd4cb9b4cc02500","githubRepo":"https://github.com/zhangzhongbo2213/Attention-DP3","githubRepoAddedBy":"user","ai_summary":"Attention-DP3 improves 3D diffusion policies by injecting object-level geometric cues via attention to stabilize performance under heavy clutter.","ai_keywords":["3D diffusion policy","open-vocabulary 2D segmentation","Tri-field Attentional Conditioning","targetness field","intra-target saliency field","backgroundness field"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":4,"organization":{"_id":"68b830782b6b404b8318fe8e","name":"dalian-university-of-technology","fullname":"DaLian University of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68b82cf0116141793335f750/N5laKTgqcFB6x_i8kzdTY.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6575702b15b1ca184b0b2700","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6575702b15b1ca184b0b2700/O9cEodqQmG-gyqMiO_edR.jpeg","isPro":false,"fullname":"Zaibin Zhang","user":"MrBean2024","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68b830782b6b404b8318fe8e","name":"dalian-university-of-technology","fullname":"DaLian University of Technology","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68b82cf0116141793335f750/N5laKTgqcFB6x_i8kzdTY.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.13318.md","query":{}}">
Papers
arxiv:2609.13318

Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning

Published on Sep 10
· Submitted by
Zaibin Zhang
on Sep 15
Authors:
,

Abstract

Attention-DP3 improves 3D diffusion policies by injecting object-level geometric cues via attention to stabilize performance under heavy clutter.

3D point-cloud observations are inherently ambiguous in complex, cluttered manipulation scenes, where target objects may be partially occluded or tightly intermingled with visually similar distractors. As a result, standard 3D diffusion policies often struggle to localize and exploit task-relevant geometry as scene complexity grows. We propose Attention-DP3, a spatially object-aware 3D diffusion policy that injects object-level geometric cues via attention while keeping the DP3 diffusion backbone unchanged. Our pipeline performs open-vocabulary 2D segmentation on RGB images, then lifts predicted target masks into 3D using calibrated camera geometry to obtain object-centric geometric priors. We incorporate these cues through Tri-field Attentional Conditioning, which constructs three complementary fields: (i) a targetness field to anchor the target object, (ii) an intra-target saliency field to emphasize task-relevant geometry within the target, and (iii) a backgroundness field to suppress distractors and clutter. Experiments on Adroit, DexArt, MetaWorld, and the real-world SO101 platform show consistent improvements over DP3, achieving state-of-the-art performance across benchmarks. Notably, as distractor objects increase, DP3 drops sharply, whereas Attention-DP3 remains stable and outperforms DP3 by up to 31\% under heavy clutter. The code is publicly available at https://github.com/zhangzhongbo2213/Attention-DP3.

Community

Paper submitter about 6 hours ago

Attention-DP3 enhances 3D diffusion policies with object-aware geometric attention, improving target localization and robustness in cluttered scenes while keeping the DP3 backbone unchanged. It consistently outperforms DP3 across simulation and real-world benchmarks, with gains of up to 31% under heavy clutter. Attention-DP3 is accepted by ECCV 2026. Code is avaliable at https://github.com/zhangzhongbo2213/Attention-DP3.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.13318
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.13318 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.13318 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.13318 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers