Hugging Face Daily Papers · · 3 min read

Vision Pretraining for Dense Spatial Perception

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A novel SSL pretraining for spatial perception with a STRONG depth model, LingBot-Depth 2.0</p>\n","updatedAt":"2026-07-07T03:29:13.887Z","author":{"_id":"6485ce5ec7f19728a49df17a","avatarUrl":"/avatars/e83966e6906c1d0f151300981e30f85a.svg","fullname":"Nan","name":"cherubicxn","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7885457277297974},"editors":["cherubicxn"],"editorAvatarUrls":["/avatars/e83966e6906c1d0f151300981e30f85a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.05247","authors":[{"_id":"6a4c71ad25849b193a83410f","name":"Zelin Fu","hidden":false},{"_id":"6a4c71ad25849b193a834110","name":"Bin Tan","hidden":false},{"_id":"6a4c71ad25849b193a834111","name":"Changjiang Sun","hidden":false},{"_id":"6a4c71ad25849b193a834112","name":"Shaohui Liu","hidden":false},{"_id":"6a4c71ad25849b193a834113","name":"Kecheng Zheng","hidden":false},{"_id":"6a4c71ad25849b193a834114","name":"Yinghao Xu","hidden":false},{"_id":"6a4c71ad25849b193a834115","name":"Xing Zhu","hidden":false},{"_id":"6a4c71ad25849b193a834116","name":"Yujun Shen","hidden":false},{"_id":"6a4c71ad25849b193a834117","user":{"_id":"6485ce5ec7f19728a49df17a","avatarUrl":"/avatars/e83966e6906c1d0f151300981e30f85a.svg","isPro":true,"fullname":"Nan","user":"cherubicxn","type":"user","name":"cherubicxn"},"name":"Nan Xue","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:11:39.054Z","hidden":false}],"publishedAt":"2026-07-06T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"Vision Pretraining for Dense Spatial Perception","submittedOnDailyBy":{"_id":"6485ce5ec7f19728a49df17a","avatarUrl":"/avatars/e83966e6906c1d0f151300981e30f85a.svg","isPro":true,"fullname":"Nan","user":"cherubicxn","type":"user","name":"cherubicxn"},"summary":"Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.","upvotes":30,"discussionId":"6a4c71ae25849b193a834118","projectPage":"https://technology.robbyant.com/lingbot-vision","githubRepo":"https://github.com/Robbyant/lingbot-vision","githubRepoAddedBy":"user","ai_summary":"Boundary modeling enables dense spatial perception by learning sub-pixel representations that enhance depth estimation and support embodied AI applications.","ai_keywords":["boundary modeling","masked boundary modeling","visual foundation models","spatial perception","depth completion","embodied artificial intelligence","DINOv3","dense visual token learning"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":290,"organization":{"_id":"69709f892cd08371c1011a2e","name":"robbyant","fullname":"Robbyant","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/ZTuImney4XzRmBHyUL47F.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6485ce5ec7f19728a49df17a","avatarUrl":"/avatars/e83966e6906c1d0f151300981e30f85a.svg","isPro":true,"fullname":"Nan","user":"cherubicxn","type":"user"},{"_id":"67aeffda7330db26f93cd62f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/ktdFt5_F05qNJ-dnD_Fk3.jpeg","isPro":false,"fullname":"jiangbonadia","user":"NadiaJiang","type":"user"},{"_id":"64252045a4f3051f54dd1d53","avatarUrl":"/avatars/0e423a3291091be3b4736a14da3ce495.svg","isPro":false,"fullname":"kecheng zheng","user":"zkcys001","type":"user"},{"_id":"64acd2ec39fcfebff8c79c00","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64acd2ec39fcfebff8c79c00/Avq66l5hO-aggNtk4Y1ss.png","isPro":false,"fullname":"Ka Leong Cheng","user":"felixcheng97","type":"user"},{"_id":"6a325017c06fb4e3163cc88d","avatarUrl":"/avatars/d7cbb528870f0ed31877a3696cfb52a5.svg","isPro":false,"fullname":"zhujiayi","user":"Posii001","type":"user"},{"_id":"697c76d2e9390e622b45c553","avatarUrl":"/avatars/83d5487d4fcab087366ded6e2663534e.svg","isPro":false,"fullname":"BlackLing","user":"BlackLing02","type":"user"},{"_id":"65250cda87ad4a39f84d482d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65250cda87ad4a39f84d482d/ryzc7hx6x09DbL4eiEBip.jpeg","isPro":false,"fullname":"Feng Chaoran","user":"Falcary","type":"user"},{"_id":"66f69b317024ed3df973d245","avatarUrl":"/avatars/35939450ae70dc7b641a209e64265dfc.svg","isPro":false,"fullname":"Jingjing Wang","user":"whalejj","type":"user"},{"_id":"63d4b843df01ef426a0f79fb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676365795587-63d4b843df01ef426a0f79fb.jpeg","isPro":false,"fullname":"Yanhong Zeng","user":"zengyh1900","type":"user"},{"_id":"68a2fb20b3ad3d518527c49d","avatarUrl":"/avatars/cb37f5b2a98cbbea6f9726552a8f60e2.svg","isPro":false,"fullname":"zq","user":"capturee","type":"user"},{"_id":"69c3c9314f60bf51a1cf57ef","avatarUrl":"/avatars/127df590b4e3f3fae2b07b5b03dbd80b.svg","isPro":false,"fullname":"xl","user":"jarstic","type":"user"},{"_id":"66f3b9d4fe3b8ef090ba0284","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f3b9d4fe3b8ef090ba0284/HuTcpFMMctnK_fb6fe25C.png","isPro":false,"fullname":"Nils DEYBACH","user":"nils-deybach","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69709f892cd08371c1011a2e","name":"robbyant","fullname":"Robbyant","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aeffda7330db26f93cd62f/ZTuImney4XzRmBHyUL47F.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.05247.md","query":{}}">
Papers
arxiv:2607.05247

Vision Pretraining for Dense Spatial Perception

Published on Jul 6
· Submitted by
Nan
on Jul 7
Authors:
,
,
,
,
,
,
,
,

Abstract

Boundary modeling enables dense spatial perception by learning sub-pixel representations that enhance depth estimation and support embodied AI applications.

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observations. Modern visual foundation models tend to prioritize semantic invariance, often at the expense of detailed spatial understanding. In this work, we study vision pretraining through a boundary-centric lens, motivated by the premise that boundaries and shape discontinuities offer essential cues for perceiving geometric properties. Concretely, we propose masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning. By scaling this framework, we develop LingBot-Vision and demonstrate its efficacy across a diverse set of downstream vision tasks with DINOv3 as a strong baseline. Remarkably, LingBot-Vision drives the progression from LingBot-Depth 1.0 to LingBot-Depth 2.0 for depth completion, and thereby yields enhanced depth estimation, a key pillar for embodied artificial intelligence. Our findings reveal that boundary modeling goes beyond simple line segments and instead serves as a scalable pretraining principle for learning spatially structured visual representations.

Community

Paper author Paper submitter about 19 hours ago

A novel SSL pretraining for spatial perception with a STRONG depth model, LingBot-Depth 2.0

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.05247
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.05247 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.05247 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.05247 in a Space README.md to link it from this page.

Collections including this paper 1

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers