Hugging Face Daily Papers · · 4 min read

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Can VLA models go beyond simple scenes and short-horizon tasks?</p>\n<p>We introduce RoboSPA, a large-scale diagnostic benchmark for evaluating VLA models on fine-grained spatial reasoning and long-horizon procedural planning. It contains 56 tasks across 10 capability categories, 5 difficulty levels, 5 robotic embodiments, and 527K+ trajectories.</p>\n<p>Evaluating RDT, GO-1, π0.5, and X-VLA reveals a clear capability gap: performance drops sharply as spatial ambiguity and task horizon increase, with all models achieving &lt;25% average success at the highest difficulty level.</p>\n<p>RoboSPA further supports step-level evaluation and failure diagnosis, enabling more fine-grained analysis beyond binary task success.</p>\n","updatedAt":"2026-09-09T06:27:53.083Z","author":{"_id":"6a881a92b8a4416fb8cf6bb4","avatarUrl":"/avatars/5fe99901bc6ce570caa6bf29ffe1da41.svg","fullname":"Zhenxuan Fan","name":"zxfan","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8563477993011475},"editors":["zxfan"],"editorAvatarUrls":["/avatars/5fe99901bc6ce570caa6bf29ffe1da41.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.05324","authors":[{"_id":"6a9fbd686c8e10537d563c37","user":{"_id":"6a881a92b8a4416fb8cf6bb4","avatarUrl":"/avatars/5fe99901bc6ce570caa6bf29ffe1da41.svg","isPro":false,"fullname":"Zhenxuan Fan","user":"zxfan","type":"user","name":"zxfan"},"name":"Zhenxuan Fan","status":"claimed_verified","statusLastChangedAt":"2026-09-08T16:45:04.207Z","hidden":false},{"_id":"6a9fbd686c8e10537d563c38","name":"Bo Zhang","hidden":false},{"_id":"6a9fbd686c8e10537d563c39","name":"Yutong Lin","hidden":false},{"_id":"6a9fbd686c8e10537d563c3a","name":"Yuqian Yuan","hidden":false},{"_id":"6a9fbd686c8e10537d563c3b","user":{"_id":"67fe5ff8957e1683eb54d05a","avatarUrl":"/avatars/fce87a8b6f84b48c6bc83e134fa9c131.svg","isPro":false,"fullname":"JueJue","user":"JackieLin0123","type":"user","name":"JackieLin0123"},"name":"Juekai Lin","status":"claimed_verified","statusLastChangedAt":"2026-09-09T09:23:13.167Z","hidden":false},{"_id":"6a9fbd686c8e10537d563c3c","name":"Liang Liang","hidden":false},{"_id":"6a9fbd686c8e10537d563c3d","name":"Zhuoyi Huang","hidden":false},{"_id":"6a9fbd686c8e10537d563c3e","name":"Wenqiao Zhang","hidden":false},{"_id":"6a9fbd686c8e10537d563c3f","name":"Juncheng Li","hidden":false},{"_id":"6a9fbd686c8e10537d563c40","name":"Siliang Tang","hidden":false},{"_id":"6a9fbd686c8e10537d563c41","name":"Jun Xiao","hidden":false},{"_id":"6a9fbd686c8e10537d563c42","name":"Yueting Zhuang","hidden":false}],"publishedAt":"2026-09-04T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?","submittedOnDailyBy":{"_id":"6a881a92b8a4416fb8cf6bb4","avatarUrl":"/avatars/5fe99901bc6ce570caa6bf29ffe1da41.svg","isPro":false,"fullname":"Zhenxuan Fan","user":"zxfan","type":"user","name":"zxfan"},"summary":"Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.","upvotes":20,"discussionId":"6a9fbd686c8e10537d563c43","projectPage":"https://fanzhenxuan.github.io/RoboSPA/","githubRepo":"https://github.com/fanzhenxuan/RoboSPA","githubRepoAddedBy":"user","ai_summary":"RoboSPA is a large-scale robotic manipulation benchmark that evaluates vision-language-action models on fine-grained spatial reasoning and long-horizon procedural planning across progressively harder task variants.","ai_keywords":["Vision-Language-Action models","embodied reasoning","spatial reasoning","procedural planning","diagnostic metrics"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":7,"organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a881a92b8a4416fb8cf6bb4","avatarUrl":"/avatars/5fe99901bc6ce570caa6bf29ffe1da41.svg","isPro":false,"fullname":"Zhenxuan Fan","user":"zxfan","type":"user"},{"_id":"67fe5ff8957e1683eb54d05a","avatarUrl":"/avatars/fce87a8b6f84b48c6bc83e134fa9c131.svg","isPro":false,"fullname":"JueJue","user":"JackieLin0123","type":"user"},{"_id":"6729bd9184e2468ea7e654f1","avatarUrl":"/avatars/a7840f2a07d08ad332e4df43473c2ab5.svg","isPro":false,"fullname":"lyt","user":"lzyrt","type":"user"},{"_id":"6912db52e5d44d45c33b46cc","avatarUrl":"/avatars/caf6ef66cd4effadc8e9eea56e385896.svg","isPro":false,"fullname":"shadow","user":"5ha0w","type":"user"},{"_id":"672378680fc8168c68cc3689","avatarUrl":"/avatars/8962db8b254acec101ad434afc07db2c.svg","isPro":false,"fullname":"Fzx","user":"FFFFFzx","type":"user"},{"_id":"675a569f7b01fc742a5adf93","avatarUrl":"/avatars/63f2be4089871bbb33f1edbd9d74eaa7.svg","isPro":false,"fullname":"Austin Reed","user":"AustinReed","type":"user"},{"_id":"6a6c7b702b8f6bcb1b61caba","avatarUrl":"/avatars/40259c5b4eb7b81d54b23737398a112f.svg","isPro":false,"fullname":"Sarah Clark","user":"sarah-clark","type":"user"},{"_id":"6a6c8bf37e229e8df66886be","avatarUrl":"/avatars/335caa28ef1d0af046d2ff2320fe5cd1.svg","isPro":false,"fullname":"Paul Martin","user":"zenithstack","type":"user"},{"_id":"6a6dc65ff9134eddf85b68ef","avatarUrl":"/avatars/80a269650c8b106ac8510bb55e54ffaf.svg","isPro":false,"fullname":"Richard Brown","user":"Meridian-Kai","type":"user"},{"_id":"6a6def554b31984754eba784","avatarUrl":"/avatars/4a0d31da2a5e59c90962fe99e0563a13.svg","isPro":false,"fullname":"Charles Harris","user":"charles-harris","type":"user"},{"_id":"6a9b502fde9d5763fb0507cb","avatarUrl":"/avatars/a73fb800f18ee6fca84928b2d605ea63.svg","isPro":false,"fullname":"胥桂芝","user":"qiangjiang21","type":"user"},{"_id":"6aa0e1feb5a40531f054767b","avatarUrl":"/avatars/65ca2e7bf2940f3f43d2a8060c754845.svg","isPro":false,"fullname":"Thomas Flynn","user":"Cobalt-Rin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61bac2af530e5c78d7b99667","name":"zju","fullname":"Zhejiang University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5e1058e9fcf41d740b69966d/7G1xjlxwCdMEmKcxNR0n5.png"},"query":{}}">
Papers
arxiv:2609.05324

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

Published on Sep 4
· Submitted by
Zhenxuan Fan
on Sep 9
Authors:

Abstract

RoboSPA is a large-scale robotic manipulation benchmark that evaluates vision-language-action models on fine-grained spatial reasoning and long-horizon procedural planning across progressively harder task variants.

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two core dimensions, Fine-Grained Spatial Reasoning and Long-Horizon Procedural Planning, covering 10 task categories and 56 base tasks. Each task is instantiated across five difficulty levels, yielding 280 variants with increasing spatial ambiguity and procedural complexity. We collect 527K trajectories across multiple embodiments and diverse scenes. Beyond binary success rate, RoboSPA introduces diagnostic metrics for more detailed evaluation. Experiments on representative VLA models show that current systems still struggle with complex spatial relations, precise low-level execution, and memory-intensive planning. These results establish RoboSPA as a challenging diagnostic benchmark for developing more capable, reliable, and generalizable embodied agents. Our data and code are available at https://github.com/fanzhenxuan/RoboSPA.

Community

Paper author Paper submitter about 8 hours ago

Can VLA models go beyond simple scenes and short-horizon tasks?

We introduce RoboSPA, a large-scale diagnostic benchmark for evaluating VLA models on fine-grained spatial reasoning and long-horizon procedural planning. It contains 56 tasks across 10 capability categories, 5 difficulty levels, 5 robotic embodiments, and 527K+ trajectories.

Evaluating RDT, GO-1, π0.5, and X-VLA reveals a clear capability gap: performance drops sharply as spatial ambiguity and task horizon increase, with all models achieving <25% average success at the highest difficulty level.

RoboSPA further supports step-level evaluation and failure diagnosis, enabling more fine-grained analysis beyond binary task success.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.05324 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.05324 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.05324 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers