Hugging Face Daily Papers · · 4 min read

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Spatial intelligence is often easier to show than tell. We introduce ProVisE, a benchmark-agnostic evaluation framework that enables image-generation models to answer spatial tasks directly in pixels. ProVisE converts these visual responses into the structured predictions required by each benchmark and scores them using its original metrics. Using SpatialGen-Bench (470 samples, 14 subtasks) and six external benchmarks, we evaluate 31 text- and image-answering systems. The two interfaces reveal complementary strengths: visual answering performs particularly well when spatial states can be expressed directly in pixels, while text VLMs remain stronger on compositional reasoning tasks. Notably, GPT Image 2 correctly solves 37% of the spatial cases missed by GPT-5.4.</p>\n","updatedAt":"2026-07-24T04:41:29.041Z","author":{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","fullname":"Wenqi Zhang","name":"zwq2018","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":14,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8618505597114563},"editors":["zwq2018"],"editorAvatarUrls":["/avatars/1427cf1a72b5db0cb263ad45885cf925.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.21072","authors":[{"_id":"6a62c5f52ee212ed0e2a13f9","user":{"_id":"690f009eb9a679e969ece716","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/690f009eb9a679e969ece716/tDU-tAo1JyHKo__e7PLgv.jpeg","isPro":false,"fullname":"xuwang","user":"wx91726","type":"user","name":"wx91726"},"name":"Xu Wang","status":"claimed_verified","statusLastChangedAt":"2026-07-24T08:45:04.294Z","hidden":false},{"_id":"6a62c5f52ee212ed0e2a13fa","name":"Kaixiang Yao","hidden":false},{"_id":"6a62c5f52ee212ed0e2a13fb","name":"Miao Pan","hidden":false},{"_id":"6a62c5f52ee212ed0e2a13fc","name":"Xiaohe Zhou","hidden":false},{"_id":"6a62c5f52ee212ed0e2a13fd","name":"Xuanyu Liu","hidden":false},{"_id":"6a62c5f52ee212ed0e2a13fe","name":"Wenqi Zhang","hidden":false},{"_id":"6a62c5f52ee212ed0e2a13ff","name":"Xuhong Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6485bd278d14bcd5cdbb7c8d/1bC31hrpClCXzgEN_SwSr.mp4"],"publishedAt":"2026-07-23T00:00:00.000Z","submittedOnDailyAt":"2026-07-24T00:00:00.000Z","title":"Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text","submittedOnDailyBy":{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","isPro":false,"fullname":"Wenqi Zhang","user":"zwq2018","type":"user","name":"zwq2018"},"summary":"Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.","upvotes":34,"discussionId":"6a62c5f62ee212ed0e2a1400","projectPage":"https://zju-omniai.github.io/ProVisE/","githubRepo":"https://github.com/ZJU-OmniAI/ProVisE","githubRepoAddedBy":"user","githubStars":16,"organization":{"_id":"696461ab2d94e9a07cdb8efd","name":"OmniAI-ZJU","fullname":"ZJU-OmniAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f2595830354d0ee043b25a/eEeRdHlGyJ148JQ6BAJ4O.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6485bd278d14bcd5cdbb7c8d","avatarUrl":"/avatars/1427cf1a72b5db0cb263ad45885cf925.svg","isPro":false,"fullname":"Wenqi Zhang","user":"zwq2018","type":"user"},{"_id":"69367c5e39abc914b3bbb2e3","avatarUrl":"/avatars/8ed0e0147bcc9c77040ad490ae8a107c.svg","isPro":false,"fullname":"Yi Pan","user":"Yi0304","type":"user"},{"_id":"67e3de6b5269fa5ef3f2bd0d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/2QfglNGb0keGsllfHaS7Q.png","isPro":false,"fullname":"sfywtdiy","user":"xnxjiu","type":"user"},{"_id":"68ca385a049f422eb68ac0a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/o0XgN0ebrQ2FAfyAh_6XF.png","isPro":false,"fullname":"kagakouko","user":"kagakouko","type":"user"},{"_id":"67543820c3af453d7b3e1d5e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67543820c3af453d7b3e1d5e/RbAZ9AQlxpy5E5Is-QN8b.jpeg","isPro":false,"fullname":"Dingming Li","user":"lidingm","type":"user"},{"_id":"6a63050b9800990ccd5da399","avatarUrl":"/avatars/e0fb2799d04535ee63363eb35291edd0.svg","isPro":false,"fullname":"wang jian min","user":"wjm7890","type":"user"},{"_id":"67dac019f1fbcfd23a623805","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7SwomV4nqu1QrJ4Bv9Z90.png","isPro":false,"fullname":"huangyouhao","user":"freshfishFF","type":"user"},{"_id":"6a630b442c9b97259e2c694f","avatarUrl":"/avatars/70d3916eff06f3895a6f95b3e6065e55.svg","isPro":false,"fullname":"ZeZhong Cheng","user":"KHecho","type":"user"},{"_id":"6358b570aff68f72ac06113b","avatarUrl":"/avatars/c1df6bab08bc1a940b0af88a3f1e58ad.svg","isPro":false,"fullname":"DawnChase","user":"DawnChase","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6485e686c7f19728a4a49b68","avatarUrl":"/avatars/8172c56c0e01163e190af3ea614e7449.svg","isPro":false,"fullname":"mayanna","user":"stena303","type":"user"},{"_id":"6690eccd6dc63c03461a1db7","avatarUrl":"/avatars/e58326dea123790e26b4769aff84c72b.svg","isPro":false,"fullname":"mayanna","user":"Nana303","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"696461ab2d94e9a07cdb8efd","name":"OmniAI-ZJU","fullname":"ZJU-OmniAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65f2595830354d0ee043b25a/eEeRdHlGyJ148JQ6BAJ4O.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.21072.md","query":{}}">
Papers
arxiv:2607.21072

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Published on Jul 23
· Submitted by
Wenqi Zhang
on Jul 24
Authors:

Abstract

Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.

Community

Paper submitter about 15 hours ago

Spatial intelligence is often easier to show than tell. We introduce ProVisE, a benchmark-agnostic evaluation framework that enables image-generation models to answer spatial tasks directly in pixels. ProVisE converts these visual responses into the structured predictions required by each benchmark and scores them using its original metrics. Using SpatialGen-Bench (470 samples, 14 subtasks) and six external benchmarks, we evaluate 31 text- and image-answering systems. The two interfaces reveal complementary strengths: visual answering performs particularly well when spatial states can be expressed directly in pixels, while text VLMs remain stronger on compositional reasoning tasks. Notably, GPT Image 2 correctly solves 37% of the spatial cases missed by GPT-5.4.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.21072
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.21072 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.21072 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers