Hugging Face Daily Papers · · 4 min read

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Project page: <a href=\"https://www.di.ens.fr/willow/research/panorama/\" rel=\"nofollow\">https://www.di.ens.fr/willow/research/panorama/</a><br> Code and data: <a href=\"https://github.com/sarapieri/panorama_grounding\" rel=\"nofollow\">https://github.com/sarapieri/panorama_grounding</a></p>\n","updatedAt":"2026-09-17T07:05:19.213Z","author":{"_id":"63fe2f588b3c5087ff8721bf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677602580834-noauth.jpeg","fullname":"Sara Pieri","name":"HuggingSara","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5038161277770996},"editors":["HuggingSara"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1677602580834-noauth.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.19143","authors":[{"_id":"6aab889e1d9cc4dec796259c","user":{"_id":"63fe2f588b3c5087ff8721bf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677602580834-noauth.jpeg","isPro":false,"fullname":"Sara Pieri","user":"HuggingSara","type":"user","name":"HuggingSara"},"name":"Sara Pieri","status":"claimed_verified","statusLastChangedAt":"2026-09-17T08:45:04.544Z","hidden":false},{"_id":"6aab889e1d9cc4dec796259d","user":{"_id":"62f38b19261bc5fb2e06652c","avatarUrl":"/avatars/0a7f7d63e1096f5d52bcb3be8e236c87.svg","isPro":false,"fullname":"Evangelos Kazakos","user":"ekazakos","type":"user","name":"ekazakos"},"name":"Evangelos Kazakos","status":"claimed_verified","statusLastChangedAt":"2026-09-17T08:45:04.551Z","hidden":false},{"_id":"6aab889e1d9cc4dec796259e","name":"Shizhe Chen","hidden":false},{"_id":"6aab889e1d9cc4dec796259f","name":"Josef Sivic","hidden":false},{"_id":"6aab889e1d9cc4dec79625a0","name":"Cordelia Schmid","hidden":false}],"publishedAt":"2026-09-16T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection","submittedOnDailyBy":{"_id":"63fe2f588b3c5087ff8721bf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677602580834-noauth.jpeg","isPro":false,"fullname":"Sara Pieri","user":"HuggingSara","type":"user","name":"HuggingSara"},"summary":"Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each referring phrase with pixel-level masks. We make three contributions. First, we introduce PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets. It provides dense captions with near-complete pixel coverage and image-text alignments at the entity level, supporting both training and evaluation. We further propose a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric that jointly evaluates textual and mask agreement. Second, we formulate phrase grounding as selection from a phrase-conditioned pool of mask proposals and introduce PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase. Training this interface jointly with caption generation enables PANORAMA to produce high-quality masks while allowing each phrase to refer to a single region or multiple instances. Third, PANORAMA achieves the best overall grounding on PanoCaps and matches or exceeds specialized models across several pixel-level grounding tasks. Experiments show that our method produces precise entity-level segmentations while maintaining detailed, mask-consistent captions. Code, data and models are available at https://www.di.ens.fr/willow/research/panorama/.","upvotes":9,"discussionId":"6aab889e1d9cc4dec79625a1","projectPage":"https://www.di.ens.fr/willow/research/panorama/","githubRepo":"https://github.com/sarapieri/panorama_grounding","githubRepoAddedBy":"user","githubStars":2,"organization":{"_id":"6a58d875b35224774a679f87","name":"Panorama-grounding","fullname":"Panorama-grounding","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63fe2f588b3c5087ff8721bf/JVObhFk5FtmytIoXMtnRd.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62f38b19261bc5fb2e06652c","avatarUrl":"/avatars/0a7f7d63e1096f5d52bcb3be8e236c87.svg","isPro":false,"fullname":"Evangelos Kazakos","user":"ekazakos","type":"user"},{"_id":"653b7be1dfce99b57b5e7a8e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/XebTuoK-LzP7e8UmVKmKR.jpeg","isPro":false,"fullname":"Javier Alejandro Lopetegui Gonzalez","user":"JavierLopetegui","type":"user"},{"_id":"67f67663c91d3b03d3818390","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ygOP7_lB_qVC3rsK3ME1X.png","isPro":false,"fullname":"Federica Spinola","user":"zampino1234","type":"user"},{"_id":"6309b6f511f104451667fb0f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6309b6f511f104451667fb0f/j-dCYceZeHFqLBVt_EjYg.png","isPro":false,"fullname":"Duc Hai","user":"haiphamcse","type":"user"},{"_id":"63fe2f588b3c5087ff8721bf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677602580834-noauth.jpeg","isPro":false,"fullname":"Sara Pieri","user":"HuggingSara","type":"user"},{"_id":"63d067a58cf6c8e4d20a57e8","avatarUrl":"/avatars/6ea9414859f4600f618547c13ed6faa6.svg","isPro":false,"fullname":"gabriel_fstr","user":"gabzouz37","type":"user"},{"_id":"668e6b47f59574a8ec2ae078","avatarUrl":"/avatars/1cbc80ee4fb4a832783bd3dbee032d6e.svg","isPro":false,"fullname":"Zeeshan Khan","user":"zk95","type":"user"},{"_id":"69c4034c98f2e0d9a031cb59","avatarUrl":"/avatars/cbc811c5d9dd0624e546300f49f796e6.svg","isPro":false,"fullname":"Bora","user":"ubgk","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a58d875b35224774a679f87","name":"Panorama-grounding","fullname":"Panorama-grounding","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63fe2f588b3c5087ff8721bf/JVObhFk5FtmytIoXMtnRd.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.19143.md","query":{}}">
Papers
arxiv:2609.19143

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Published on Sep 16
· Submitted by
Sara Pieri
on Sep 17

Abstract

Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each referring phrase with pixel-level masks. We make three contributions. First, we introduce PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets. It provides dense captions with near-complete pixel coverage and image-text alignments at the entity level, supporting both training and evaluation. We further propose a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric that jointly evaluates textual and mask agreement. Second, we formulate phrase grounding as selection from a phrase-conditioned pool of mask proposals and introduce PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase. Training this interface jointly with caption generation enables PANORAMA to produce high-quality masks while allowing each phrase to refer to a single region or multiple instances. Third, PANORAMA achieves the best overall grounding on PanoCaps and matches or exceeds specialized models across several pixel-level grounding tasks. Experiments show that our method produces precise entity-level segmentations while maintaining detailed, mask-consistent captions. Code, data and models are available at https://www.di.ens.fr/willow/research/panorama/.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.19143
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.19143 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.19143 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers