Hugging Face Daily Papers · · 4 min read

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.</p>\n","updatedAt":"2026-07-31T03:27:26.139Z","author":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","fullname":"Yang Shi","name":"DogNeverSleep","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":14,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8145877718925476},"editors":["DogNeverSleep"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28509","authors":[{"_id":"6a6c1606202e2d9e3ffdb772","name":"Tengfei Liu","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb773","user":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user","name":"DogNeverSleep"},"name":"Yang Shi","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.606Z","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb774","name":"Yuran Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb775","name":"Xiaohan Zhang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb776","name":"Yuqing Wen","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb777","name":"Yuqi Tang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb778","name":"Qixun Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb779","name":"Zhuoran Zhang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77a","name":"Xuanyu Zhu","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77b","name":"Weihong Lin","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77c","name":"Xinlei Yu","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77d","name":"Yujie Wei","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77e","name":"Xinwei Long","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77f","name":"Fengxiang Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb780","name":"Xinlong Chen","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb781","name":"Yue Ding","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb782","name":"Jialu Chen","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb783","name":"Haotian Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb784","name":"Yuanxing Zhang","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"RefCaptioner: Multi-Reference Image-Grounded Video Captioning","submittedOnDailyBy":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user","name":"DogNeverSleep"},"summary":"Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.","upvotes":20,"discussionId":"6a6c1606202e2d9e3ffdb785","githubRepo":"https://github.com/pkucs-Ltf/RefCaptioner","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"662c559b322afcbae51b3c8b","name":"KlingTeam","fullname":"Kling Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60e272ca6c78a8c122b12127/ZQV1aKLUDPf2rUcxxAqj6.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user"},{"_id":"67d63e228d5c7a132cbcf39b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png","isPro":false,"fullname":"neil yu","user":"yxl66666","type":"user"},{"_id":"667e577139b49eba118d569f","avatarUrl":"/avatars/1a26dd96b4b352b8968561750ecae9a7.svg","isPro":false,"fullname":"Xinwei Long","user":"xinwei666","type":"user"},{"_id":"6700b2b6bff0e8b51d07fa00","avatarUrl":"/avatars/6cd7e243b7bc37ae9d308c175cbe6f05.svg","isPro":false,"fullname":"asdasd","user":"asdjghh","type":"user"},{"_id":"661cd9a47c7339263b11d71a","avatarUrl":"/avatars/4ca6ea300a010b60c6a51792f92a0538.svg","isPro":false,"fullname":"jacuzzi","user":"2kxx","type":"user"},{"_id":"69a6eaebc2ba11d926d4ac61","avatarUrl":"/avatars/6cec959b72c65ceed2dd5232752f2a16.svg","isPro":false,"fullname":"Leo","user":"Fakerookie","type":"user"},{"_id":"68be5e706da9b1623542c717","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/N0Zh1lvLYQhhhXp5XYRG9.png","isPro":false,"fullname":"sikichen","user":"sikickchen","type":"user"},{"_id":"637f70d6fab5db9101c3dfc8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637f70d6fab5db9101c3dfc8/NgkYNXWLDavLbrnCby2Fl.jpeg","isPro":false,"fullname":"Yujie Wei","user":"weilllllls","type":"user"},{"_id":"68803774326d963ec15ac76c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68803774326d963ec15ac76c/ZSb23DRhjH3Mjf17K2zyz.jpeg","isPro":false,"fullname":"Wanshun Su","user":"RowanSu","type":"user"},{"_id":"66d43d619e358c92da0adacc","avatarUrl":"/avatars/f878387e2d778ce915b6bd52b192f30c.svg","isPro":false,"fullname":"Zhou","user":"Descartyes","type":"user"},{"_id":"644d2532d185572dd1e48f90","avatarUrl":"/avatars/5831acebb02d8bc8f80f56b7b11c7c69.svg","isPro":false,"fullname":"Zhu","user":"zzzhu","type":"user"},{"_id":"653fd9d807faf7b0bc57cda4","avatarUrl":"/avatars/7efe7f095270927a0f13e5b72a0d231a.svg","isPro":false,"fullname":"HouHuaWei","user":"HouHuaWei","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"662c559b322afcbae51b3c8b","name":"KlingTeam","fullname":"Kling Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60e272ca6c78a8c122b12127/ZQV1aKLUDPf2rUcxxAqj6.jpeg"},"query":{}}">
Papers
arxiv:2607.28509

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Published on Jul 30
· Submitted by
Yang Shi
on Jul 31
Authors:
,

Abstract

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.

Community

Paper author Paper submitter about 7 hours ago

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.28509 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers