Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.</p>\n","updatedAt":"2026-07-31T03:27:26.139Z","author":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","fullname":"Yang Shi","name":"DogNeverSleep","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":14,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8145877718925476},"editors":["DogNeverSleep"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28509","authors":[{"_id":"6a6c1606202e2d9e3ffdb772","name":"Tengfei Liu","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb773","user":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user","name":"DogNeverSleep"},"name":"Yang Shi","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.606Z","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb774","name":"Yuran Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb775","name":"Xiaohan Zhang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb776","name":"Yuqing Wen","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb777","name":"Yuqi Tang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb778","name":"Qixun Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb779","name":"Zhuoran Zhang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77a","name":"Xuanyu Zhu","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77b","name":"Weihong Lin","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77c","name":"Xinlei Yu","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77d","name":"Yujie Wei","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77e","name":"Xinwei Long","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb77f","name":"Fengxiang Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb780","name":"Xinlong Chen","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb781","name":"Yue Ding","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb782","name":"Jialu Chen","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb783","name":"Haotian Wang","hidden":false},{"_id":"6a6c1606202e2d9e3ffdb784","name":"Yuanxing Zhang","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"RefCaptioner: Multi-Reference Image-Grounded Video Captioning","submittedOnDailyBy":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user","name":"DogNeverSleep"},"summary":"Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.","upvotes":20,"discussionId":"6a6c1606202e2d9e3ffdb785","githubRepo":"https://github.com/pkucs-Ltf/RefCaptioner","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"662c559b322afcbae51b3c8b","name":"KlingTeam","fullname":"Kling Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60e272ca6c78a8c122b12127/ZQV1aKLUDPf2rUcxxAqj6.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user"},{"_id":"67d63e228d5c7a132cbcf39b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png","isPro":false,"fullname":"neil yu","user":"yxl66666","type":"user"},{"_id":"667e577139b49eba118d569f","avatarUrl":"/avatars/1a26dd96b4b352b8968561750ecae9a7.svg","isPro":false,"fullname":"Xinwei Long","user":"xinwei666","type":"user"},{"_id":"6700b2b6bff0e8b51d07fa00","avatarUrl":"/avatars/6cd7e243b7bc37ae9d308c175cbe6f05.svg","isPro":false,"fullname":"asdasd","user":"asdjghh","type":"user"},{"_id":"661cd9a47c7339263b11d71a","avatarUrl":"/avatars/4ca6ea300a010b60c6a51792f92a0538.svg","isPro":false,"fullname":"jacuzzi","user":"2kxx","type":"user"},{"_id":"69a6eaebc2ba11d926d4ac61","avatarUrl":"/avatars/6cec959b72c65ceed2dd5232752f2a16.svg","isPro":false,"fullname":"Leo","user":"Fakerookie","type":"user"},{"_id":"68be5e706da9b1623542c717","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/N0Zh1lvLYQhhhXp5XYRG9.png","isPro":false,"fullname":"sikichen","user":"sikickchen","type":"user"},{"_id":"637f70d6fab5db9101c3dfc8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637f70d6fab5db9101c3dfc8/NgkYNXWLDavLbrnCby2Fl.jpeg","isPro":false,"fullname":"Yujie Wei","user":"weilllllls","type":"user"},{"_id":"68803774326d963ec15ac76c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68803774326d963ec15ac76c/ZSb23DRhjH3Mjf17K2zyz.jpeg","isPro":false,"fullname":"Wanshun Su","user":"RowanSu","type":"user"},{"_id":"66d43d619e358c92da0adacc","avatarUrl":"/avatars/f878387e2d778ce915b6bd52b192f30c.svg","isPro":false,"fullname":"Zhou","user":"Descartyes","type":"user"},{"_id":"644d2532d185572dd1e48f90","avatarUrl":"/avatars/5831acebb02d8bc8f80f56b7b11c7c69.svg","isPro":false,"fullname":"Zhu","user":"zzzhu","type":"user"},{"_id":"653fd9d807faf7b0bc57cda4","avatarUrl":"/avatars/7efe7f095270927a0f13e5b72a0d231a.svg","isPro":false,"fullname":"HouHuaWei","user":"HouHuaWei","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"662c559b322afcbae51b3c8b","name":"KlingTeam","fullname":"Kling Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60e272ca6c78a8c122b12127/ZQV1aKLUDPf2rUcxxAqj6.jpeg"},"query":{}}">
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Abstract
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.
Community
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.28509 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.