Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.</p>\n","updatedAt":"2026-07-17T03:18:39.262Z","author":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","fullname":"Yang Shi","name":"DogNeverSleep","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":14,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.867752194404602},"editors":["DogNeverSleep"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.14202","authors":[{"_id":"6a599ee36c2e371e6ca3813b","name":"Yuqi Tang","hidden":false},{"_id":"6a599ee36c2e371e6ca3813c","name":"Tengfei Liu","hidden":false},{"_id":"6a599ee36c2e371e6ca3813d","name":"Yizheng Lai","hidden":false},{"_id":"6a599ee36c2e371e6ca3813e","name":"Yuran Wang","hidden":false},{"_id":"6a599ee36c2e371e6ca3813f","name":"Yang Shi","hidden":false},{"_id":"6a599ee36c2e371e6ca38140","name":"Wanshun Su","hidden":false},{"_id":"6a599ee36c2e371e6ca38141","name":"Zhuoran Zhang","hidden":false},{"_id":"6a599ee36c2e371e6ca38142","name":"Qixun Wang","hidden":false},{"_id":"6a599ee36c2e371e6ca38143","name":"Xiaohan Zhang","hidden":false},{"_id":"6a599ee36c2e371e6ca38144","name":"Xinlei Yu","hidden":false},{"_id":"6a599ee36c2e371e6ca38145","name":"Xuehai Bai","hidden":false},{"_id":"6a599ee36c2e371e6ca38146","name":"Xuanyu Zhu","hidden":false},{"_id":"6a599ee36c2e371e6ca38147","name":"Bohan Zeng","hidden":false},{"_id":"6a599ee36c2e371e6ca38148","name":"Bozhou Li","hidden":false},{"_id":"6a599ee36c2e371e6ca38149","name":"Shujie Li","hidden":false},{"_id":"6a599ee36c2e371e6ca3814a","name":"Yifan Dai","hidden":false},{"_id":"6a599ee36c2e371e6ca3814b","name":"Yujie Wei","hidden":false},{"_id":"6a599ee36c2e371e6ca3814c","name":"Shixuan Liu","hidden":false},{"_id":"6a599ee36c2e371e6ca3814d","name":"Haotian Wang","hidden":false},{"_id":"6a599ee36c2e371e6ca3814e","name":"Jialu Chen","hidden":false},{"_id":"6a599ee36c2e371e6ca3814f","name":"Yuanxing Zhang","hidden":false}],"publishedAt":"2026-07-15T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation","submittedOnDailyBy":{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user","name":"DogNeverSleep"},"summary":"Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.","upvotes":30,"discussionId":"6a599ee36c2e371e6ca38150","githubRepo":"https://github.com/cactusqq/KeyFrame-Compass","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"662c559b322afcbae51b3c8b","name":"KlingTeam","fullname":"Kling Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60e272ca6c78a8c122b12127/ZQV1aKLUDPf2rUcxxAqj6.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"673c7319d11b1c2e246ead9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673c7319d11b1c2e246ead9c/IjFIO--N7Hm_BOEafhEQv.jpeg","isPro":false,"fullname":"Yang Shi","user":"DogNeverSleep","type":"user"},{"_id":"67d63e228d5c7a132cbcf39b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png","isPro":false,"fullname":"neil yu","user":"yxl66666","type":"user"},{"_id":"6671214c92412fd4640714eb","avatarUrl":"/avatars/48fa84e7bc3bb92ad0192aa26b32de10.svg","isPro":false,"fullname":"Bohan Zeng","user":"zbhpku","type":"user"},{"_id":"661e62c6bac5d981f886f77b","avatarUrl":"/avatars/f1eb51ed4499ca434c8939573dfbd5e2.svg","isPro":false,"fullname":"Bozhou Li","user":"zooblastlbz","type":"user"},{"_id":"634fdf85f32062e9eb7a6c94","avatarUrl":"/avatars/24e932e1da9aa21af4c36bfddefca760.svg","isPro":false,"fullname":"Zhuoran Zhang","user":"SuperCarryZZR","type":"user"},{"_id":"6700b2b6bff0e8b51d07fa00","avatarUrl":"/avatars/6cd7e243b7bc37ae9d308c175cbe6f05.svg","isPro":false,"fullname":"asdasd","user":"asdjghh","type":"user"},{"_id":"69ce390201d713064aea5864","avatarUrl":"/avatars/af3977aeb5432599fb6b576c3f64a46b.svg","isPro":false,"fullname":"Bohan Zeng","user":"zbhpku1","type":"user"},{"_id":"660781a450d2b7a71091240d","avatarUrl":"/avatars/da9439b8920605d8427893d0ebc32dfa.svg","isPro":false,"fullname":"Bohan Zeng","user":"zbh0217","type":"user"},{"_id":"644d2532d185572dd1e48f90","avatarUrl":"/avatars/5831acebb02d8bc8f80f56b7b11c7c69.svg","isPro":false,"fullname":"Zhu","user":"zzzhu","type":"user"},{"_id":"674e77fa59a127e4eacf5dba","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/674e77fa59a127e4eacf5dba/W7qr94Buvvaio8zhKrEha.jpeg","isPro":false,"fullname":"Yifan Dai","user":"Moonwines","type":"user"},{"_id":"637f70d6fab5db9101c3dfc8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/637f70d6fab5db9101c3dfc8/NgkYNXWLDavLbrnCby2Fl.jpeg","isPro":false,"fullname":"Yujie Wei","user":"weilllllls","type":"user"},{"_id":"65e71ef39cf349af2940b317","avatarUrl":"/avatars/fc1cd8d3510946fc947d67b16b51834b.svg","isPro":false,"fullname":"Yuran Wang","user":"Ryann829","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"662c559b322afcbae51b3c8b","name":"KlingTeam","fullname":"Kling Team","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60e272ca6c78a8c122b12127/ZQV1aKLUDPf2rUcxxAqj6.jpeg"},"query":{}}">
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
Abstract
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.
Community
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.14202 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.14202 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.14202 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.