Hugging Face Daily Papers · · 5 min read

TempCloze: Can Video-LLMs Identify the Missing Middle?

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.</p>\n","updatedAt":"2026-09-11T06:45:31.711Z","author":{"_id":"65e7d63b14856e8859f1924c","avatarUrl":"/avatars/bc58ab252c7b4d95ff99e4506fb8d3e9.svg","fullname":"Pei Wenqi","name":"CedPei","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8728842735290527},"editors":["CedPei"],"editorAvatarUrls":["/avatars/bc58ab252c7b4d95ff99e4506fb8d3e9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.01515","authors":[{"_id":"6aa2f09f47a406da7901e5f4","user":{"_id":"65e7d63b14856e8859f1924c","avatarUrl":"/avatars/bc58ab252c7b4d95ff99e4506fb8d3e9.svg","isPro":false,"fullname":"Pei Wenqi","user":"CedPei","type":"user","name":"CedPei"},"name":"Wenqi Pei","status":"claimed_verified","statusLastChangedAt":"2026-09-11T08:45:04.689Z","hidden":false},{"_id":"6aa2f09f47a406da7901e5f5","name":"Henry Hengyuan Zhao","hidden":false},{"_id":"6aa2f09f47a406da7901e5f6","name":"Yilai Liu","hidden":false},{"_id":"6aa2f09f47a406da7901e5f7","user":{"_id":"65a28e129acab19980226731","avatarUrl":"/avatars/abc3828f807efc4e03837b0eae063f98.svg","isPro":false,"fullname":"Jiahao Meng","user":"marinero4972","type":"user","name":"marinero4972"},"name":"Jiahao Meng","status":"claimed_verified","statusLastChangedAt":"2026-09-11T09:53:35.230Z","hidden":false},{"_id":"6aa2f09f47a406da7901e5f8","user":{"_id":"6399c67bf78f75ae73146760","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6399c67bf78f75ae73146760/LAZxoSRD-hte-S9736iyg.jpeg","isPro":false,"fullname":"CHEN Han","user":"Concyclics","type":"user","name":"Concyclics"},"name":"Han Chen","status":"claimed_verified","statusLastChangedAt":"2026-09-11T09:53:37.834Z","hidden":false},{"_id":"6aa2f09f47a406da7901e5f9","name":"Ziyu Wang","hidden":false},{"_id":"6aa2f09f47a406da7901e5fa","name":"Hongyang Du","hidden":false}],"publishedAt":"2026-09-01T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"TempCloze: Can Video-LLMs Identify the Missing Middle?","submittedOnDailyBy":{"_id":"65e7d63b14856e8859f1924c","avatarUrl":"/avatars/bc58ab252c7b4d95ff99e4506fb8d3e9.svg","isPro":false,"fullname":"Pei Wenqi","user":"CedPei","type":"user","name":"CedPei"},"summary":"Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.","upvotes":6,"discussionId":"6aa2f09f47a406da7901e5fb","githubRepo":"https://github.com/CedricPei/Temporal-Cloze","githubRepoAddedBy":"user","ai_summary":"TempCloze evaluates visual temporal reasoning in Video-LLMs by requiring identification of missing video segments from distractors targeting semantics, alignment, and progression.","ai_keywords":["Video-LLMs","temporal reasoning","video cloze","same-source distractors","semantic alignment","progression","test-time scaling"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"67ea9ecfc234715db8dbf339","name":"hkuhk","fullname":"The University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ea9e8d2d95c10a0da11b0c/FNnR4M7YqKRuG43N5771B.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65e7d63b14856e8859f1924c","avatarUrl":"/avatars/bc58ab252c7b4d95ff99e4506fb8d3e9.svg","isPro":false,"fullname":"Pei Wenqi","user":"CedPei","type":"user"},{"_id":"663ccf6156e0690cb3acb1b5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/663ccf6156e0690cb3acb1b5/rZFLJtaIr9UP3v9m0N1u3.jpeg","isPro":false,"fullname":"Yilai Liu","user":"YilaiLiu-HKU","type":"user"},{"_id":"6399c67bf78f75ae73146760","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6399c67bf78f75ae73146760/LAZxoSRD-hte-S9736iyg.jpeg","isPro":false,"fullname":"CHEN Han","user":"Concyclics","type":"user"},{"_id":"656d767f02a56b531acf1293","avatarUrl":"/avatars/c428b20952aef82486174497f325bd95.svg","isPro":false,"fullname":"Fish","user":"TropicalFatFish","type":"user"},{"_id":"65a28e129acab19980226731","avatarUrl":"/avatars/abc3828f807efc4e03837b0eae063f98.svg","isPro":false,"fullname":"Jiahao Meng","user":"marinero4972","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67ea9ecfc234715db8dbf339","name":"hkuhk","fullname":"The University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ea9e8d2d95c10a0da11b0c/FNnR4M7YqKRuG43N5771B.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.01515.md","query":{}}">
Papers
arxiv:2609.01515

TempCloze: Can Video-LLMs Identify the Missing Middle?

Published on Sep 1
· Submitted by
Pei Wenqi
on Sep 11

Abstract

TempCloze evaluates visual temporal reasoning in Video-LLMs by requiring identification of missing video segments from distractors targeting semantics, alignment, and progression.

Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.

Community

Paper author Paper submitter about 7 hours ago

Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.01515
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.01515 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.01515 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers