Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.</p>\n","updatedAt":"2026-08-07T01:59:02.621Z","author":{"_id":"65e387095132c2edd193ae49","avatarUrl":"/avatars/39278e5b026bcbdde88c560fc54018c5.svg","fullname":"Yifan Shen","name":"SivanSX","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8634002804756165},"editors":["SivanSX"],"editorAvatarUrls":["/avatars/39278e5b026bcbdde88c560fc54018c5.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05631","authors":[{"_id":"6a753b9ee1228e04b3238150","user":{"_id":"65e387095132c2edd193ae49","avatarUrl":"/avatars/39278e5b026bcbdde88c560fc54018c5.svg","isPro":false,"fullname":"Yifan Shen","user":"SivanSX","type":"user","name":"SivanSX"},"name":"Yifan Shen","status":"claimed_verified","statusLastChangedAt":"2026-08-07T08:45:04.487Z","hidden":false},{"_id":"6a753b9ee1228e04b3238151","name":"Jian Xu","hidden":false},{"_id":"6a753b9ee1228e04b3238152","user":{"_id":"67eb59446b0871640462bfb0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/fRRrOjm10id376KdVOsAO.png","isPro":false,"fullname":"Boyi Li","user":"Resurgammm","type":"user","name":"Resurgammm"},"name":"Boyi Li","status":"claimed_verified","statusLastChangedAt":"2026-08-07T08:45:04.492Z","hidden":false},{"_id":"6a753b9ee1228e04b3238153","name":"Yuner Zhang","hidden":false},{"_id":"6a753b9ee1228e04b3238154","name":"Tianjiao Yu","hidden":false},{"_id":"6a753b9ee1228e04b3238155","name":"Bingxuan Li","hidden":false},{"_id":"6a753b9ee1228e04b3238156","name":"Houze Yang","hidden":false},{"_id":"6a753b9ee1228e04b3238157","name":"Rushi Wang","hidden":false},{"_id":"6a753b9ee1228e04b3238158","name":"Xu Cao","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"ChronoVision: Temporal Reasoning via Latent State Reconstruction","submittedOnDailyBy":{"_id":"65e387095132c2edd193ae49","avatarUrl":"/avatars/39278e5b026bcbdde88c560fc54018c5.svg","isPro":false,"fullname":"Yifan Shen","user":"SivanSX","type":"user","name":"SivanSX"},"summary":"Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.","upvotes":24,"discussionId":"6a753b9fe1228e04b3238159","organization":{"_id":"6843ed589d4c30b7ca665513","name":"PediaMedAI","fullname":"PediaMed AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5fc9f05d52770aca770bd3d9/Iw8E1A9k4Id51m8x90zOR.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67eb59446b0871640462bfb0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/fRRrOjm10id376KdVOsAO.png","isPro":false,"fullname":"Boyi Li","user":"Resurgammm","type":"user"},{"_id":"6a129c23980b93ff94814d70","avatarUrl":"/avatars/e81565743d1214ab314bc139188d770d.svg","isPro":false,"fullname":"yuner zhang","user":"yuner123","type":"user"},{"_id":"63f2ddc17ddf724fbcc6c3c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f2ddc17ddf724fbcc6c3c1/LkNUsHzAcSHA-o_gzFVPR.jpeg","isPro":true,"fullname":"Tianjiao Yu","user":"Tianjiao-Yu","type":"user"},{"_id":"64b9e7d144ade326861379a3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cBDthxj4dqUU6q9bgjP6_.png","isPro":false,"fullname":"Jiateng Liu","user":"Lumos-Jiateng","type":"user"},{"_id":"65e387095132c2edd193ae49","avatarUrl":"/avatars/39278e5b026bcbdde88c560fc54018c5.svg","isPro":false,"fullname":"Yifan Shen","user":"SivanSX","type":"user"},{"_id":"6840105d01d706fdb861a440","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/G8VGI8PEm-y7d8Hx9VB0Z.jpeg","isPro":false,"fullname":"Jian Xu","user":"Grapesoda08","type":"user"},{"_id":"6a6a831d726441725a2ed822","avatarUrl":"/avatars/45ba67b0e43f5f2e65aff93418dcdba2.svg","isPro":false,"fullname":"Charles Thompson","user":"charlesthompson","type":"user"},{"_id":"6a6a9449144c90a0d41692f0","avatarUrl":"/avatars/6af977a850fe2dfd55c11f7100d92251.svg","isPro":false,"fullname":"David Jackson","user":"lunarbeacon","type":"user"},{"_id":"6a6a95368acf46140bae93a3","avatarUrl":"/avatars/c64673a23a306cc73ed00382ec370e10.svg","isPro":false,"fullname":"Kevin Thompson","user":"rapidFox","type":"user"},{"_id":"6a6aa6de726441725a30a027","avatarUrl":"/avatars/2e777075d9d0d6dcdaaf0e97860f551e.svg","isPro":false,"fullname":"Sarah Clark","user":"granitecore","type":"user"},{"_id":"6a6c7f02101ebc51fc3dba6a","avatarUrl":"/avatars/a0fbbf918e016065a0871bf717a7afbe.svg","isPro":false,"fullname":"Michael Sanchez","user":"Indigo-Michael","type":"user"},{"_id":"6a6c84510e22708c334af29d","avatarUrl":"/avatars/59a86460c683848fd2575630eda71690.svg","isPro":false,"fullname":"Elizabeth Gonzalez","user":"IndigoElizabeth","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6843ed589d4c30b7ca665513","name":"PediaMedAI","fullname":"PediaMed AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/5fc9f05d52770aca770bd3d9/Iw8E1A9k4Id51m8x90zOR.jpeg"},"query":{}}">
ChronoVision: Temporal Reasoning via Latent State Reconstruction
Abstract
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.
Community
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.05631 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.05631 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.05631 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.