Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: <a href=\"https://cdfan0627.github.io/LongE2V-page/\" rel=\"nofollow\">https://cdfan0627.github.io/LongE2V-page/</a></p>\n","updatedAt":"2026-07-10T06:52:12.679Z","author":{"_id":"6459d5da3b6fafd9664807ab","avatarUrl":"/avatars/57430d1bbde3a2fe5586e5fbcafb0e74.svg","fullname":"Yu-Lun Liu","name":"yulunliu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8744712471961975},"editors":["yulunliu"],"editorAvatarUrls":["/avatars/57430d1bbde3a2fe5586e5fbcafb0e74.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.08770","authors":[{"_id":"6a50965675fd3d966bd45e95","user":{"_id":"64ea1e12925565abda02b17b","avatarUrl":"/avatars/b2bc33d95a147c6c8cf6b54672eb5a97.svg","isPro":false,"fullname":"Cheng-De Fan","user":"fansam39","type":"user","name":"fansam39"},"name":"Cheng-De Fan","status":"claimed_verified","statusLastChangedAt":"2026-07-10T07:44:32.961Z","hidden":false},{"_id":"6a50965675fd3d966bd45e96","name":"Chun-Wei Tuan Mu","hidden":false},{"_id":"6a50965675fd3d966bd45e97","name":"Chen-Wei Chang","hidden":false},{"_id":"6a50965675fd3d966bd45e98","name":"Chin-Yang Lin","hidden":false},{"_id":"6a50965675fd3d966bd45e99","name":"Kun-Ru Wu","hidden":false},{"_id":"6a50965675fd3d966bd45e9a","name":"Yu-Chee Tseng","hidden":false},{"_id":"6a50965675fd3d966bd45e9b","name":"Yu-Lun Liu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6459d5da3b6fafd9664807ab/q7t0w7KPRxkUz_5wjypED.mp4"],"publishedAt":"2026-07-09T00:00:00.000Z","submittedOnDailyAt":"2026-07-10T00:00:00.000Z","title":"LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models","submittedOnDailyBy":{"_id":"6459d5da3b6fafd9664807ab","avatarUrl":"/avatars/57430d1bbde3a2fe5586e5fbcafb0e74.svg","isPro":false,"fullname":"Yu-Lun Liu","user":"yulunliu","type":"user","name":"yulunliu"},"summary":"Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/","upvotes":19,"discussionId":"6a50965775fd3d966bd45e9c","projectPage":"https://cdfan0627.github.io/LongE2V-page/","githubRepo":"https://github.com/cdfan0627/LongE2V","githubRepoAddedBy":"user","ai_summary":"LongE2V enables high-quality video recovery from sparse event streams by leveraging pre-trained video diffusion priors and addressing temporal stability and frame interpolation challenges.","ai_keywords":["video diffusion priors","event-based video reconstruction","frame interpolation","temporal drift","autoregressive unrolling","adaptive context switching","reencoding alignment","cross residual correction","event voxel density augmentation","zero-shot generalization"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":20,"organization":{"_id":"63e39e6499a032b1c950403d","name":"NYCU","fullname":"National Yang Ming Chiao Tung University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e39df6c65f975b436bb6b8/WLWf1bSpvrXBYYKEdXbgU.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6459d5da3b6fafd9664807ab","avatarUrl":"/avatars/57430d1bbde3a2fe5586e5fbcafb0e74.svg","isPro":false,"fullname":"Yu-Lun Liu","user":"yulunliu","type":"user"},{"_id":"6672fe26c33b5004b69a1d6a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Ff8cOS6Y0TPUSihx_hOMe.png","isPro":false,"fullname":"YouZhe","user":"YouZhe","type":"user"},{"_id":"64cdecee2f1f9578a0e701c8","avatarUrl":"/avatars/95a51dd4e1b7b9366ebcbd6028ad148b.svg","isPro":false,"fullname":"Yi-Ruei Liu","user":"Shigon","type":"user"},{"_id":"68a41489d9b513a884bca475","avatarUrl":"/avatars/e4caeb16f3c4e7c36835cf26c8cb0d2c.svg","isPro":false,"fullname":"You-Zhe Xie","user":"YouZhe0305","type":"user"},{"_id":"676ce504027822ead2b5f193","avatarUrl":"/avatars/91de797fc4a971a481b2dce82b579f66.svg","isPro":false,"fullname":"YuanKangNeilLee","user":"NeilLeeNTU","type":"user"},{"_id":"6672ebc506b6d49dda7598c5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6672ebc506b6d49dda7598c5/9yUeKzZZVtBoy2L-dNPMf.png","isPro":false,"fullname":"Sytwu","user":"Sytwu","type":"user"},{"_id":"64ea1e12925565abda02b17b","avatarUrl":"/avatars/b2bc33d95a147c6c8cf6b54672eb5a97.svg","isPro":false,"fullname":"Cheng-De Fan","user":"fansam39","type":"user"},{"_id":"696519a7c4eb6cb05630a14b","avatarUrl":"/avatars/3c9848c7e634a61785524439a44b79a0.svg","isPro":false,"fullname":"yeh chi-yang","user":"chi-yang","type":"user"},{"_id":"68e0130b12b8c1ecde3383ae","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/gptMaTL5NfK5gLqVH_1zm.png","isPro":false,"fullname":"You-Zhe Xie","user":"Beck0305","type":"user"},{"_id":"684afb68f144221f28256461","avatarUrl":"/avatars/48c3d76057f78e1ca4abb2b121a2d089.svg","isPro":false,"fullname":"Zhenjun Zhao","user":"rickyeric","type":"user"},{"_id":"666afb91e936f6cbcfc8b50c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/666afb91e936f6cbcfc8b50c/_lcbPagwDTn02TaOSDxUq.jpeg","isPro":false,"fullname":"Chin-Yang Lin","user":"linjohnss","type":"user"},{"_id":"662798ac599e147a00b1ba67","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/hWKN5EZXNjdLH4Z3mAEaG.png","isPro":false,"fullname":"Samynhn","user":"Samynhn","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63e39e6499a032b1c950403d","name":"NYCU","fullname":"National Yang Ming Chiao Tung University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e39df6c65f975b436bb6b8/WLWf1bSpvrXBYYKEdXbgU.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.08770.md","query":{}}">
LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
Abstract
LongE2V enables high-quality video recovery from sparse event streams by leveraging pre-trained video diffusion priors and addressing temporal stability and frame interpolation challenges.
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/
Community
Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation. By fine-tuning a foundational video model, our approach achieves high data efficiency and superior perceptual quality. We introduce Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences. We also propose Reencoding Alignment with Cross Residual Correction to ensure precise bidirectional consistency during frame interpolation. Furthermore, Event Voxel Density Augmentation ensures robustness across varying sensor resolutions. Extensive experiments on real-world benchmarks demonstrate that LongE2V outperforms state-of-the-art methods across all three tasks, exhibiting exceptional temporal coherence and zero-shot generalization. Project page: https://cdfan0627.github.io/LongE2V-page/
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.08770 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.08770 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.08770 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.