<a href=\"https://cdn-uploads.huggingface.co/production/uploads/67f87529318a17cc80365190/mOm1eNwWKUtkCNixbsukM.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/67f87529318a17cc80365190/mOm1eNwWKUtkCNixbsukM.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-07-17T16:45:42.583Z","author":{"_id":"67f87529318a17cc80365190","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f87529318a17cc80365190/kv4cAvD5BrWQFRXKG4FXg.jpeg","fullname":"Maijunxian Wang","name":"Mark7121983123","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.39387303590774536},"editors":["Mark7121983123"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/67f87529318a17cc80365190/kv4cAvD5BrWQFRXKG4FXg.jpeg"],"reactions":[],"isReport":false}},{"id":"6a5afa00a1a4165c04edb25a","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":376,"isUserFollowing":false},"createdAt":"2026-07-18T03:58:56.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [Next Forcing: Causal World Modeling with Multi-Chunk Prediction](https://huggingface.co/papers/2606.11187) (2026)\n* [Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model](https://huggingface.co/papers/2607.03509) (2026)\n* [BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression](https://huggingface.co/papers/2606.10135) (2026)\n* [Latent Visual Cache for Video Reasoning](https://huggingface.co/papers/2607.02607) (2026)\n* [UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation](https://huggingface.co/papers/2607.05133) (2026)\n* [MoWorld: A Flash World Model](https://huggingface.co/papers/2607.06216) (2026)\n* [Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning](https://huggingface.co/papers/2606.02842) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.11187\">Next Forcing: Causal World Modeling with Multi-Chunk Prediction</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.03509\">Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.10135\">BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.02607\">Latent Visual Cache for Video Reasoning</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.05133\">UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.06216\">MoWorld: A Flash World Model</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.02842\">Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-07-18T03:58:56.519Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":376,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7143896818161011},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.15278","authors":[{"_id":"6a5a5baad6689d7651ea876d","name":"Zezhong Qian","hidden":false},{"_id":"6a5a5baad6689d7651ea876e","name":"Xiaowei Chi","hidden":false},{"_id":"6a5a5baad6689d7651ea876f","name":"Chak-Wing Mak","hidden":false},{"_id":"6a5a5baad6689d7651ea8770","name":"Tianze Zhou","hidden":false},{"_id":"6a5a5baad6689d7651ea8771","name":"Ruibin Yuan","hidden":false},{"_id":"6a5a5baad6689d7651ea8772","name":"Yuhan Rui","hidden":false},{"_id":"6a5a5baad6689d7651ea8773","name":"Hengzhe Sun","hidden":false},{"_id":"6a5a5baad6689d7651ea8774","name":"Zhuoqun Wu","hidden":false},{"_id":"6a5a5baad6689d7651ea8775","name":"Yuming Li","hidden":false},{"_id":"6a5a5baad6689d7651ea8776","name":"Siyuan Qian","hidden":false},{"_id":"6a5a5baad6689d7651ea8777","name":"Sirui Han","hidden":false},{"_id":"6a5a5baad6689d7651ea8778","name":"Shanghang Zhang","hidden":false}],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-17T00:00:00.000Z","title":"Hierarchical Denoising For Multi-Step Visual Reasoning","submittedOnDailyBy":{"_id":"67f87529318a17cc80365190","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f87529318a17cc80365190/kv4cAvD5BrWQFRXKG4FXg.jpeg","isPro":false,"fullname":"Maijunxian Wang","user":"Mark7121983123","type":"user","name":"Mark7121983123"},"summary":"Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to achieve logical consistency and low-latency streaming for complex reasoning tasks. We propose HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning. HDR organizes video latents into a tree-structured hierarchy, enabling coarse-to-fine reasoning before streaming output. Coarse denoising layers preserve uncertain hypotheses for global planning, while finer layers progressively refine them into concrete visual states. A sparse hierarchical attention pattern (SHAP) further reduces temporal attention costs. We introduce a level-stratified multi-step video reasoning benchmark with out-of-distribution cases, covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring. Compared with streaming autoregressive diffusion baselines, HDR improves success from 34.22 to 60.29 (76.2% relative gain) and increases average progress from 76.00 to 89.56, demonstrating more consistent reasoning trajectories. HDR maintains low-latency streaming at 0.70 seconds per latent, achieving 54.2 times faster inference than bidirectional diffusion. It also retains 82.9% of full-data performance with only 2% training data, compared with 52.0% for bidirectional diffusion. Real-world robot experiments further demonstrate HDR's potential for physical interaction and world modeling. Project demo: https://hierarchical-diffusion-reasoning.github.io/.","upvotes":3,"discussionId":"6a5a5baad6689d7651ea8779","projectPage":"https://hierarchical-diffusion-reasoning.github.io/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"66bde8b33f0c44697d666f8f","avatarUrl":"/avatars/063f9727d80bf50a38d030866d99eb2d.svg","isPro":false,"fullname":"marco","user":"marcozzxx810","type":"user"},{"_id":"69ccaf5baa7e18f7efae4997","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Q9fpe3ZxK0-yt5QgZOd3U.jpeg","isPro":false,"fullname":"Мария Яковлев","user":"liamclar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.15278.md","query":{}}">
Hierarchical Denoising For Multi-Step Visual Reasoning
Abstract
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to achieve logical consistency and low-latency streaming for complex reasoning tasks. We propose HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning. HDR organizes video latents into a tree-structured hierarchy, enabling coarse-to-fine reasoning before streaming output. Coarse denoising layers preserve uncertain hypotheses for global planning, while finer layers progressively refine them into concrete visual states. A sparse hierarchical attention pattern (SHAP) further reduces temporal attention costs. We introduce a level-stratified multi-step video reasoning benchmark with out-of-distribution cases, covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring. Compared with streaming autoregressive diffusion baselines, HDR improves success from 34.22 to 60.29 (76.2% relative gain) and increases average progress from 76.00 to 89.56, demonstrating more consistent reasoning trajectories. HDR maintains low-latency streaming at 0.70 seconds per latent, achieving 54.2 times faster inference than bidirectional diffusion. It also retains 82.9% of full-data performance with only 2% training data, compared with 52.0% for bidirectional diffusion. Real-world robot experiments further demonstrate HDR's potential for physical interaction and world modeling. Project demo: https://hierarchical-diffusion-reasoning.github.io/.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.15278 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.15278 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.15278 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.