Hugging Face Daily Papers · · 4 min read

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🚀 Sol-Attn is a training-free sparse attention method that accelerates video generation while better preserving quality.</p>\n<p>Sol-Attn unifies dynamic routing, sparse computation, and approximate correction in a single online-softmax pass:<br>• On-the-fly block thresholding for dynamic yet controllable budgets<br>• Proxy-score reuse to approximate unselected blocks</p>\n<p>Results (vs dense FlashAttention-3):<br>• Wan 2.1-14B: 2.02× end-to-end<br>• HunyuanVideo-13B: 2.12× end-to-end<br>• LTX 2.3: up to 2.4× end-to-end</p>\n<p>When integrated into Sol-Engine (with kernel fusion + caching):<br>• Wan 2.1-14B: 3.48× end-to-end<br>• HunyuanVideo-13B: 5.08× end-to-end<br>Already available in Sol-Engine.</p>\n<p>The B200 kernel is still under further optimization.</p>\n<p>🎬 Project: <a href=\"http://nvlabs.github.io/Sana/Sol-Attn/\" rel=\"nofollow\">http://nvlabs.github.io/Sana/Sol-Attn/</a><br>📄 Paper: <a href=\"https://arxiv.org/abs/2607.24027\" rel=\"nofollow\">https://arxiv.org/abs/2607.24027</a><br>🔗 Code: <a href=\"https://github.com/NVlabs/Sana/tree/sol-engine\" rel=\"nofollow\">https://github.com/NVlabs/Sana/tree/sol-engine</a></p>\n","updatedAt":"2026-07-28T06:57:09.715Z","author":{"_id":"66015e8aa4d296af07de538e","avatarUrl":"/avatars/a1295c631cc2646282c545859975ce4c.svg","fullname":"Owen","name":"Owen777","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":91,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7568889260292053},"editors":["Owen777"],"editorAvatarUrls":["/avatars/a1295c631cc2646282c545859975ce4c.svg"],"reactions":[],"isReport":false}},{"id":"6a6856d6f6537b28f9ae9ab7","author":{"_id":"64e86fbd0c2413c3571ef7a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e86fbd0c2413c3571ef7a6/KDpJ0UpjICfQKUr14ekcR.png","fullname":"Haopeng Li","name":"hp-l33","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false},"createdAt":"2026-07-28T07:14:30.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\n\n\nhttps://cdn-uploads.huggingface.co/production/uploads/64e86fbd0c2413c3571ef7a6/KDjVcDOg2MCFMgiSkvEJg.mp4\n","html":"<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/64e86fbd0c2413c3571ef7a6/KDjVcDOg2MCFMgiSkvEJg.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-07-28T07:14:30.072Z","author":{"_id":"64e86fbd0c2413c3571ef7a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e86fbd0c2413c3571ef7a6/KDpJ0UpjICfQKUr14ekcR.png","fullname":"Haopeng Li","name":"hp-l33","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5270630121231079},"editors":["hp-l33"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64e86fbd0c2413c3571ef7a6/KDpJ0UpjICfQKUr14ekcR.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.24027","authors":[{"_id":"6a684a8973f69d5af2bec959","user":{"_id":"64e86fbd0c2413c3571ef7a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e86fbd0c2413c3571ef7a6/KDpJ0UpjICfQKUr14ekcR.png","isPro":false,"fullname":"Haopeng Li","user":"hp-l33","type":"user","name":"hp-l33"},"name":"Haopeng Li","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:45:04.657Z","hidden":false},{"_id":"6a684a8973f69d5af2bec95a","name":"Yitong Li","hidden":false},{"_id":"6a684a8973f69d5af2bec95b","user":{"_id":"645b5b09bc7518912e1f9733","avatarUrl":"/avatars/4d35f728b41f93881a9b67c337f4d1df.svg","isPro":false,"fullname":"Chen","user":"Lawrence-cj","type":"user","name":"Lawrence-cj"},"name":"Junsong Chen","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:58:18.202Z","hidden":false},{"_id":"6a684a8973f69d5af2bec95c","user":{"_id":"66015e8aa4d296af07de538e","avatarUrl":"/avatars/a1295c631cc2646282c545859975ce4c.svg","isPro":false,"fullname":"Owen","user":"Owen777","type":"user","name":"Owen777"},"name":"Tian Ye","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:58:19.938Z","hidden":false},{"_id":"6a684a8973f69d5af2bec95d","name":"Haozhe Liu","hidden":false},{"_id":"6a684a8973f69d5af2bec95e","name":"Jincheng Yu","hidden":false},{"_id":"6a684a8973f69d5af2bec95f","name":"Duomin Wang","hidden":false},{"_id":"6a684a8973f69d5af2bec960","name":"Ruihua Zhang","hidden":false},{"_id":"6a684a8973f69d5af2bec961","name":"Zeke Xie","hidden":false},{"_id":"6a684a8973f69d5af2bec962","name":"Enze Xie","hidden":false},{"_id":"6a684a8973f69d5af2bec963","user":{"_id":"63797f727df2fefdcaf3ff7e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1668906853549-noauth.jpeg","isPro":false,"fullname":"Song","user":"songhan","type":"user","name":"songhan"},"name":"Song Han","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:58:16.424Z","hidden":false}],"publishedAt":"2026-07-27T00:00:00.000Z","submittedOnDailyAt":"2026-07-28T00:00:00.000Z","title":"Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification","submittedOnDailyBy":{"_id":"66015e8aa4d296af07de538e","avatarUrl":"/avatars/a1295c631cc2646282c545859975ce4c.svg","isPro":false,"fullname":"Owen","user":"Owen777","type":"user","name":"Owen777"},"summary":"Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify attention both efficiently and accurately for two reasons: (1) Rigid, unpredictable, and costly routing: selecting a fixed fraction of top-ranked blocks by proxy score imposes fixed budgets, whereas retaining blocks to reach a target cumulative proxy probability mass yields dynamic but potentially imbalanced budgets; both incur non-negligible overhead from computing and materializing proxy scores. (2) Lossy keep-or-drop sparsification: unselected blocks are discarded entirely, degrading accuracy under aggressive sparsity. These limitations motivate cheaper dynamic-budget routing while limiting accuracy degradation. In this paper, we introduce training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention. The core of Sol-Attn is on-the-fly block thresholding with proxy-score reuse, which selects critical blocks by comparing block proxy scores against a threshold during online softmax. This design enables dynamic yet controllable block budgets without materializing the proxy map, while directly reusing the proxy scores of unselected blocks to approximate their contribution. Experiments across image and video generation tasks show that Sol-Attn advances the quality-efficiency frontier of training-free sparse attention, delivering 2.1 times and 2.3 times end-to-end speedups for video generation and editing, respectively, while preserving visual quality.","upvotes":24,"discussionId":"6a684a8973f69d5af2bec964","projectPage":"http://nvlabs.github.io/Sana/Sol-Attn/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"603bdba23249b99991dbcbc4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/603bdba23249b99991dbcbc4/cxCnN1H-RXOhojHY3Wcxo.jpeg","isPro":false,"fullname":"Tolga Cangöz","user":"tolgacangoz","type":"user"},{"_id":"66015e8aa4d296af07de538e","avatarUrl":"/avatars/a1295c631cc2646282c545859975ce4c.svg","isPro":false,"fullname":"Owen","user":"Owen777","type":"user"},{"_id":"64e86fbd0c2413c3571ef7a6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e86fbd0c2413c3571ef7a6/KDpJ0UpjICfQKUr14ekcR.png","isPro":false,"fullname":"Haopeng Li","user":"hp-l33","type":"user"},{"_id":"67136093d2e50f1e8c9fad52","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/0q49MyGuav8lJ9CIeyLhu.png","isPro":false,"fullname":"Donghao Zhou","user":"donghao-zhou","type":"user"},{"_id":"64ae9b88a22a179fc4d07992","avatarUrl":"/avatars/c9065f04a1188ea3129e56a90328ffd3.svg","isPro":false,"fullname":"wang","user":"dorni","type":"user"},{"_id":"644e24ced6001776ed74b0fb","avatarUrl":"/avatars/1833d681e1c63266f85e4c3697221b70.svg","isPro":false,"fullname":"L","user":"Hao-Zhe","type":"user"},{"_id":"645b5b09bc7518912e1f9733","avatarUrl":"/avatars/4d35f728b41f93881a9b67c337f4d1df.svg","isPro":false,"fullname":"Chen","user":"Lawrence-cj","type":"user"},{"_id":"61af81009f77f7b669578f95","avatarUrl":"/avatars/fb50773ac49948940eb231834ee6f2fd.svg","isPro":false,"fullname":"rotem israeli","user":"irotem98","type":"user"},{"_id":"6479925ab77e18dbf640bd67","avatarUrl":"/avatars/bb52ecd22ca4b49157f8668be35409e7.svg","isPro":false,"fullname":"Zhiheng Liu","user":"Johanan0528","type":"user"},{"_id":"6696755fd26a65bd255184d3","avatarUrl":"/avatars/8d46c21a7b23f0100a7e3385fea61edf.svg","isPro":false,"fullname":"Kele Shao (SII)","user":"keleshao","type":"user"},{"_id":"65e7be93a5deaa480d51a88c","avatarUrl":"/avatars/44bf42dbde5ccc64114a36f1cfdad635.svg","isPro":false,"fullname":"Qihang Fan","user":"aldjalkdf","type":"user"},{"_id":"67443f7c2f6a94e877c95b32","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/RbMBDBKyHHRjjqn7_t22G.png","isPro":false,"fullname":"Zixuan Chen","user":"zxchen00","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Papers
arxiv:2607.24027

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Published on Jul 27
· Submitted by
Owen
on Jul 28

Abstract

Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify attention both efficiently and accurately for two reasons: (1) Rigid, unpredictable, and costly routing: selecting a fixed fraction of top-ranked blocks by proxy score imposes fixed budgets, whereas retaining blocks to reach a target cumulative proxy probability mass yields dynamic but potentially imbalanced budgets; both incur non-negligible overhead from computing and materializing proxy scores. (2) Lossy keep-or-drop sparsification: unselected blocks are discarded entirely, degrading accuracy under aggressive sparsity. These limitations motivate cheaper dynamic-budget routing while limiting accuracy degradation. In this paper, we introduce training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention. The core of Sol-Attn is on-the-fly block thresholding with proxy-score reuse, which selects critical blocks by comparing block proxy scores against a threshold during online softmax. This design enables dynamic yet controllable block budgets without materializing the proxy map, while directly reusing the proxy scores of unselected blocks to approximate their contribution. Experiments across image and video generation tasks show that Sol-Attn advances the quality-efficiency frontier of training-free sparse attention, delivering 2.1 times and 2.3 times end-to-end speedups for video generation and editing, respectively, while preserving visual quality.

Community

Paper author Paper submitter about 14 hours ago

🚀 Sol-Attn is a training-free sparse attention method that accelerates video generation while better preserving quality.

Sol-Attn unifies dynamic routing, sparse computation, and approximate correction in a single online-softmax pass:
• On-the-fly block thresholding for dynamic yet controllable budgets
• Proxy-score reuse to approximate unselected blocks

Results (vs dense FlashAttention-3):
• Wan 2.1-14B: 2.02× end-to-end
• HunyuanVideo-13B: 2.12× end-to-end
• LTX 2.3: up to 2.4× end-to-end

When integrated into Sol-Engine (with kernel fusion + caching):
• Wan 2.1-14B: 3.48× end-to-end
• HunyuanVideo-13B: 5.08× end-to-end
Already available in Sol-Engine.

The B200 kernel is still under further optimization.

🎬 Project: http://nvlabs.github.io/Sana/Sol-Attn/
📄 Paper: https://arxiv.org/abs/2607.24027
🔗 Code: https://github.com/NVlabs/Sana/tree/sol-engine

Paper author about 14 hours ago

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.24027 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.24027 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.24027 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers