Hugging Face Daily Papers · · 4 min read

Self Gradient Forcing: Native Long Video Extrapolation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Self Gradient Forcing: Native Long Video Extrapolation<br>Code: <a href=\"https://github.com/zhuang2002/Self_Gradient_Forcing\" rel=\"nofollow\">https://github.com/zhuang2002/Self_Gradient_Forcing</a><br>Page: <a href=\"https://zhuang2002.github.io/SelfGradientForcing/\" rel=\"nofollow\">https://zhuang2002.github.io/SelfGradientForcing/</a><br>HF paper: <a href=\"https://huggingface.co/papers/2607.20368\">https://huggingface.co/papers/2607.20368</a><br>Arxiv: <a href=\"https://arxiv.org/abs/2607.20368\" rel=\"nofollow\">https://arxiv.org/abs/2607.20368</a></p>\n","updatedAt":"2026-07-23T03:47:25.174Z","author":{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","fullname":"JunhaoZhuang","name":"JunhaoZhuang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":87,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3890930414199829},"editors":["JunhaoZhuang"],"editorAvatarUrls":["/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.20368","authors":[{"_id":"6a617e47c3792c34f5a040b0","name":"Junhao Zhuang","hidden":false},{"_id":"6a617e47c3792c34f5a040b1","name":"Shiyi Zhang","hidden":false},{"_id":"6a617e47c3792c34f5a040b2","name":"Yuxuan Bian","hidden":false},{"_id":"6a617e47c3792c34f5a040b3","name":"Yaowei Li","hidden":false},{"_id":"6a617e47c3792c34f5a040b4","name":"Yawen Luo","hidden":false},{"_id":"6a617e47c3792c34f5a040b5","name":"Yijun Liu","hidden":false},{"_id":"6a617e47c3792c34f5a040b6","name":"Weiyang Jin","hidden":false},{"_id":"6a617e47c3792c34f5a040b7","name":"Songchun Zhang","hidden":false},{"_id":"6a617e47c3792c34f5a040b8","name":"Xianglong He","hidden":false},{"_id":"6a617e47c3792c34f5a040b9","name":"Xuying Zhang","hidden":false},{"_id":"6a617e47c3792c34f5a040ba","name":"Haoran Li","hidden":false},{"_id":"6a617e47c3792c34f5a040bb","name":"Haoyang Huang","hidden":false},{"_id":"6a617e47c3792c34f5a040bc","name":"Zeyue Xue","hidden":false},{"_id":"6a617e47c3792c34f5a040bd","name":"Nan Duan","hidden":false}],"publishedAt":"2026-07-22T00:00:00.000Z","submittedOnDailyAt":"2026-07-23T00:00:00.000Z","title":"Self Gradient Forcing: Native Long Video Extrapolation","submittedOnDailyBy":{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","isPro":false,"fullname":"JunhaoZhuang","user":"JunhaoZhuang","type":"user","name":"JunhaoZhuang"},"summary":"Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.","upvotes":19,"discussionId":"6a617e47c3792c34f5a040be","projectPage":"https://zhuang2002.github.io/SelfGradientForcing/","githubRepo":"https://github.com/zhuang2002/Self_Gradient_Forcing","githubRepoAddedBy":"user","githubStars":17},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64970d3d9c3b29dca8633f87","avatarUrl":"/avatars/11e3c9c66d28490d6d09925f9aa47cd1.svg","isPro":false,"fullname":"JunhaoZhuang","user":"JunhaoZhuang","type":"user"},{"_id":"65ec2463df813b9c1571da7d","avatarUrl":"/avatars/731c90738d01f4e46dd4503d1a5e5137.svg","isPro":false,"fullname":"yaoruchang","user":"ruchangyao","type":"user"},{"_id":"683971013e5dd928f059bade","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/683971013e5dd928f059bade/0LyqSxoz-cm-JQuIs9ZAx.jpeg","isPro":false,"fullname":"Ruikang","user":"Lyricccco","type":"user"},{"_id":"6362801380c1a705a6ea54ac","avatarUrl":"/avatars/041ad5abf9be42e336938f51ebb8746c.svg","isPro":false,"fullname":"Yaowei Li","user":"Yw22","type":"user"},{"_id":"65e92ea71a58734e13ca709e","avatarUrl":"/avatars/e9b6e0e38814ebdbe9a1c60d73214749.svg","isPro":false,"fullname":"Linhua Huang","user":"Linhua-Huang","type":"user"},{"_id":"65781534e390cfd40998d7af","avatarUrl":"/avatars/85d3be7f74b9d959ba9b3ccc04398536.svg","isPro":false,"fullname":"Hongyang Wei","user":"nonwhy","type":"user"},{"_id":"66743477ab975c859114d410","avatarUrl":"/avatars/ac692cc336e383fb2cb53db6d1e3fe8c.svg","isPro":false,"fullname":"yawenluo","user":"yawenluo","type":"user"},{"_id":"66608add236f958513d21d2e","avatarUrl":"/avatars/53eca0891c98cbb93be899885160a983.svg","isPro":false,"fullname":"Weiyang Jin","user":"Wayne-King","type":"user"},{"_id":"683d173214785f6d5902f9c0","avatarUrl":"/avatars/2357d97bbcecd3a51279442bd99a0c50.svg","isPro":false,"fullname":"Haoran Li","user":"jahnsonblack","type":"user"},{"_id":"67344a21db744d70cb9be933","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Z9Eh-asE3ZISNGXOTzTFQ.png","isPro":false,"fullname":"Haoyu Wang","user":"why986","type":"user"},{"_id":"646eac510867c99c2d3fde08","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/646eac510867c99c2d3fde08/fvIiW7zj4aTNbp16kTNBA.jpeg","isPro":false,"fullname":"Yaofeng Su","user":"Exploration","type":"user"},{"_id":"63721f5ada3183d9d53cfe1f","avatarUrl":"/avatars/593c14c907848da7dbc9e5418751bd94.svg","isPro":false,"fullname":"Xue Zeyue","user":"xzyhku","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.20368.md","query":{}}">
Papers
arxiv:2607.20368

Self Gradient Forcing: Native Long Video Extrapolation

Published on Jul 22
· Submitted by
JunhaoZhuang
on Jul 23
#2 Paper of the day
Authors:
,

Abstract

Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the historical context-gradient gap. We propose Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing supervision signal without backpropagating through the full serial rollout. Pass 1 performs a no-gradient autoregressive rollout matching inference and, at a sampled denoising exit step, records both the self-generated context and the noisy latents fed to the model. Pass 2 performs parallel context-gradient reconstruction for the recorded exit step. The generated context is used as stop-gradient clean-latent input, while the model recomputes the context KV representations and future-to-context causal attention. Thus, SGF provides the missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory. Across extensive long-horizon frame-wise and chunk-wise experiments under different initializations, SGF achieves stronger native long-video extrapolation than Self Forcing, especially in subject identity, background/layout consistency, and temporal stability. Remarkably, using only a 5-second training window, SGF can extrapolate to videos lasting several minutes. Code and models will be released to advance research on autoregressive video generation.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.20368
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.20368 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.20368 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers