TL;DR: Training-free appearance-indexed memory for long autoregressive video generation — retain content by <em>what appears</em>, not <em>when</em> it appeared.</p>\n<p>Sliding-window memory forgets early content (subjects drift), while pinning reference frames freezes motion. RECAP-Forcing resolves this with two plug-in mechanisms on a frozen backbone:</p>\n<ul>\n<li><strong>Reinforced attention sink</strong> — a single pre-softmax bias turns the first frames into memory the model actually retrieves;</li>\n<li><strong>Optical-flow novelty bank</strong> — new appearances (entrances, disocclusions) are detected via flow and their original KV states stored verbatim in a fixed-size bank, retained by novelty rather than recency.</li>\n</ul>\n<p>Zero learnable parameters. On Self-Forcing (Wan2.1-1.3B), Dynamic Degree jumps 27.5 → 58.1 while VBench Total improves 75.9 → 79.5; ~79% human preference over 216 blind pairwise judgments; identity stays stable in 5-minute rollouts. Generalizes across 4 causal video backbones with constant ~1.6× overhead.</p>\n","updatedAt":"2026-09-01T20:36:14.395Z","author":{"_id":"632a994b4ee2960462ecbf7e","avatarUrl":"/avatars/994e812be5a30d5908141215fdf18625.svg","fullname":"Haiyang Xu","name":"xuhaiyang3110","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8545826077461243},"editors":["xuhaiyang3110"],"editorAvatarUrls":["/avatars/994e812be5a30d5908141215fdf18625.svg"],"reactions":[],"isReport":false}},{"id":"6a977b0cb9c98429e282be5a","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false},"createdAt":"2026-09-02T01:25:32.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [Surprise Forcing: What to Remember, When to Skip in Long Video Generation](https://huggingface.co/papers/2607.18436) (2026)\n* [Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation](https://huggingface.co/papers/2608.26902) (2026)\n* [LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation](https://huggingface.co/papers/2608.28460) (2026)\n* [HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation](https://huggingface.co/papers/2607.20125) (2026)\n* [DensityKV: Density-Guided KV Cache Compression for Long Video Generation](https://huggingface.co/papers/2608.27922) (2026)\n* [Self Gradient Forcing: Native Long Video Extrapolation](https://huggingface.co/papers/2607.20368) (2026)\n* [AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report](https://huggingface.co/papers/2607.18367) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.18436\">Surprise Forcing: What to Remember, When to Skip in Long Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.26902\">Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.28460\">LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.20125\">HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.27922\">DensityKV: Density-Guided KV Cache Compression for Long Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.20368\">Self Gradient Forcing: Native Long Video Extrapolation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.18367\">AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-09-02T01:25:32.069Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6972655653953552},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26671","authors":[{"_id":"6a9735c8fe3c2f89286c3835","user":{"_id":"632a994b4ee2960462ecbf7e","avatarUrl":"/avatars/994e812be5a30d5908141215fdf18625.svg","isPro":true,"fullname":"Haiyang Xu","user":"xuhaiyang3110","type":"user","name":"xuhaiyang3110"},"name":"Haiyang Xu","status":"claimed_verified","statusLastChangedAt":"2026-09-02T00:45:04.549Z","hidden":false},{"_id":"6a9735c8fe3c2f89286c3836","name":"Zheng Ding","hidden":false},{"_id":"6a9735c8fe3c2f89286c3837","name":"Zhuowen Tu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/632a994b4ee2960462ecbf7e/AkcLp_k-xZGS91XTf6BuU.png"],"publishedAt":"2026-08-27T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"RECAP-Forcing: Retaining Content Appearances for Long Video Generation","submittedOnDailyBy":{"_id":"632a994b4ee2960462ecbf7e","avatarUrl":"/avatars/994e812be5a30d5908141215fdf18625.svg","isPro":true,"fullname":"Haiyang Xu","user":"xuhaiyang3110","type":"user","name":"xuhaiyang3110"},"summary":"Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content--such as entering subjects, disoccluded regions, and newly introduced scenes--at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance-indexed memory makes long-range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical-flow-based novelty bank extends the same principle by selectively retaining newly revealed content. As a training-free inference method with no additional learnable parameters, RECAP-Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.","upvotes":1,"discussionId":"6a9735c8fe3c2f89286c3838","projectPage":"https://xxuhaiyang.github.io/RECAP-Forcing/","githubRepo":"https://github.com/xXuHaiyang/RECAP-Forcing","githubRepoAddedBy":"user","ai_summary":"RECAP-Forcing improves long video generation by indexing memory according to appearance novelty rather than recency, preserving key-value caches for newly visible content to maintain long-range consistency without extra training.","ai_keywords":["autoregressive video generation","attention window","KV cache","appearance novelty","RECAP-Forcing","attention sink","optical-flow-based novelty bank","long-range consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8,"organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64ed876a74d9b58eabc769a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ed876a74d9b58eabc769a4/K4bJVW0FlqRtAAxJBJifR.jpeg","isPro":true,"fullname":"Boyang Wang","user":"HikariDawn","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26671.md","query":{}}">
RECAP-Forcing: Retaining Content Appearances for Long Video Generation
Abstract
RECAP-Forcing improves long video generation by indexing memory according to appearance novelty rather than recency, preserving key-value caches for newly visible content to maintain long-range consistency without extra training.
Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content--such as entering subjects, disoccluded regions, and newly introduced scenes--at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance-indexed memory makes long-range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical-flow-based novelty bank extends the same principle by selectively retaining newly revealed content. As a training-free inference method with no additional learnable parameters, RECAP-Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.
Community
TL;DR: Training-free appearance-indexed memory for long autoregressive video generation — retain content by what appears, not when it appeared.
Sliding-window memory forgets early content (subjects drift), while pinning reference frames freezes motion. RECAP-Forcing resolves this with two plug-in mechanisms on a frozen backbone:
- Reinforced attention sink — a single pre-softmax bias turns the first frames into memory the model actually retrieves;
- Optical-flow novelty bank — new appearances (entrances, disocclusions) are detected via flow and their original KV states stored verbatim in a fixed-size bank, retained by novelty rather than recency.
Zero learnable parameters. On Self-Forcing (Wan2.1-1.3B), Dynamic Degree jumps 27.5 → 58.1 while VBench Total improves 75.9 → 79.5; ~79% human preference over 216 blind pairwise judgments; identity stays stable in 5-minute rollouts. Generalizes across 4 causal video backbones with constant ~1.6× overhead.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.26671 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.26671 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.26671 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.