Hugging Face Daily Papers · · 7 min read

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

TL;DR: Training-free appearance-indexed memory for long autoregressive video generation — retain content by <em>what appears</em>, not <em>when</em> it appeared.</p>\n<p>Sliding-window memory forgets early content (subjects drift), while pinning reference frames freezes motion. RECAP-Forcing resolves this with two plug-in mechanisms on a frozen backbone:</p>\n<ul>\n<li><strong>Reinforced attention sink</strong> — a single pre-softmax bias turns the first frames into memory the model actually retrieves;</li>\n<li><strong>Optical-flow novelty bank</strong> — new appearances (entrances, disocclusions) are detected via flow and their original KV states stored verbatim in a fixed-size bank, retained by novelty rather than recency.</li>\n</ul>\n<p>Zero learnable parameters. On Self-Forcing (Wan2.1-1.3B), Dynamic Degree jumps 27.5 → 58.1 while VBench Total improves 75.9 → 79.5; ~79% human preference over 216 blind pairwise judgments; identity stays stable in 5-minute rollouts. Generalizes across 4 causal video backbones with constant ~1.6× overhead.</p>\n","updatedAt":"2026-09-01T20:36:14.395Z","author":{"_id":"632a994b4ee2960462ecbf7e","avatarUrl":"/avatars/994e812be5a30d5908141215fdf18625.svg","fullname":"Haiyang Xu","name":"xuhaiyang3110","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8545826077461243},"editors":["xuhaiyang3110"],"editorAvatarUrls":["/avatars/994e812be5a30d5908141215fdf18625.svg"],"reactions":[],"isReport":false}},{"id":"6a977b0cb9c98429e282be5a","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false},"createdAt":"2026-09-02T01:25:32.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [Surprise Forcing: What to Remember, When to Skip in Long Video Generation](https://huggingface.co/papers/2607.18436) (2026)\n* [Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation](https://huggingface.co/papers/2608.26902) (2026)\n* [LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation](https://huggingface.co/papers/2608.28460) (2026)\n* [HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation](https://huggingface.co/papers/2607.20125) (2026)\n* [DensityKV: Density-Guided KV Cache Compression for Long Video Generation](https://huggingface.co/papers/2608.27922) (2026)\n* [Self Gradient Forcing: Native Long Video Extrapolation](https://huggingface.co/papers/2607.20368) (2026)\n* [AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report](https://huggingface.co/papers/2607.18367) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.18436\">Surprise Forcing: What to Remember, When to Skip in Long Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.26902\">Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.28460\">LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.20125\">HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.27922\">DensityKV: Density-Guided KV Cache Compression for Long Video Generation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.20368\">Self Gradient Forcing: Native Long Video Extrapolation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.18367\">AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-09-02T01:25:32.069Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6972655653953552},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26671","authors":[{"_id":"6a9735c8fe3c2f89286c3835","user":{"_id":"632a994b4ee2960462ecbf7e","avatarUrl":"/avatars/994e812be5a30d5908141215fdf18625.svg","isPro":true,"fullname":"Haiyang Xu","user":"xuhaiyang3110","type":"user","name":"xuhaiyang3110"},"name":"Haiyang Xu","status":"claimed_verified","statusLastChangedAt":"2026-09-02T00:45:04.549Z","hidden":false},{"_id":"6a9735c8fe3c2f89286c3836","name":"Zheng Ding","hidden":false},{"_id":"6a9735c8fe3c2f89286c3837","name":"Zhuowen Tu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/632a994b4ee2960462ecbf7e/AkcLp_k-xZGS91XTf6BuU.png"],"publishedAt":"2026-08-27T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"RECAP-Forcing: Retaining Content Appearances for Long Video Generation","submittedOnDailyBy":{"_id":"632a994b4ee2960462ecbf7e","avatarUrl":"/avatars/994e812be5a30d5908141215fdf18625.svg","isPro":true,"fullname":"Haiyang Xu","user":"xuhaiyang3110","type":"user","name":"xuhaiyang3110"},"summary":"Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content--such as entering subjects, disoccluded regions, and newly introduced scenes--at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance-indexed memory makes long-range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical-flow-based novelty bank extends the same principle by selectively retaining newly revealed content. As a training-free inference method with no additional learnable parameters, RECAP-Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.","upvotes":1,"discussionId":"6a9735c8fe3c2f89286c3838","projectPage":"https://xxuhaiyang.github.io/RECAP-Forcing/","githubRepo":"https://github.com/xXuHaiyang/RECAP-Forcing","githubRepoAddedBy":"user","ai_summary":"RECAP-Forcing improves long video generation by indexing memory according to appearance novelty rather than recency, preserving key-value caches for newly visible content to maintain long-range consistency without extra training.","ai_keywords":["autoregressive video generation","attention window","KV cache","appearance novelty","RECAP-Forcing","attention sink","optical-flow-based novelty bank","long-range consistency"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8,"organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64ed876a74d9b58eabc769a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ed876a74d9b58eabc769a4/K4bJVW0FlqRtAAxJBJifR.jpeg","isPro":true,"fullname":"Boyang Wang","user":"HikariDawn","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"697e87d12cc19315a8497001","name":"UCSanDiego","fullname":"University of California at San Diego","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/697e8687c00f332cf492d29e/KUQpvngxP4r9oBSDZwIwZ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26671.md","query":{}}">
Papers
arxiv:2608.26671

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

Published on Aug 27
· Submitted by
Haiyang Xu
on Sep 1
Authors:

Abstract

RECAP-Forcing improves long video generation by indexing memory according to appearance novelty rather than recency, preserving key-value caches for newly visible content to maintain long-range consistency without extra training.

Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly appearing content--such as entering subjects, disoccluded regions, and newly introduced scenes--at the moment it first becomes visible, prioritizing novelty over recency. Memory should scale with the amount of newly introduced content, rather than with video length. This appearance-indexed memory makes long-range consistency an explicit property of the memory structure. Our framework unifies two mechanisms under this single principle. At the beginning of a video, when all visible content is novel, an attention sink preserves the initial scene. As the video evolves, an optical-flow-based novelty bank extends the same principle by selectively retaining newly revealed content. As a training-free inference method with no additional learnable parameters, RECAP-Forcing consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

Community

Paper author Paper submitter about 5 hours ago

TL;DR: Training-free appearance-indexed memory for long autoregressive video generation — retain content by what appears, not when it appeared.

Sliding-window memory forgets early content (subjects drift), while pinning reference frames freezes motion. RECAP-Forcing resolves this with two plug-in mechanisms on a frozen backbone:

  • Reinforced attention sink — a single pre-softmax bias turns the first frames into memory the model actually retrieves;
  • Optical-flow novelty bank — new appearances (entrances, disocclusions) are detected via flow and their original KV states stored verbatim in a fixed-size bank, retained by novelty rather than recency.

Zero learnable parameters. On Self-Forcing (Wan2.1-1.3B), Dynamic Degree jumps 27.5 → 58.1 while VBench Total improves 75.9 → 79.5; ~79% human preference over 216 blind pairwise judgments; identity stays stable in 5-minute rollouts. Generalizes across 4 causal video backbones with constant ~1.6× overhead.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.26671
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.26671 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.26671 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.26671 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers