LatentStream advances streaming video understanding from external “store-and-retrieve” memory toward “retrieve-and-internalize,” progressively consolidating historical evidence into a compact, evolving latent working memory. By combining hierarchical memory consolidation, expanding latent receptive fields, and confidence-guided optimization, it achieves state-of-the-art performance across online and offline video benchmarks under bounded memory. Our code will be available at <a href=\"https://github.com/quhongyu/LatentStream\" rel=\"nofollow\">https://github.com/quhongyu/LatentStream</a>.</p>\n","updatedAt":"2026-09-04T04:27:13.938Z","author":{"_id":"6a4358768ae6ed7930e6b72e","avatarUrl":"/avatars/f62d7aad1650b3d091c73a7a6727146d.svg","fullname":"Qu Hongyu","name":"quhongyu123","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.8220832943916321},"editors":["quhongyu123"],"editorAvatarUrls":["/avatars/f62d7aad1650b3d091c73a7a6727146d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.04131","authors":[{"_id":"6a9a41c78f7c3b75572394ef","name":"Hongyu Qu","hidden":false},{"_id":"6a9a41c78f7c3b75572394f0","name":"Guangming Yao","hidden":false},{"_id":"6a9a41c78f7c3b75572394f1","name":"Ling Xing","hidden":false},{"_id":"6a9a41c78f7c3b75572394f2","name":"Xiaobin Hu","hidden":false},{"_id":"6a9a41c78f7c3b75572394f3","name":"Rongxing Ding","hidden":false},{"_id":"6a9a41c78f7c3b75572394f4","name":"Guibin Zhang","hidden":false},{"_id":"6a9a41c78f7c3b75572394f5","name":"Fan Zhang","hidden":false},{"_id":"6a9a41c78f7c3b75572394f6","name":"Yi Yuan","hidden":false},{"_id":"6a9a41c78f7c3b75572394f7","name":"Xiangbo Shu","hidden":false},{"_id":"6a9a41c78f7c3b75572394f8","name":"Shuicheng Yan","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding","submittedOnDailyBy":{"_id":"6a4358768ae6ed7930e6b72e","avatarUrl":"/avatars/f62d7aad1650b3d091c73a7a6727146d.svg","isPro":false,"fullname":"Qu Hongyu","user":"quhongyu123","type":"user","name":"quhongyu123"},"summary":"Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical observations into an external memory bank and retrieve query-relevant evidence as additional visual context. Though effective, this store-and-retrieve paradigm keeps historical evidence as external visual context, preventing it from being internalized into a compact, evolving latent memory that can continuously guide streaming reasoning. To bridge this gap, we introduce LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize. Specifically, LatentStream comprises three coordinated components. First, Query-agnostic Hierarchical Streaming Memory organizes visual history into short-, mid-, and long-term levels under a fixed memory budget through Jenks-guided adaptive consolidation. Once a query arrives, Hierarchical Latent Memory Evolution equips groups of latent memory tokens with progressively expanding memory receptive fields, enabling them to iteratively retrieve historical evidence from their corresponding scopes and internalize it into a compact, fixed-length latent memory. Finally, Progressive Confidence-guided Latent Memory Optimization constructs a hierarchical progression reward from group-wise predictive entropy and jointly refines the latent memory tokens and retrieved evidence, encouraging increasingly confident streaming reasoning. Extensive experiments demonstrate that LatentStream achieves new state-of-the-art results on existing online and offline video benchmarks.","upvotes":15,"discussionId":"6a9a41c78f7c3b75572394f9","projectPage":"https://github.com/quhongyu/LatentStream","ai_summary":"LatentStream introduces a progressive latent working memory framework that internalizes streaming visual evidence into compact evolving tokens for continuous reasoning.","ai_keywords":["multimodal large language models","streaming video understanding","latent working memory","hierarchical streaming memory","Jenks-guided adaptive consolidation","latent memory tokens","memory receptive fields","predictive entropy","confidence-guided optimization"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66ebce4f322dcfe976c30460","avatarUrl":"/avatars/aafb9e5e639861aa711679d76211a44e.svg","isPro":false,"fullname":"Ling Xing","user":"ling441","type":"user"},{"_id":"6a4f70ed2d71767867d66876","avatarUrl":"/avatars/bee31513c3f8a9a73d6e43e1fff3223e.svg","isPro":false,"fullname":"Levinia Bu","user":"buwenli","type":"user"},{"_id":"6a4f08d8567298c12d8b6f6c","avatarUrl":"/avatars/0707ff4ee2e65d00ce3556d257ad518d.svg","isPro":false,"fullname":"z","user":"hugsvip","type":"user"},{"_id":"680f0a912b588ca79dce7754","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/680f0a912b588ca79dce7754/Ipvi8ZEpCusH3gqLbMOlt.png","isPro":false,"fullname":"GeorgeHu","user":"GeorgeHu6","type":"user"},{"_id":"6639ad487c0ab4fd9df1dde5","avatarUrl":"/avatars/8cc99f6ed8f8c1b2a14dde797a991a8c.svg","isPro":false,"fullname":"Fan Zhang","user":"Karl28","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"},{"_id":"68cd08739e0a3f5e6b6e7575","avatarUrl":"/avatars/69122f8eea1b737f33b7bd74c0f063b4.svg","isPro":false,"fullname":"sherry","user":"Jack-sherry","type":"user"},{"_id":"6a4358768ae6ed7930e6b72e","avatarUrl":"/avatars/f62d7aad1650b3d091c73a7a6727146d.svg","isPro":false,"fullname":"Qu Hongyu","user":"quhongyu123","type":"user"},{"_id":"6a4fa3f1bd4a917d8496634c","avatarUrl":"/avatars/3d5f173fa510e2d977b3ac0627643f61.svg","isPro":false,"fullname":"xiongzihao","user":"xiongzihao","type":"user"},{"_id":"68e745d93f652a3945dd8d64","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/iXVHnYYfWLJdAUpMeLho4.png","isPro":false,"fullname":"刘洋","user":"DimXion","type":"user"},{"_id":"664567f02894e9815819710b","avatarUrl":"/avatars/cfac9a4372f1be5ac9a24194b93d767d.svg","isPro":false,"fullname":"yichufan","user":"yichufan","type":"user"},{"_id":"6a4f1de5bcdb9d49a5203866","avatarUrl":"/avatars/df15c55cf208bde52b3380e893b01403.svg","isPro":false,"fullname":"xiao duan","user":"chenzzzzx","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.04131.md","query":{}}">
Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
Abstract
LatentStream introduces a progressive latent working memory framework that internalizes streaming visual evidence into compact evolving tokens for continuous reasoning.
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical observations into an external memory bank and retrieve query-relevant evidence as additional visual context. Though effective, this store-and-retrieve paradigm keeps historical evidence as external visual context, preventing it from being internalized into a compact, evolving latent memory that can continuously guide streaming reasoning. To bridge this gap, we introduce LatentStream, a progressive latent working memory framework that shifts streaming memory from store-and-retrieve to retrieve-and-internalize. Specifically, LatentStream comprises three coordinated components. First, Query-agnostic Hierarchical Streaming Memory organizes visual history into short-, mid-, and long-term levels under a fixed memory budget through Jenks-guided adaptive consolidation. Once a query arrives, Hierarchical Latent Memory Evolution equips groups of latent memory tokens with progressively expanding memory receptive fields, enabling them to iteratively retrieve historical evidence from their corresponding scopes and internalize it into a compact, fixed-length latent memory. Finally, Progressive Confidence-guided Latent Memory Optimization constructs a hierarchical progression reward from group-wise predictive entropy and jointly refines the latent memory tokens and retrieved evidence, encouraging increasingly confident streaming reasoning. Extensive experiments demonstrate that LatentStream achieves new state-of-the-art results on existing online and offline video benchmarks.
Community
LatentStream advances streaming video understanding from external “store-and-retrieve” memory toward “retrieve-and-internalize,” progressively consolidating historical evidence into a compact, evolving latent working memory. By combining hierarchical memory consolidation, expanding latent receptive fields, and confidence-guided optimization, it achieves state-of-the-art performance across online and offline video benchmarks under bounded memory. Our code will be available at https://github.com/quhongyu/LatentStream.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.04131 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.04131 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.04131 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.