Hugging Face Daily Papers · · 4 min read

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments.<br>Although searchable, such fragments are not generation-ready: language models must reconstruct cross-modal and temporal alignments at inference time, when context is limited and attribution is difficult.<br>We propose EM^2Mem, an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors during memory construction.<br>Each event-indexed memory cell aligns multimodal records, temporal context, graph-linked relations, semantic facts, and provenance, enabling compact evidence readout over grounded multimodal events rather than modality-specific fragments.<br>Across three long-video QA benchmarks, EM^2Mem improves average accuracy over the strongest memory baseline by 2.0, 2.4, and 3.7 points, improves strict event-level Top-5 evidence recall by 7.0 points, and reduces per-query latency by 4.67 times and total inference tokens by 63.66% (The code will be integrated into \\url{<a href=\"https://github.com/zjunlp/LightMem%7D\" rel=\"nofollow\">https://github.com/zjunlp/LightMem}</a>).</p>\n","updatedAt":"2026-09-02T01:49:51.697Z","author":{"_id":"6441f1d2603214724ec0c1c2","avatarUrl":"/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg","fullname":"Shumin Deng","name":"231sm","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8738232851028442},"editors":["231sm"],"editorAvatarUrls":["/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00551","authors":[{"_id":"6a97808cfe3c2f89286c38df","name":"Yijun Chen","hidden":false},{"_id":"6a97808cfe3c2f89286c38e0","name":"Yaqi Zheng","hidden":false},{"_id":"6a97808cfe3c2f89286c38e1","name":"Yanya Li","hidden":false},{"_id":"6a97808cfe3c2f89286c38e2","name":"Boyi Xiao","hidden":false},{"_id":"6a97808cfe3c2f89286c38e3","name":"Buqiang Xu","hidden":false},{"_id":"6a97808cfe3c2f89286c38e4","name":"Shuofei Qiao","hidden":false},{"_id":"6a97808cfe3c2f89286c38e5","name":"Jizhan Fang","hidden":false},{"_id":"6a97808cfe3c2f89286c38e6","name":"Xinle Deng","hidden":false},{"_id":"6a97808cfe3c2f89286c38e7","name":"Yunzhi Yao","hidden":false},{"_id":"6a97808cfe3c2f89286c38e8","name":"Xuehai Wang","hidden":false},{"_id":"6a97808cfe3c2f89286c38e9","name":"Liuxin Zhang","hidden":false},{"_id":"6a97808cfe3c2f89286c38ea","name":"Hui Li","hidden":false},{"_id":"6a97808cfe3c2f89286c38eb","name":"Huajun Chen","hidden":false},{"_id":"6a97808cfe3c2f89286c38ec","name":"Shumin Deng","hidden":false}],"publishedAt":"2026-09-01T00:00:00.000Z","submittedOnDailyAt":"2026-09-02T00:00:00.000Z","title":"EM^2Mem: Event-Centric Multimodal Memory for Large Language Models","submittedOnDailyBy":{"_id":"6441f1d2603214724ec0c1c2","avatarUrl":"/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg","isPro":false,"fullname":"Shumin Deng","user":"231sm","type":"user","name":"231sm"},"summary":"Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragments are not generation-ready: language models must reconstruct cross-modal and temporal alignments at inference time, when context is limited and attribution is difficult. We propose EM^2Mem, an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors during memory construction. Each event-indexed memory cell aligns multimodal records, temporal context, graph-linked relations, semantic facts, and provenance, enabling compact evidence readout over grounded multimodal events rather than modality-specific fragments. Across three long-video QA benchmarks, EM^2Mem improves average accuracy over the strongest memory baseline by 2.0, 2.4, and 3.7 points, improves strict event-level Top-5 evidence recall by 7.0 points, and reduces per-query latency by 4.67 times and total inference tokens by 63.66% (The code will be integrated into https://github.com/zjunlp/LightMem).","upvotes":6,"discussionId":"6a97808cfe3c2f89286c38ed","ai_summary":"EM²Mem binds multimodal evidence to event anchors for compact, generation-ready memory in long-video question answering.","ai_keywords":["multimodal memory","event-centric memory","cross-modal alignment","temporal context","graph-linked relations","provenance","evidence recall","long-video QA"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"620a6fcd8d5e5dfed284bc91","name":"zjunlp","fullname":"ZJUNLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1644851027419-620a61cba53066560e226d30.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6441f1d2603214724ec0c1c2","avatarUrl":"/avatars/d3c4b759e6a5635e37ff715fae52e5ba.svg","isPro":false,"fullname":"Shumin Deng","user":"231sm","type":"user"},{"_id":"6a8196b509dddac72f26fab1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8196b509dddac72f26fab1/y0jcikLGLHKlMJErdMnB5.jpeg","isPro":false,"fullname":"Leon","user":"itsle-onguo","type":"user"},{"_id":"620b3bbb0668e435407c8d0a","avatarUrl":"/avatars/e0fccbb2577d76088e09f054c35cffbc.svg","isPro":true,"fullname":"Ningyu Zhang","user":"Ningyu","type":"user"},{"_id":"6a1443f02a9759cfbdf80a48","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a1443f02a9759cfbdf80a48/dIoHGMzJ3U6k7VTOaBd4Y.jpeg","isPro":false,"fullname":"Haoxiong Wang","user":"WangHX2026","type":"user"},{"_id":"65535b54140fc44a74d43635","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/MIrD8OzDKF2aI38i7ZPjR.jpeg","isPro":false,"fullname":"Zhisong Qiu","user":"consultantQ","type":"user"},{"_id":"679a01a99893a68681ef1847","avatarUrl":"/avatars/17fe173acda467df2b90cca9e5f3c656.svg","isPro":false,"fullname":"ye","user":"haohaojun","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"620a6fcd8d5e5dfed284bc91","name":"zjunlp","fullname":"ZJUNLP","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1644851027419-620a61cba53066560e226d30.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00551.md","query":{}}">
Papers
arxiv:2609.00551

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

Published on Sep 1
· Submitted by
Shumin Deng
on Sep 2
Authors:
,

Abstract

EM²Mem binds multimodal evidence to event anchors for compact, generation-ready memory in long-video question answering.

Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments. Although searchable, such fragments are not generation-ready: language models must reconstruct cross-modal and temporal alignments at inference time, when context is limited and attribution is difficult. We propose EM^2Mem, an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors during memory construction. Each event-indexed memory cell aligns multimodal records, temporal context, graph-linked relations, semantic facts, and provenance, enabling compact evidence readout over grounded multimodal events rather than modality-specific fragments. Across three long-video QA benchmarks, EM^2Mem improves average accuracy over the strongest memory baseline by 2.0, 2.4, and 3.7 points, improves strict event-level Top-5 evidence recall by 7.0 points, and reduces per-query latency by 4.67 times and total inference tokens by 63.66% (The code will be integrated into https://github.com/zjunlp/LightMem).

Community

Paper submitter about 6 hours ago

Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, summaries, or graph facts as isolated fragments.
Although searchable, such fragments are not generation-ready: language models must reconstruct cross-modal and temporal alignments at inference time, when context is limited and attribution is difficult.
We propose EM^2Mem, an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors during memory construction.
Each event-indexed memory cell aligns multimodal records, temporal context, graph-linked relations, semantic facts, and provenance, enabling compact evidence readout over grounded multimodal events rather than modality-specific fragments.
Across three long-video QA benchmarks, EM^2Mem improves average accuracy over the strongest memory baseline by 2.0, 2.4, and 3.7 points, improves strict event-level Top-5 evidence recall by 7.0 points, and reduces per-query latency by 4.67 times and total inference tokens by 63.66% (The code will be integrated into \url{https://github.com/zjunlp/LightMem}).

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.00551
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.00551 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.00551 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.00551 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers