Hugging Face Daily Papers · · 4 min read

Memory for Large Language Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

\n\t<a id=\"💡-a-comprehensive-survey-on-architectural-level-memory-in-large-language-models\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#💡-a-comprehensive-survey-on-architectural-level-memory-in-large-language-models\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\t💡 A Comprehensive Survey on Architectural-Level Memory in Large Language Models\n\t</span>\n</h3>\n<p><strong>Summary:</strong><br>This survey from Tsinghua, NUS, and Bosch AI provides a unified theoretical framework for understanding architectural-level memory in LLMs, explicitly distinguishing it from external agent-based memory systems. The authors introduce a novel 3D taxonomy categorizing memory mechanisms by Representation (Implicit vs. Explicit), Update Dynamics (Offline vs. Online), and Persistence (Short-term vs. Long-term). It systematically maps the paradigm shift from computational byproducts like attention KV Caches and recurrent hidden states to explicitly addressable modules, such as Titans, TTT, and Engram. Highly recommended for researchers focusing on long-context scaling, hybrid architectures, and algorithm-hardware co-design.</p>\n","updatedAt":"2026-07-30T10:24:50.550Z","author":{"_id":"6621d759c367a8f13d00ad57","avatarUrl":"/avatars/575cedf7249943c021be22638f8e84aa.svg","fullname":"Sining Zhoubian","name":"SiningZhou","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8200857639312744},"editors":["SiningZhou"],"editorAvatarUrls":["/avatars/575cedf7249943c021be22638f8e84aa.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.25380","authors":[{"_id":"6a6b1fb7b2106777884ab058","name":"Sining Zhoubian","hidden":false},{"_id":"6a6b1fb7b2106777884ab059","name":"Dan Zhang","hidden":false},{"_id":"6a6b1fb7b2106777884ab05a","name":"Evgeny Kharlamov","hidden":false},{"_id":"6a6b1fb7b2106777884ab05b","name":"Jie Tang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6621d759c367a8f13d00ad57/NsLy7qgM7XRYHvkSzg2BJ.png","https://cdn-uploads.huggingface.co/production/uploads/6621d759c367a8f13d00ad57/LEJMwMjFAwvczOhvCXGrI.png","https://cdn-uploads.huggingface.co/production/uploads/6621d759c367a8f13d00ad57/Xp5LJft620ZBIylCyQgKk.png"],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"Memory for Large Language Models","submittedOnDailyBy":{"_id":"6621d759c367a8f13d00ad57","avatarUrl":"/avatars/575cedf7249943c021be22638f8e84aa.svg","isPro":false,"fullname":"Sining Zhoubian","user":"SiningZhou","type":"user","name":"SiningZhou"},"summary":"Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention, recurrent state dynamics, parameter-efficient adaptations, and scalable lookup storage---this rapid evolution has led to a highly fragmented research landscape. In this survey, we present a systematic, architecture-centric taxonomy of memory in LLMs. Our framework characterizes memory along three orthogonal axes: representation (implicit versus explicit), update dynamics (offline versus online), and persistence (short-term versus long-term). We further formalize the granular mechanisms dictating memory writing, routing, state transitions, and consolidation. This unified perspective elucidates the conceptual boundaries between computation-coupled and independently addressable memory, effectively bridging disparate architectural paradigms. Additionally, we critically analyze hybrid memory architectures, system-level efficiency trade-offs, and multi-dimensional evaluation methodologies. By consolidating these scattered advancements into a cohesive framework, this survey charts the trajectory of memory-centric LLM design and provides a principled foundation for future innovations in scalable and adaptive language modeling.","upvotes":4,"discussionId":"6a6b1fb7b2106777884ab05c","projectPage":"https://arxiv.org/abs/2607.25380","organization":{"_id":"64db4fc57266618e854318f4","name":"THU-KEG","fullname":"Knowledge Engineer Group @ Tsinghua University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648c4b46e549be47af1aafcd/5atqdE9AUWvYAHm9FNkG_.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6621d759c367a8f13d00ad57","avatarUrl":"/avatars/575cedf7249943c021be22638f8e84aa.svg","isPro":false,"fullname":"Sining Zhoubian","user":"SiningZhou","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6351e5bb3734c6e8a5c1bec1","avatarUrl":"/avatars/a784a51b369b197398575c3afbd5ceab.svg","isPro":false,"fullname":"Han-Bit Kang","user":"hbkang","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64db4fc57266618e854318f4","name":"THU-KEG","fullname":"Knowledge Engineer Group @ Tsinghua University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/648c4b46e549be47af1aafcd/5atqdE9AUWvYAHm9FNkG_.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.25380.md","query":{}}">
Papers
arxiv:2607.25380

Memory for Large Language Models

Published on Jul 28
· Submitted by
Sining Zhoubian
on Jul 30
Authors:
,

Abstract

Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a spectrum of explicit, controllable mechanisms. While recent advances introduce diverse strategies---spanning transient attention, recurrent state dynamics, parameter-efficient adaptations, and scalable lookup storage---this rapid evolution has led to a highly fragmented research landscape. In this survey, we present a systematic, architecture-centric taxonomy of memory in LLMs. Our framework characterizes memory along three orthogonal axes: representation (implicit versus explicit), update dynamics (offline versus online), and persistence (short-term versus long-term). We further formalize the granular mechanisms dictating memory writing, routing, state transitions, and consolidation. This unified perspective elucidates the conceptual boundaries between computation-coupled and independently addressable memory, effectively bridging disparate architectural paradigms. Additionally, we critically analyze hybrid memory architectures, system-level efficiency trade-offs, and multi-dimensional evaluation methodologies. By consolidating these scattered advancements into a cohesive framework, this survey charts the trajectory of memory-centric LLM design and provides a principled foundation for future innovations in scalable and adaptive language modeling.

Community

Paper submitter about 16 hours ago

💡 A Comprehensive Survey on Architectural-Level Memory in Large Language Models

Summary:
This survey from Tsinghua, NUS, and Bosch AI provides a unified theoretical framework for understanding architectural-level memory in LLMs, explicitly distinguishing it from external agent-based memory systems. The authors introduce a novel 3D taxonomy categorizing memory mechanisms by Representation (Implicit vs. Explicit), Update Dynamics (Offline vs. Online), and Persistence (Short-term vs. Long-term). It systematically maps the paradigm shift from computational byproducts like attention KV Caches and recurrent hidden states to explicitly addressable modules, such as Titans, TTT, and Engram. Highly recommended for researchers focusing on long-context scaling, hybrid architectures, and algorithm-hardware co-design.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.25380
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.25380 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.25380 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.25380 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers