Hugging Face Daily Papers · · 5 min read

MemHarness: Memory Is Reconstructed, Not Replayed

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Most memory-augmented LLM agents follow a retrieve-and-replay paradigm, directly inserting retrieved trajectories into the context. This conflates retrieval relevance with action-level applicability: a memory may be semantically related yet inappropriate for the current state, since it was formed under different conditions. Inspired by the reconstructive nature of human memory—recall reorganizes past experience using present cues rather than reproducing it verbatim—we recast memory-guided decision-making as a five-stage process: observation, retrieval, critique, reconstruction, and action. We introduce MemHarness, which inserts explicit memory critique and reconstruction between retrieval and action, turning static records into context-sensitive guidance while retaining their traceability. Crucially, this reconstructive ability requires no human annotation: it emerges end-to-end via GRPO by optimizing for task success. On ALFWorld and WebShop, MemHarness consistently outperforms both pure RL and static memory-augmented baselines, and ablations confirm that adaptive reconstruction—not retrieval alone—is the primary driver of the gains.</p>\n","updatedAt":"2026-07-31T01:56:44.599Z","author":{"_id":"66fa9fa56b620c614cdb1b57","avatarUrl":"/avatars/907c101ba4528de05d2ce4bb1ad82005.svg","fullname":"Rong Wu","name":"Edaizi","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8808912634849548},"editors":["Edaizi"],"editorAvatarUrls":["/avatars/907c101ba4528de05d2ce4bb1ad82005.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28272","authors":[{"_id":"6a6bff297bd25d8874c070ad","name":"Rong Wu","hidden":false},{"_id":"6a6bff297bd25d8874c070ae","name":"Daocheng Fu","hidden":false},{"_id":"6a6bff297bd25d8874c070af","name":"Licheng Wen","hidden":false},{"_id":"6a6bff297bd25d8874c070b0","name":"Xuemeng Yang","hidden":false},{"_id":"6a6bff297bd25d8874c070b1","name":"Shu Zou","hidden":false},{"_id":"6a6bff297bd25d8874c070b2","name":"Jianbiao Mei","hidden":false},{"_id":"6a6bff297bd25d8874c070b3","name":"Yuxin Wang","hidden":false},{"_id":"6a6bff297bd25d8874c070b4","name":"Hairong Zhang","hidden":false},{"_id":"6a6bff297bd25d8874c070b5","name":"Yu Yang","hidden":false},{"_id":"6a6bff297bd25d8874c070b6","name":"Tao Hu","hidden":false},{"_id":"6a6bff297bd25d8874c070b7","name":"Cong Zhang","hidden":false},{"_id":"6a6bff297bd25d8874c070b8","name":"Botian Shi","hidden":false},{"_id":"6a6bff297bd25d8874c070b9","name":"Pinlong Cai","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"MemHarness: Memory Is Reconstructed, Not Replayed","submittedOnDailyBy":{"_id":"66fa9fa56b620c614cdb1b57","avatarUrl":"/avatars/907c101ba4528de05d2ce4bb1ad82005.svg","isPro":false,"fullname":"Rong Wu","user":"Edaizi","type":"user","name":"Edaizi"},"summary":"Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's current situation. This ``replay'' paradigm ignores the gap between the abstract, general nature of stored experience and the concrete, ever-changing states encountered at decision time, frequently causing negative transfer. In contrast, humans rarely recall past experiences verbatim; instead, they reorganize and adapt retrieved memories to fit the present context. Inspired by this, we propose MemHarness, a framework that equips LLM agents to actively harness and reconstruct past experiences based on the present context. At each decision step, a unified policy model critiques and reconstructs the retrieved experience conditioned on the current state, producing context-grounded guidance before acting. This reconstructive ability emerges naturally through end-to-end training with GRPO. Experiments on ALFWorld and WebShop show that MemHarness substantially outperforms pure RL and static memory-augmented baselines, demonstrating strong robustness in out-of-distribution (OOD) scenarios. Furthermore, our analyses reveal that this reconstruction objective not only prevents negative transfer but also serves as latent guidance during training, fundamentally improving the agent's intrinsic reasoning capabilities.","upvotes":12,"discussionId":"6a6bff297bd25d8874c070ba","projectPage":"https://arxiv.org/abs/2607.28272","githubRepo":"https://github.com/KnowledgeXLab/MemHarness","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"68e8716fa897e565aaec87ca","name":"KnowledgeXLab","fullname":"KnowledgeXLab@Shanghai AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a7c43ae940d769194055df/oPtAVxSeVl8XHm5mo0Nq8.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66fa9fa56b620c614cdb1b57","avatarUrl":"/avatars/907c101ba4528de05d2ce4bb1ad82005.svg","isPro":false,"fullname":"Rong Wu","user":"Edaizi","type":"user"},{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64b289681917256b41140bf7","avatarUrl":"/avatars/22309802f680ad38310d53d4b008ab21.svg","isPro":false,"fullname":"YcHades","user":"YcHades","type":"user"},{"_id":"67e41575baa288f7f82af73e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67e41575baa288f7f82af73e/0Wx_4SmB6XsjUghQZTRLj.jpeg","isPro":false,"fullname":"YANG YU","user":"yuyang-cloud","type":"user"},{"_id":"65827953d73d6402f7ae332d","avatarUrl":"/avatars/d57eae62171ece7a237eaf4ea938a31c.svg","isPro":false,"fullname":"Mei","user":"Jianbiao","type":"user"},{"_id":"656882da5025f8e01b424eda","avatarUrl":"/avatars/7fda244634c90b4647b49f83fdb25a9b.svg","isPro":false,"fullname":"yxm","user":"jokester-yxm","type":"user"},{"_id":"6350c89759bfa9a85d434138","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1666238674117-6350c89759bfa9a85d434138.jpeg","isPro":false,"fullname":"Yang Lee","user":"innovation64","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"},{"_id":"6a046a20df2cee54a73927bf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a046a20df2cee54a73927bf/XO_-R8gfimjsGa8TpDJpA.jpeg","isPro":false,"fullname":"Katelyn Powers","user":"KatelynPowers","type":"user"},{"_id":"64d4615cf8082bf19b916492","avatarUrl":"/avatars/8e1b59565ec5e4b31090cf1b911781b9.svg","isPro":false,"fullname":"wongyukim","user":"wongyukim","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68e8716fa897e565aaec87ca","name":"KnowledgeXLab","fullname":"KnowledgeXLab@Shanghai AI Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a7c43ae940d769194055df/oPtAVxSeVl8XHm5mo0Nq8.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.28272.md","query":{}}">
Papers
arxiv:2607.28272

MemHarness: Memory Is Reconstructed, Not Replayed

Published on Jul 30
· Submitted by
Rong Wu
on Jul 31
Authors:
,

Abstract

Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's current situation. This ``replay'' paradigm ignores the gap between the abstract, general nature of stored experience and the concrete, ever-changing states encountered at decision time, frequently causing negative transfer. In contrast, humans rarely recall past experiences verbatim; instead, they reorganize and adapt retrieved memories to fit the present context. Inspired by this, we propose MemHarness, a framework that equips LLM agents to actively harness and reconstruct past experiences based on the present context. At each decision step, a unified policy model critiques and reconstructs the retrieved experience conditioned on the current state, producing context-grounded guidance before acting. This reconstructive ability emerges naturally through end-to-end training with GRPO. Experiments on ALFWorld and WebShop show that MemHarness substantially outperforms pure RL and static memory-augmented baselines, demonstrating strong robustness in out-of-distribution (OOD) scenarios. Furthermore, our analyses reveal that this reconstruction objective not only prevents negative transfer but also serves as latent guidance during training, fundamentally improving the agent's intrinsic reasoning capabilities.

Community

Paper submitter about 8 hours ago

Most memory-augmented LLM agents follow a retrieve-and-replay paradigm, directly inserting retrieved trajectories into the context. This conflates retrieval relevance with action-level applicability: a memory may be semantically related yet inappropriate for the current state, since it was formed under different conditions. Inspired by the reconstructive nature of human memory—recall reorganizes past experience using present cues rather than reproducing it verbatim—we recast memory-guided decision-making as a five-stage process: observation, retrieval, critique, reconstruction, and action. We introduce MemHarness, which inserts explicit memory critique and reconstruction between retrieval and action, turning static records into context-sensitive guidance while retaining their traceability. Crucially, this reconstructive ability requires no human annotation: it emerges end-to-end via GRPO by optimizing for task success. On ALFWorld and WebShop, MemHarness consistently outperforms both pure RL and static memory-augmented baselines, and ablations confirm that adaptive reconstruction—not retrieval alone—is the primary driver of the gains.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.28272
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.28272 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.28272 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.28272 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers