Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility space using a fixed-dimensional per-task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced-order parameterization increases the average feedback received by each utility coordinate and characterize the steady-state occupancy of erroneous coordinates under a generic coordinate-transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold-Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced-order utility states support efficient self-evolving agent memory while limiting persistent reward contamination. Code is available at: <a href=\"https://github.com/YOUNG-fnxm/RoMeRL\" rel=\"nofollow\">https://github.com/YOUNG-fnxm/RoMeRL</a></p>\n","updatedAt":"2026-08-11T02:38:58.412Z","author":{"_id":"68637e35b380b30a22ad00b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7Tf0RWoNFHwQv599ECMhD.png","fullname":"yi yang","name":"yangyiking","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8732135891914368},"editors":["yangyiking"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7Tf0RWoNFHwQv599ECMhD.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.02508","authors":[{"_id":"6a77feda8e9301703eaa5c45","user":{"_id":"68637e35b380b30a22ad00b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7Tf0RWoNFHwQv599ECMhD.png","isPro":false,"fullname":"yi yang","user":"yangyiking","type":"user","name":"yangyiking"},"name":"Yi Yang","status":"claimed_verified","statusLastChangedAt":"2026-08-09T08:45:04.518Z","hidden":false},{"_id":"6a77feda8e9301703eaa5c46","name":"Zhennan Chen","hidden":false},{"_id":"6a77feda8e9301703eaa5c47","user":{"_id":"673b5f24e863f1d28b402efc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673b5f24e863f1d28b402efc/TIs4lUcJ8lLlAeylr0Ang.jpeg","isPro":false,"fullname":"YiHong Zhuang","user":"utdawn","type":"user","name":"utdawn"},"name":"Yihong Zhuang","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.407Z","hidden":false},{"_id":"6a77feda8e9301703eaa5c48","name":"Tiehan Fan","hidden":false},{"_id":"6a77feda8e9301703eaa5c49","name":"Yinan Chen","hidden":false},{"_id":"6a77feda8e9301703eaa5c4a","name":"Jian Li","hidden":false},{"_id":"6a77feda8e9301703eaa5c4b","name":"Jian Yang","hidden":false},{"_id":"6a77feda8e9301703eaa5c4c","name":"Ying Tai","hidden":false}],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States","submittedOnDailyBy":{"_id":"68637e35b380b30a22ad00b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7Tf0RWoNFHwQv599ECMhD.png","isPro":false,"fullname":"yi yang","user":"yangyiking","type":"user","name":"yangyiking"},"summary":"Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility space using a fixed-dimensional per-task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced-order parameterization increases the average feedback received by each utility coordinate and characterize the steady-state occupancy of erroneous coordinates under a generic coordinate-transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold-Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced-order utility states support efficient self-evolving agent memory while limiting persistent reward contamination. Code is available at: https://github.com/YOUNG-fnxm/RoMeRL","upvotes":9,"discussionId":"6a77fedb8e9301703eaa5c4d","githubRepo":"https://github.com/YOUNG-fnxm/RoMeRL","githubRepoAddedBy":"user","ai_summary":"RoMeRL reduces trajectory-indexed memory utilities to fixed-dimensional per-task states to concentrate feedback, avoid reward contamination, and improve self-evolving LLM agent performance.","ai_keywords":["Reduced-Order Memory Reinforcement Learning","trajectory-indexed utilities","memory-reward trap","per-task memory state","outcome polarity","semantic coordinates","feedback density","Cold-Q ratio"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":8},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"68637e35b380b30a22ad00b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/7Tf0RWoNFHwQv599ECMhD.png","isPro":false,"fullname":"yi yang","user":"yangyiking","type":"user"},{"_id":"66449e619ff401732687f013","avatarUrl":"/avatars/251897d1324a70a9bf761513871c5841.svg","isPro":false,"fullname":"chen","user":"zhen-nan","type":"user"},{"_id":"65927f3b754092f6b1e187a7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65927f3b754092f6b1e187a7/gUrNvIQHmsl1vLwSUxpmL.jpeg","isPro":false,"fullname":"tiehan fan","user":"AnonMegumi","type":"user"},{"_id":"673b5f24e863f1d28b402efc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/673b5f24e863f1d28b402efc/TIs4lUcJ8lLlAeylr0Ang.jpeg","isPro":false,"fullname":"YiHong Zhuang","user":"utdawn","type":"user"},{"_id":"6486e362c9d77c91be1ee2bf","avatarUrl":"/avatars/c8a3f4d0cf59676f0ae0f49f5e478670.svg","isPro":false,"fullname":"Chengxi Li","user":"MrZixi","type":"user"},{"_id":"674d5e814e69219448bb2535","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/PaS4hKJ1s-wgcWIuuFxXP.png","isPro":false,"fullname":"ganjiarun","user":"gan777","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"69c9391eccb5f3b29e286351","avatarUrl":"/avatars/4965f0a84deb20b9ae7529fc76a6b281.svg","isPro":false,"fullname":"Yanyan","user":"testhugginglmf","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.02508.md","query":{}}">
RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
Abstract
RoMeRL reduces trajectory-indexed memory utilities to fixed-dimensional per-task states to concentrate feedback, avoid reward contamination, and improve self-evolving LLM agent performance.
Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility space using a fixed-dimensional per-task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced-order parameterization increases the average feedback received by each utility coordinate and characterize the steady-state occupancy of erroneous coordinates under a generic coordinate-transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold-Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced-order utility states support efficient self-evolving agent memory while limiting persistent reward contamination. Code is available at: https://github.com/YOUNG-fnxm/RoMeRL
Community
Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedback over an ever-expanding state space. Second, because trajectory-level rewards are jointly assigned to co-retrieved memories, irrelevant experiences may receive misleading utility updates and consequently enter the memory-reward trap. To address these challenges, we introduce Reduced-Order Memory Reinforcement Learning (RoMeRL), which represents the growing trajectory-indexed utility space using a fixed-dimensional per-task memory state factorized by outcome polarity and memory dynamics. RoMeRL incorporates new experiences through a fixed set of semantic coordinates whose contents are updated or replaced over time, thereby concentrating feedback over a bounded utility support. Theoretically, we show that this reduced-order parameterization increases the average feedback received by each utility coordinate and characterize the steady-state occupancy of erroneous coordinates under a generic coordinate-transition model. Empirically, across ALFWorld and LifelongAgentBench, RoMeRL improves task performance, reduces the Cold-Q ratio by 80.0%, increases feedback density by approximately 6.0 times, reduces the maintained memory size by 84.4%, and cuts LLM calls by 21.1%. These results show that reduced-order utility states support efficient self-evolving agent memory while limiting persistent reward contamination. Code is available at: https://github.com/YOUNG-fnxm/RoMeRL
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.02508 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.02508 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.02508 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.