Hugging Face Daily Papers · · 4 min read

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Project Page: <a href=\"https://sizhezhao.github.io/projects/MaP-WAM/\" rel=\"nofollow\">https://sizhezhao.github.io/projects/MaP-WAM/</a><br>GitHub: <a href=\"https://github.com/aipixel/MaP-WAM\" rel=\"nofollow\">https://github.com/aipixel/MaP-WAM</a></p>\n","updatedAt":"2026-09-11T12:04:56.247Z","author":{"_id":"63f47b5321eb234ab739e91a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f47b5321eb234ab739e91a/vWfFNVtMkHl8gieha5PPd.jpeg","fullname":"Haozhe Xie","name":"hzxie","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":27,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7038599252700806},"editors":["hzxie"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/63f47b5321eb234ab739e91a/vWfFNVtMkHl8gieha5PPd.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.11561","authors":[{"_id":"6aa3de9c47a406da7901e99b","user":{"_id":"6769003931100198233acfc9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6769003931100198233acfc9/4HOr-exIRu9lyBt8JQWZ2.jpeg","isPro":false,"fullname":"SizheZhao","user":"SizheZhao","type":"user","name":"SizheZhao"},"name":"Sizhe Zhao","status":"claimed_verified","statusLastChangedAt":"2026-09-11T12:32:31.516Z","hidden":false},{"_id":"6aa3de9c47a406da7901e99c","user":{"_id":"63f47b5321eb234ab739e91a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f47b5321eb234ab739e91a/vWfFNVtMkHl8gieha5PPd.jpeg","isPro":false,"fullname":"Haozhe Xie","user":"hzxie","type":"user","name":"hzxie"},"name":"Haozhe Xie","status":"claimed_verified","statusLastChangedAt":"2026-09-11T12:32:29.662Z","hidden":false},{"_id":"6aa3de9c47a406da7901e99d","name":"Weiyu Zhao","hidden":false},{"_id":"6aa3de9c47a406da7901e99e","name":"Chenchu Zhang","hidden":false},{"_id":"6aa3de9c47a406da7901e99f","name":"Huan Wang","hidden":false},{"_id":"6aa3de9c47a406da7901e9a0","name":"Chenyang Wang","hidden":false},{"_id":"6aa3de9c47a406da7901e9a1","name":"Qinglin Liu","hidden":false},{"_id":"6aa3de9c47a406da7901e9a2","name":"Shengping Zhang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/63f47b5321eb234ab739e91a/clT7ifvITIQhw6ni64AqX.mp4"],"publishedAt":"2026-09-10T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"Memory as Plans: World-Action Modeling with Memory-Grounded Planning","submittedOnDailyBy":{"_id":"63f47b5321eb234ab739e91a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f47b5321eb234ab739e91a/vWfFNVtMkHl8gieha5PPd.jpeg","isPro":false,"fullname":"Haozhe Xie","user":"hzxie","type":"user","name":"hzxie"},"summary":"Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.","upvotes":7,"discussionId":"6aa3de9c47a406da7901e9a3","projectPage":"https://sizhezhao.github.io/projects/MaP-WAM/","githubRepo":"https://github.com/aipixel/MaP-WAM","githubRepoAddedBy":"user","ai_summary":"MaP-WAM improves non-Markovian robotic manipulation by separating memory-grounded planning from plan-conditioned execution, using compact episodic segment records and progress-calibrated action chunks to maintain fixed inference latency.","ai_keywords":["MaP-WAM","Memory-as-Plans","non-Markovian","world-action modeling","episodic memory","segment records","visual guidance","World-Action-Progress model","action chunks","execution progress","structured attention","key-value caching"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63f47b5321eb234ab739e91a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63f47b5321eb234ab739e91a/vWfFNVtMkHl8gieha5PPd.jpeg","isPro":false,"fullname":"Haozhe Xie","user":"hzxie","type":"user"},{"_id":"6769003931100198233acfc9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6769003931100198233acfc9/4HOr-exIRu9lyBt8JQWZ2.jpeg","isPro":false,"fullname":"SizheZhao","user":"SizheZhao","type":"user"},{"_id":"63fc70edb9db84750cea5ff5","avatarUrl":"/avatars/d27de31ba820d98a2542a1fe8fc6ed35.svg","isPro":false,"fullname":"Chenyang Wang","user":"cy-xq","type":"user"},{"_id":"65faee5d679333f5e0281bdc","avatarUrl":"/avatars/2ab67d9e106b9249a4732f378cf68ccc.svg","isPro":false,"fullname":"zhaoweiyu","user":"zhaoweiyu","type":"user"},{"_id":"6925a2d7d298877c448b20a0","avatarUrl":"/avatars/d9bec970a632517217d8c82bb22fdba5.svg","isPro":false,"fullname":"Colin Yoo","user":"Coldswamp","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6aa3f78017bb7af55ca557d9","avatarUrl":"/avatars/1939801f014176b7d762d921a24d4f63.svg","isPro":false,"fullname":"ZLY","user":"Hazestar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.11561.md","query":{}}">
Papers
arxiv:2609.11561

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

Published on Sep 10
· Submitted by
Haozhe Xie
on Sep 11
Authors:

Abstract

MaP-WAM improves non-Markovian robotic manipulation by separating memory-grounded planning from plan-conditioned execution, using compact episodic segment records and progress-calibrated action chunks to maintain fixed inference latency.

Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their combinations, and may therefore lose fine-grained visual evidence or face a trade-off between history coverage and execution efficiency. We introduce MaP-WAM, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history. MaP-WAM represents memory as completed segment records containing language instructions and sparse visual context, and converts this episodic memory into compact plans comprising the next segment-level language plan and corresponding visual guidance. A World-Action-Progress (WAP) model executes each plan over an unknown duration by jointly predicting action chunks and corresponding execution progress at inference time, calibrating predicted progress through plan-observation alignment for adaptive segment transitions and closed-loop context updates. MaP-WAM keeps the executor context length fixed, while structured attention further enables key-value caching in both planning and execution. MaP-WAM achieves state-of-the-art performance on RMBench with an 83.3% success rate and attains 78.0% success on real-robot tasks, while maintaining approximately constant executor inference latency as task history grows.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.11561
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.11561 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.11561 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.11561 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers