Hugging Face Daily Papers · · 3 min read

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation</p>\n","updatedAt":"2026-07-09T02:13:58.960Z","author":{"_id":"67d63e228d5c7a132cbcf39b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png","fullname":"neil yu","name":"yxl66666","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4796733558177948},"editors":["yxl66666"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.07608","authors":[{"_id":"6a4efdfe48d70828b718ddee","name":"Hongyu Qu","hidden":false},{"_id":"6a4efdfe48d70828b718ddef","name":"Jianzhe Gao","hidden":false},{"_id":"6a4efdfe48d70828b718ddf0","name":"Xiaobin Hu","hidden":false},{"_id":"6a4efdfe48d70828b718ddf1","name":"Shaohuan Yang","hidden":false},{"_id":"6a4efdfe48d70828b718ddf2","name":"Xinlei Yu","hidden":false},{"_id":"6a4efdfe48d70828b718ddf3","name":"Rui Yan","hidden":false},{"_id":"6a4efdfe48d70828b718ddf4","name":"Wenguan Wang","hidden":false},{"_id":"6a4efdfe48d70828b718ddf5","name":"Xiangbo Shu","hidden":false},{"_id":"6a4efdfe48d70828b718ddf6","name":"Shuicheng Yan","hidden":false}],"publishedAt":"2026-07-08T00:00:00.000Z","submittedOnDailyAt":"2026-07-09T00:00:00.000Z","title":"Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation","submittedOnDailyBy":{"_id":"67d63e228d5c7a132cbcf39b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png","isPro":false,"fullname":"neil yu","user":"yxl66666","type":"user","name":"yxl66666"},"summary":"Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.","upvotes":37,"discussionId":"6a4efdff48d70828b718ddf7","githubRepo":"https://github.com/quhongyu/LaMem-VLA","githubRepoAddedBy":"user","ai_summary":"LaMem-VLA introduces a latent-memory-native framework that integrates historical experience into vision-language-action reasoning through coordinated memory components operating in the same latent space.","ai_keywords":["Vision-Language-Action models","Markovian assumption","memory-augmented VLAs","latent embedding space","latent memory tokens","short-term memory vaults","long-term memory vaults","multimodal cognition","context-relevant evidence","compact latent memory tokens","continuous embedding sequence","bounded context"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67d63e228d5c7a132cbcf39b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png","isPro":false,"fullname":"neil yu","user":"yxl66666","type":"user"},{"_id":"65ed5d4433c279253910f144","avatarUrl":"/avatars/d91cec60a7b32e872ac8e93c472bf3db.svg","isPro":false,"fullname":"Binqian Xu","user":"BinqianXu","type":"user"},{"_id":"670320fa1b322aa32e18b5ef","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/670320fa1b322aa32e18b5ef/bBkhdADWsP_p4TMNwzksr.png","isPro":false,"fullname":"LingqiKong","user":"Tibbersk","type":"user"},{"_id":"6a4f08d8567298c12d8b6f6c","avatarUrl":"/avatars/0707ff4ee2e65d00ce3556d257ad518d.svg","isPro":false,"fullname":"z","user":"hugsvip","type":"user"},{"_id":"6a4f0bf311bcc94e4ee90e50","avatarUrl":"/avatars/d237fccf29556cb1dc5f2502be71981a.svg","isPro":false,"fullname":"lijing","user":"Lijing419","type":"user"},{"_id":"68cbbb8be421635b1fcd939a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/j0YRMV76frv5qplcJDx_c.png","isPro":false,"fullname":"boli","user":"cnjdml","type":"user"},{"_id":"653926288838e131acdeaee6","avatarUrl":"/avatars/ea72d59d913287e67e371940cc82ab0c.svg","isPro":false,"fullname":"WWZ","user":"WWsirius","type":"user"},{"_id":"69b7afa1549161c6c7a866a0","avatarUrl":"/avatars/59018927d81d3b6163700245d9b480de.svg","isPro":false,"fullname":"nykolcharilus","user":"nykol","type":"user"},{"_id":"680f0a912b588ca79dce7754","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/680f0a912b588ca79dce7754/Ipvi8ZEpCusH3gqLbMOlt.png","isPro":false,"fullname":"GeorgeHu","user":"GeorgeHu6","type":"user"},{"_id":"69bcee16bf15b7207894f9bc","avatarUrl":"/avatars/c33f98251b2ab27252254f4e979846dd.svg","isPro":false,"fullname":"程佳伟","user":"sakuaus","type":"user"},{"_id":"6a4f1b5211b93b862c1ec4e0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a4f1b5211b93b862c1ec4e0/c3g43qxXNcUTegR_bHUEQ.jpeg","isPro":false,"fullname":"aniya","user":"yang210","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.07608.md","query":{}}">
Papers
arxiv:2607.07608

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Published on Jul 8
· Submitted by
neil yu
on Jul 9
#2 Paper of the day
Authors:
,

Abstract

LaMem-VLA introduces a latent-memory-native framework that integrates historical experience into vision-language-action reasoning through coordinated memory components operating in the same latent space.

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.

Community

Paper submitter about 8 hours ago

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.07608
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.07608 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.07608 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers