Hugging Face Daily Papers · · 4 min read

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65377c30e48353201e6fdda0/2XjeKWrewB9yse0gi-N77.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65377c30e48353201e6fdda0/2XjeKWrewB9yse0gi-N77.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-07-14T03:24:14.213Z","author":{"_id":"65377c30e48353201e6fdda0","avatarUrl":"/avatars/a8f803b6f2e598eaee9c52c0d2ddfc16.svg","fullname":"Jiaheng Liu","name":"CheeryLJH","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":29,"isUserFollowing":false}},"numEdits":0,"editors":["CheeryLJH"],"editorAvatarUrls":["/avatars/a8f803b6f2e598eaee9c52c0d2ddfc16.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.10350","authors":[{"_id":"6a55aaf7a9d74d6e65bbd28f","name":"Jiayi Tian","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd290","name":"Shiao Liu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd291","name":"Yuting Xu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd292","name":"Jia Lu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd293","name":"Zihao Guan","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd294","name":"Honglin Han","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd295","name":"Di Yang","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd296","name":"Minqi Gu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd297","name":"Yifei Qian","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd298","name":"Tianlin Zhang","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd299","name":"Yanqing Zhu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd29a","name":"Zeqian Ye","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd29b","name":"Menglin Yang","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd29c","name":"Fei Wang","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd29d","name":"Xu Hu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd29e","name":"Xiuxian Li","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd29f","name":"Wei Zhang","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a0","name":"Shihui Su","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a1","name":"Yiyan Ji","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a2","name":"Jingbo Wang","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a3","name":"Ziteng Feng","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a4","name":"Jiaheng Liu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a5","name":"Zhaoxiang Zhang","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a6","name":"Xiaolong Wu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a7","name":"Mingyang Yin","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a8","name":"Zedong Chu","hidden":false},{"_id":"6a55aaf7a9d74d6e65bbd2a9","name":"Mu Xu","hidden":false}],"publishedAt":"2026-07-11T15:24:43.000Z","submittedOnDailyAt":"2026-07-14T00:00:00.000Z","title":"ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory","submittedOnDailyBy":{"_id":"65377c30e48353201e6fdda0","avatarUrl":"/avatars/a8f803b6f2e598eaee9c52c0d2ddfc16.svg","isPro":false,"fullname":"Jiaheng Liu","user":"CheeryLJH","type":"user","name":"CheeryLJH"},"summary":"Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.","upvotes":67,"discussionId":"6a55aaf7a9d74d6e65bbd2aa","organization":{"_id":"641415d08900ef6afa2fcb73","name":"acvlab","fullname":"Alibaba AMAP CV Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6414106ce7d5f817d204e160/dfveRtrRy8Xn7QpG684zl.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65377c30e48353201e6fdda0","avatarUrl":"/avatars/a8f803b6f2e598eaee9c52c0d2ddfc16.svg","isPro":false,"fullname":"Jiaheng Liu","user":"CheeryLJH","type":"user"},{"_id":"66d8539136aa505569fbd0b2","avatarUrl":"/avatars/bd2681946de858a342bab3e9a16dde1a.svg","isPro":false,"fullname":"jyyyyy","user":"jyyyyy67","type":"user"},{"_id":"647b2c7b2d27d3541decf7a1","avatarUrl":"/avatars/003ea8e46fec9869876280679c7bbbcf.svg","isPro":false,"fullname":"Zedong Chu","user":"jellyczd","type":"user"},{"_id":"64a7eaa9ef22f9c793e055a4","avatarUrl":"/avatars/91d71d9ecc49b4caad99a5dbaf3563e7.svg","isPro":false,"fullname":"tian","user":"tianjy05","type":"user"},{"_id":"62653e93e25d2b490db759d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1650802315262-noauth.jpeg","isPro":false,"fullname":"Hengda Xu","user":"DaDaMrX","type":"user"},{"_id":"662b5ee41ce08a97563c3353","avatarUrl":"/avatars/d4dc86b8c79f8f06060aac1136e7a1a0.svg","isPro":false,"fullname":"Shen Jiajun","user":"NOBULON","type":"user"},{"_id":"69325cf26129af2269cfeed6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69325cf26129af2269cfeed6/0MN0Hysgf64iAMAb_X_n1.jpeg","isPro":false,"fullname":"feng","user":"fengshore04","type":"user"},{"_id":"6a290d1112c6c39ec9069be7","avatarUrl":"/avatars/709c93b61a17bb28cf3e859a30531c46.svg","isPro":false,"fullname":"David Wang","user":"youziwang","type":"user"},{"_id":"616d275d7aa28205c1148fa6","avatarUrl":"/avatars/d8fb008fe45e33053b28a5abc027b98b.svg","isPro":false,"fullname":"Dingqiuyu","user":"Dingqiuyu","type":"user"},{"_id":"66a9a55d7cda19fabeedbb89","avatarUrl":"/avatars/8e7acdd3a9c3552fbeff882bf32f245e.svg","isPro":false,"fullname":"lxp","user":"lxpp","type":"user"},{"_id":"622174f943826d6f261f8206","avatarUrl":"/avatars/a0d4646ecc51d117534b9264bb08ae35.svg","isPro":false,"fullname":"Xiuxian Li","user":"xiuxian","type":"user"},{"_id":"677e133ee86d0754dc7ce296","avatarUrl":"/avatars/4a16414db6f2ff30e32068cab31fab24.svg","isPro":false,"fullname":"mingchenlin","user":"mingchenlin2025","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"641415d08900ef6afa2fcb73","name":"acvlab","fullname":"Alibaba AMAP CV Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6414106ce7d5f817d204e160/dfveRtrRy8Xn7QpG684zl.png"},"query":{}}">
Papers
arxiv:2607.10350

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory

Published on Jul 11
· Submitted by
Jiaheng Liu
on Jul 14
#3 Paper of the day
Authors:
,

Abstract

Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification, and cross-embodiment execution. We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring. ABot-AgentOS further introduces Universal Multi-modal Graph Memory, a persistent source-grounded substrate that converts dialogue, visual observations, spatial context, temporal relations, and task traces into typed nodes and edges. A failure-driven self-evolution loop converts diagnosed memory failures into gated runtime evo-assets that are promoted only to later evaluation splits, preventing current-split ground-truth leakage while enabling continual improvement. On an initial EmbodiedWorldBench subset, ABot-AgentOS improves over a single-controller baseline in both task success and goal completion. Across memory benchmarks, ABot-AgentOS Static achieves 87.5 on LoCoMo, 59.9 on OpenEQA EM-EQA, 88.6 on Mem-Gallery, and 76.5 Acc@All on NExT-QA; self-evolution further improves LoCoMo to 88.7, OpenEQA to 60.4, and Mem-Gallery to 89.0. These results suggest that a general Agent OS layer can improve long-horizon embodied execution while providing persistent, auditable memory for continual interaction.

Community

Paper submitter about 17 hours ago

We present ABot-AgentOS, a general robotic Agent Operating System that sits above low-level controllers and provides a deliberative agent layer for scene-conditioned planning, context-isolated skill execution, multi-stage verification, multi-modal memory, and edge-cloud collaboration. To evaluate such systems, we introduce EmbodiedWorldBench, an executable benchmark with 16 indoor, outdoor, and hybrid scenes, four difficulty levels, and over 200 tasks involving navigation, object search, NPC dialogue, dynamic events, and trace-grounded scoring.

image

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.10350 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.10350 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.10350 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers