Hugging Face Daily Papers · · 5 min read

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.</p>\n","updatedAt":"2026-08-13T04:08:06.669Z","author":{"_id":"691c08008411a45dc9ff4530","avatarUrl":"/avatars/c39c2c9818e96c00825bf12c9dbc912d.svg","fullname":"ltguo","name":"CASIA-IVAer","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8640090227127075},"editors":["CASIA-IVAer"],"editorAvatarUrls":["/avatars/c39c2c9818e96c00825bf12c9dbc912d.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06729","authors":[{"_id":"6a7d430b0ac8bee77474eef4","name":"Guiyu Zhao","hidden":false},{"_id":"6a7d430b0ac8bee77474eef5","name":"Longteng Guo","hidden":false},{"_id":"6a7d430b0ac8bee77474eef6","name":"Yanghong Mei","hidden":false},{"_id":"6a7d430b0ac8bee77474eef7","name":"Zilin Zhu","hidden":false},{"_id":"6a7d430b0ac8bee77474eef8","name":"Yu Zhang","hidden":false},{"_id":"6a7d430b0ac8bee77474eef9","name":"Bin Cao","hidden":false},{"_id":"6a7d430b0ac8bee77474eefa","name":"Mingming Yu","hidden":false},{"_id":"6a7d430b0ac8bee77474eefb","name":"Xingjian He","hidden":false},{"_id":"6a7d430b0ac8bee77474eefc","name":"Jie Jiang","hidden":false},{"_id":"6a7d430b0ac8bee77474eefd","name":"Jing Liu","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models","submittedOnDailyBy":{"_id":"691c08008411a45dc9ff4530","avatarUrl":"/avatars/c39c2c9818e96c00825bf12c9dbc912d.svg","isPro":false,"fullname":"ltguo","user":"CASIA-IVAer","type":"user","name":"CASIA-IVAer"},"summary":"While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.","upvotes":1,"discussionId":"6a7d430b0ac8bee77474eefe","ai_summary":"AtlasVLA improves embodied AI by replacing reactive control with proactive reasoning via persistent world-ego memory, enabling robust long-horizon manipulation from a single wrist camera.","ai_keywords":["Vision-Language-Action models","4D Persistent World State Memory","voxel-hashed spatial state","Ego-Working State Memory","diffusion transformer","DiT","world-ego state"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"691c31f046a68e3d1d8dcfac","name":"CASIA-IVA-Lab","fullname":"CASIA-IVA-Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/691c08008411a45dc9ff4530/Tf9QH8MysDlD_FgaJAf_t.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"691c08008411a45dc9ff4530","avatarUrl":"/avatars/c39c2c9818e96c00825bf12c9dbc912d.svg","isPro":false,"fullname":"ltguo","user":"CASIA-IVAer","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"691c31f046a68e3d1d8dcfac","name":"CASIA-IVA-Lab","fullname":"CASIA-IVA-Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/691c08008411a45dc9ff4530/Tf9QH8MysDlD_FgaJAf_t.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06729.md","query":{}}">
Papers
arxiv:2608.06729

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

Published on Aug 7
· Submitted by
ltguo
on Aug 13
Authors:
,

Abstract

AtlasVLA improves embodied AI by replacing reactive control with proactive reasoning via persistent world-ego memory, enabling robust long-horizon manipulation from a single wrist camera.

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.

Community

Paper submitter about 8 hours ago

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.06729
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.06729 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.06729 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.06729 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers