Hugging Face Daily Papers · · 5 min read

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, interface-dependent sequences, cannot directly access device capabilities, and make execution boundaries difficult to define. We present PalmClaw, an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device. PalmClaw exposes device capabilities as device tools with explicit arguments, structured results, and clearly defined execution boundaries. This design enables agents to use mobile capabilities directly while keeping each action explicit and controlled. Experiments show an 11.5% relative improvement in task success and a 94.9% reduction in completion time over the strongest baseline, with lower setup burden and traces illustrating how execution boundaries are applied.</p>\n","updatedAt":"2026-07-16T08:54:45.916Z","author":{"_id":"653a111eee5888edef9182cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/653a111eee5888edef9182cf/7jQ08JDk2UEBla91At8QR.jpeg","fullname":"Hongru Cai","name":"HongruCai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9044077396392822},"editors":["HongruCai"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/653a111eee5888edef9182cf/7jQ08JDk2UEBla91At8QR.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.13027","authors":[{"_id":"6a58872769aa0f8878bddc85","name":"Hongru Cai","hidden":false},{"_id":"6a58872769aa0f8878bddc86","name":"Yongqi Li","hidden":false},{"_id":"6a58872769aa0f8878bddc87","name":"Ran Wei","hidden":false},{"_id":"6a58872769aa0f8878bddc88","name":"Wenjie Li","hidden":false}],"publishedAt":"2026-07-14T00:00:00.000Z","submittedOnDailyAt":"2026-07-16T00:00:00.000Z","title":"PalmClaw: A Native On-Device Agent Framework for Mobile Phones","submittedOnDailyBy":{"_id":"653a111eee5888edef9182cf","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/653a111eee5888edef9182cf/7jQ08JDk2UEBla91At8QR.jpeg","isPro":false,"fullname":"Hongru Cai","user":"HongruCai","type":"user","name":"HongruCai"},"summary":"Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, interface-dependent sequences, cannot directly access device capabilities, and make execution boundaries difficult to define. We present PalmClaw, an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device. PalmClaw exposes device capabilities as device tools with explicit arguments, structured results, and clearly defined execution boundaries. This design enables agents to use mobile capabilities directly while keeping each action explicit and controlled. Experiments show an 11.5\\% relative improvement in task success and a 94.9\\% reduction in completion time over the strongest baseline, with lower setup burden and traces illustrating how execution boundaries are applied. Code is available at https://github.com/ModalityDance/PalmClaw.","upvotes":6,"discussionId":"6a58872869aa0f8878bddc89","projectPage":"https://modalitydance.github.io/PalmClaw/","githubRepo":"https://github.com/ModalityDance/PalmClaw","githubRepoAddedBy":"user","githubStars":1120,"organization":{"_id":"69396d0f6ef210a3d45ac4b7","name":"ModalityDance","fullname":"ModalityDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/653a111eee5888edef9182cf/7BPn5_PnfH27PkAaLQnxW.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a1476fb6126d28ecccf43f4","avatarUrl":"/avatars/23b3ae9c0f1967a8ab0dbbbe7adb6d36.svg","isPro":false,"fullname":"조 예은","user":"averydavis","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"69cd3869d15cd2055b9cabf3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/K9lwYohX1BOCnJ0du7PRF.png","isPro":false,"fullname":"Olivia PEREZ","user":"matthewlewis196","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"619f9755da83161f25840698","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/619f9755da83161f25840698/FM421pE1mz5v1YhrxA8ZA.jpeg","isPro":false,"fullname":"Muhammad Umair","user":"umair894","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"69396d0f6ef210a3d45ac4b7","name":"ModalityDance","fullname":"ModalityDance","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/653a111eee5888edef9182cf/7BPn5_PnfH27PkAaLQnxW.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.13027.md","query":{}}">
Papers
arxiv:2607.13027

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

Published on Jul 14
· Submitted by
Hongru Cai
on Jul 16
Authors:
,

Abstract

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, interface-dependent sequences, cannot directly access device capabilities, and make execution boundaries difficult to define. We present PalmClaw, an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device. PalmClaw exposes device capabilities as device tools with explicit arguments, structured results, and clearly defined execution boundaries. This design enables agents to use mobile capabilities directly while keeping each action explicit and controlled. Experiments show an 11.5\% relative improvement in task success and a 94.9\% reduction in completion time over the strongest baseline, with lower setup burden and traces illustrating how execution boundaries are applied. Code is available at https://github.com/ModalityDance/PalmClaw.

Community

Paper submitter about 11 hours ago

Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, interface-dependent sequences, cannot directly access device capabilities, and make execution boundaries difficult to define. We present PalmClaw, an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device. PalmClaw exposes device capabilities as device tools with explicit arguments, structured results, and clearly defined execution boundaries. This design enables agents to use mobile capabilities directly while keeping each action explicit and controlled. Experiments show an 11.5% relative improvement in task success and a 94.9% reduction in completion time over the strongest baseline, with lower setup burden and traces illustrating how execution boundaries are applied.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.13027
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.13027 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.13027 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.13027 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers