Agent as Policy (AGP) connects a general purpose coding agent directly to a physical robot. The agent interprets camera observations, writes programs, requests motions, and adjusts its actions using physical feedback while model weights stay fixed. Evaluated tasks include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. AGP achieves at least 8 successful trials out of 10 for each evaluated configuration in assembly, block construction, and dice flipping. Reusing saved procedures and programs reduces execution time across repeated trials.</p>\n","updatedAt":"2026-09-15T04:39:28.723Z","author":{"_id":"642bca844d7a550711e7beac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg","fullname":"JillJia","name":"JillJia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9059349894523621},"editors":["JillJia"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg"],"reactions":[],"isReport":false}},{"id":"6aa8db1932cabad686bf8db2","author":{"_id":"642bca844d7a550711e7beac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg","fullname":"JillJia","name":"JillJia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-09-15T05:43:53.000Z","type":"comment","data":{"edited":true,"hidden":false,"latest":{"raw":"The agent trajectories are available on Hugging Face.\n\n[Agent trajectories](https://huggingface.co/datasets/Agent-as-Policy/agent-as-policy)","html":"<p>The agent trajectories are available on Hugging Face.</p>\n<p><a href=\"https://huggingface.co/datasets/Agent-as-Policy/agent-as-policy\">Agent trajectories</a></p>\n","updatedAt":"2026-09-15T05:45:21.462Z","author":{"_id":"642bca844d7a550711e7beac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg","fullname":"JillJia","name":"JillJia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.7169563174247742},"editors":["JillJia"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg"],"reactions":[],"isReport":false}},{"id":"6aa8e0f15495580f70e63f09","author":{"_id":"642bca844d7a550711e7beac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg","fullname":"JillJia","name":"JillJia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-09-15T06:08:49.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"A short video showing AGP on real robot tasks and how the agent controls the robot through tool calls and visual feedback.\n\n\nhttps://cdn-uploads.huggingface.co/production/uploads/642bca844d7a550711e7beac/61B4lv9vtgEo97l1jkDec.mp4\n","html":"<p>A short video showing AGP on real robot tasks and how the agent controls the robot through tool calls and visual feedback.</p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/642bca844d7a550711e7beac/61B4lv9vtgEo97l1jkDec.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-09-15T06:08:49.072Z","author":{"_id":"642bca844d7a550711e7beac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg","fullname":"JillJia","name":"JillJia","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7135363817214966},"editors":["JillJia"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.12541","authors":[{"_id":"6aa877455dd4cb9b4cc0245d","name":"Mengzhao Jia","hidden":false},{"_id":"6aa877455dd4cb9b4cc0245e","name":"Yang Lin","hidden":false},{"_id":"6aa877455dd4cb9b4cc0245f","name":"Xixin Zhang","hidden":false},{"_id":"6aa877455dd4cb9b4cc02460","name":"Zhihan Zhang","hidden":false},{"_id":"6aa877455dd4cb9b4cc02461","name":"Xiaobai Liu","hidden":false},{"_id":"6aa877455dd4cb9b4cc02462","name":"Meng Jiang","hidden":false}],"publishedAt":"2026-09-11T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"Agent as Policy for Robotic Manipulation","submittedOnDailyBy":{"_id":"642bca844d7a550711e7beac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg","isPro":false,"fullname":"JillJia","user":"JillJia","type":"user","name":"JillJia"},"summary":"We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, die reorientation, targeted throwing, and bimanual towel folding. AGP achieves success rates of 100%, 100%, and 80% on three block construction configurations. These findings establish a path for general-purpose agents to act as robotic policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.","upvotes":5,"discussionId":"6aa877455dd4cb9b4cc02463","projectPage":"https://agent-as-policy-2026.github.io/","ai_summary":"A general-purpose agent directly controls a physical robot by interpreting visuals, writing executable programs, and revising actions based on physical feedback across diverse manipulation tasks.","ai_keywords":["Agent as Policy","executable programs","motion commands","physical outcomes","real-world manipulation","precision manipulation","deformable objects","bimanual towel folding","robotic policies","runtime reasoning"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6356ef35fe4ffe942db2460b","name":"notredame","fullname":"University of Notre Dame","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/RJJ94XCJw7R0WkOyrvXIU.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"642bca844d7a550711e7beac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642bca844d7a550711e7beac/qE0btwUskAYeAmuAe82rh.jpeg","isPro":false,"fullname":"JillJia","user":"JillJia","type":"user"},{"_id":"63bf9695da08ed054400205e","avatarUrl":"/avatars/5223b212431337c6186db96ac70b7ae7.svg","isPro":false,"fullname":"Zhihan Zhang","user":"zhihz0535","type":"user"},{"_id":"64dbec79c8d191a54ae936e1","avatarUrl":"/avatars/3b50117c7bb7ab294952bd3f59615e65.svg","isPro":false,"fullname":"Xixin Zhang","user":"XixinZhang","type":"user"},{"_id":"6a06be7048e108aedff98143","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a06be7048e108aedff98143/KNOjEvrnC6Vyf6tsStSGH.jpeg","isPro":false,"fullname":"YANG LIN","user":"elsannaly","type":"user"},{"_id":"68ecaf459490559e1ecdb59d","avatarUrl":"/avatars/e7c97693960519a9270452da95e718c0.svg","isPro":false,"fullname":"gjb970312.","user":"gjb970312","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6356ef35fe4ffe942db2460b","name":"notredame","fullname":"University of Notre Dame","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/RJJ94XCJw7R0WkOyrvXIU.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.12541.md","query":{}}">
Agent as Policy for Robotic Manipulation
Abstract
A general-purpose agent directly controls a physical robot by interpreting visuals, writing executable programs, and revising actions based on physical feedback across diverse manipulation tasks.
We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, die reorientation, targeted throwing, and bimanual towel folding. AGP achieves success rates of 100%, 100%, and 80% on three block construction configurations. These findings establish a path for general-purpose agents to act as robotic policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.
Community
Agent as Policy (AGP) connects a general purpose coding agent directly to a physical robot. The agent interprets camera observations, writes programs, requests motions, and adjusts its actions using physical feedback while model weights stay fixed. Evaluated tasks include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. AGP achieves at least 8 successful trials out of 10 for each evaluated configuration in assembly, block construction, and dice flipping. Reusing saved procedures and programs reduces execution time across repeated trials.
A short video showing AGP on real robot tasks and how the agent controls the robot through tool calls and visual feedback.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.12541 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.12541 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.