Hugging Face Daily Papers · · 6 min read

In-Context Robot Learning with VLM Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/64e5ee4b93d04e3439f4e988/x4BiHbsJTlvqwHkDn-xV3.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/64e5ee4b93d04e3439f4e988/x4BiHbsJTlvqwHkDn-xV3.png\" alt=\"Figure1_page2\"></a><br>Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.</p>\n","updatedAt":"2026-09-17T11:06:39.888Z","author":{"_id":"64e5ee4b93d04e3439f4e988","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e5ee4b93d04e3439f4e988/g-THJvunoSPUyQ2v_fChU.jpeg","fullname":"taoranyi","name":"thewhole","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8720060586929321},"editors":["thewhole"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64e5ee4b93d04e3439f4e988/g-THJvunoSPUyQ2v_fChU.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.19138","authors":[{"_id":"6aabc4a1d880715a3d023eba","name":"Dongzhou Cheng","hidden":false},{"_id":"6aabc4a1d880715a3d023ebb","name":"Taoran Yi","hidden":false},{"_id":"6aabc4a1d880715a3d023ebc","name":"Ye Fang","hidden":false},{"_id":"6aabc4a1d880715a3d023ebd","name":"Xingwu Zhang","hidden":false},{"_id":"6aabc4a1d880715a3d023ebe","name":"Fan Feng","hidden":false},{"_id":"6aabc4a1d880715a3d023ebf","name":"Yixuan Li","hidden":false},{"_id":"6aabc4a1d880715a3d023ec0","name":"Gengxiong Zhuang","hidden":false},{"_id":"6aabc4a1d880715a3d023ec1","name":"Rongze Wang","hidden":false},{"_id":"6aabc4a1d880715a3d023ec2","name":"Shuai Yang","hidden":false},{"_id":"6aabc4a1d880715a3d023ec3","name":"Wei Song","hidden":false},{"_id":"6aabc4a1d880715a3d023ec4","name":"Weizhi Xue","hidden":false},{"_id":"6aabc4a1d880715a3d023ec5","name":"Minyan Wu","hidden":false},{"_id":"6aabc4a1d880715a3d023ec6","name":"Jie Gui","hidden":false},{"_id":"6aabc4a1d880715a3d023ec7","name":"Jiaqi Wang","hidden":false},{"_id":"6aabc4a1d880715a3d023ec8","name":"Tong Wu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64e5ee4b93d04e3439f4e988/qS2gNBZG3aCBfd2DYDnzP.mp4"],"publishedAt":"2026-09-16T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"In-Context Robot Learning with VLM Agents","submittedOnDailyBy":{"_id":"64e5ee4b93d04e3439f4e988","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e5ee4b93d04e3439f4e988/g-THJvunoSPUyQ2v_fChU.jpeg","isPro":false,"fullname":"taoranyi","user":"thewhole","type":"user","name":"thewhole"},"summary":"Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.","upvotes":13,"discussionId":"6aabc4a1d880715a3d023ec9","projectPage":"https://cheng-haha.github.io/GPT-Policy/","githubRepo":"https://github.com/cheng-haha/GPT-Policy","githubRepoAddedBy":"user","githubStars":178},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64e5ee4b93d04e3439f4e988","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64e5ee4b93d04e3439f4e988/g-THJvunoSPUyQ2v_fChU.jpeg","isPro":false,"fullname":"taoranyi","user":"thewhole","type":"user"},{"_id":"68ef94c055d780acc05005d6","avatarUrl":"/avatars/bbb5192e78f40cd9f6fe2ad0d67271e0.svg","isPro":false,"fullname":"SII-CDZ","user":"SII-CDZ","type":"user"},{"_id":"68f0d98585dbea8d4dd263f8","avatarUrl":"/avatars/797d93a337221a7c34373cd8f5af7825.svg","isPro":false,"fullname":"Weizhi Xue","user":"roseblooming","type":"user"},{"_id":"687b095fed4ab7635a2f9332","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/R9AOoTrglRG8ADxrTEh-X.png","isPro":false,"fullname":"ggengx","user":"ggengx","type":"user"},{"_id":"6310929fda0aaad2f70397a1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6310929fda0aaad2f70397a1/YoHeYZezmCOXwL-qsiZOc.png","isPro":false,"fullname":"Tong Wu","user":"wutong16","type":"user"},{"_id":"64b4eec4faa3181a5eab9c46","avatarUrl":"/avatars/bcc9bf5cbf67546ad2b4c9ec8b96ac96.svg","isPro":true,"fullname":"Jiaqi Wang","user":"myownskyW7","type":"user"},{"_id":"6805f6647593cbf4c6ab23e2","avatarUrl":"/avatars/43b3b9fa612e5461b5a728e6de849bdc.svg","isPro":false,"fullname":"Xingwu Zhang","user":"XingwuZhang","type":"user"},{"_id":"695a04e7d53ec182639f31ed","avatarUrl":"/avatars/66d62a0f32dfd98e242347e392aa0e43.svg","isPro":false,"fullname":"ff","user":"ff0726","type":"user"},{"_id":"665eccf5ffd59344a22533a8","avatarUrl":"/avatars/2ae2710753ce34a04937384bc6dddf70.svg","isPro":false,"fullname":"Wei Song (SII)","user":"Songweii","type":"user"},{"_id":"679b2fbf3c0102760f054a57","avatarUrl":"/avatars/a05a9944992bee0ba2bfd56ac2d4a627.svg","isPro":false,"fullname":"Eric Lee","user":"Salmonnn","type":"user"},{"_id":"6454aad9fbe00f9e73bcd569","avatarUrl":"/avatars/657468e4fcc6905fa95add0af0fad876.svg","isPro":false,"fullname":"Yixuan Li","user":"liyixxxuan","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.19138.md","query":{}}">
Papers
arxiv:2609.19138

In-Context Robot Learning with VLM Agents

Published on Sep 16
· Submitted by
taoranyi
on Sep 17
Authors:
,

Abstract

Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.

Community

Paper submitter about 4 hours ago

Figure1_page2
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.19138
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.19138 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.19138 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.19138 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers