This survey reframes modern robot learning around <strong>what actually ships—frozen policy weights versus executable skills—and introduces a five-rung taxonomy of code-as-policy systems based on increasingly powerful combinations of execution feedback, persistent memory, and program search, culminating in autonomous self-improving robot skill loops.</strong> </p>\n<p>➡️ 𝐊𝐞𝐲 𝐇𝐢𝐠𝐡𝐥𝐢𝐠𝐡𝐭𝐬 𝐨𝐟 𝐭𝐡𝐞 𝐖𝐞𝐢𝐠𝐡𝐭𝐬-𝐯𝐬-𝐒𝐤𝐢𝐥𝐥𝐬 𝐅𝐫𝐚𝐦𝐞𝐰𝐨𝐫𝐤:</p>\n<p>🧭 𝑾𝒆𝒊𝒈𝒉𝒕𝒔 𝒗𝒔. 𝑺𝒌𝒊𝒍𝒍𝒔 𝑻𝒂𝒙𝒐𝒏𝒐𝒎𝒚: Introduces a unified taxonomy spanning <strong>77 core systems across six robot-learning families</strong>—code-as-policy, end-to-end VLA, reward synthesis, skill libraries, sim-to-real/transfer, and benchmarks—plus 225 landscape works. Its key architectural distinction is whether competence is encoded in <strong>frozen neural weights</strong> (e.g., VLA backbone + action head) or represented as <strong>inspectable, executable programs/skills</strong> that can be edited and recombined after deployment. The taxonomy on page 3 makes this decomposition explicit. </p>\n<p>🔄 𝑭𝒊𝒗𝒆-𝑹𝒖𝒏𝒈 𝑺𝒆𝒍𝒇-𝑰𝒎𝒑𝒓𝒐𝒗𝒆𝒎𝒆𝒏𝒕 𝑳𝒂𝒅𝒅𝒆𝒓 (𝑭 + 𝑴 + 𝑺): The paper's main analytical novelty is decomposing code-as-policy agents by three operational mechanisms—<strong>Feedback (F)</strong> from execution, <strong>Memory (M)</strong> persisted across tasks, and <strong>Search (S)</strong> over multiple candidate programs—and arranging systems from <strong>zero-shot synthesis → closed-loop repair → skill-library accumulation → evolutionary search → full F+M+S self-improvement</strong>. Crucially, it distinguishes sequential debugging from genuine search and frozen model parameters from runtime memory, making “self-improvement” technically testable rather than a loose label. </p>\n<p>🧠 𝑭𝒖𝒍𝒍 𝑺𝒆𝒍𝒇-𝑰𝒎𝒑𝒓𝒐𝒗𝒊𝒏𝒈 𝑹𝒐𝒃𝒐𝒕 𝑳𝒐𝒐𝒑 + 𝑺𝒌𝒊𝒍𝒍 𝑬𝒄𝒐𝒏𝒐𝒎𝒚: Identifies a sparsely populated frontier—represented by <strong>ASPIRE, ENPIRE, and RoboClaw</strong>—where an agent executes skills, obtains grounded traces/feedback, stores validated skills in persistent memory, and searches/mutates candidate programs, feeding accumulated competence into future tasks. The architecture is summarized in the <strong>page-11 diagram as Actor Agent → Execution Engine (F) → Skill Memory (M) → Evolutionary Search (S)</strong> with skills recursively returned to future tasks. The survey argues this is the missing adaptation layer between today's static robot-skill marketplaces and genuinely deployable, continually improving robot ecosystems.</p>\n","updatedAt":"2026-08-07T21:49:16.372Z","author":{"_id":"63a4754927f1f64ed7238dac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63a4754927f1f64ed7238dac/aH-eJF-31g4vof9jv2gmI.jpeg","fullname":"Aman Chadha","name":"amanchadha","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8180489540100098},"editors":["amanchadha"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/63a4754927f1f64ed7238dac/aH-eJF-31g4vof9jv2gmI.jpeg"],"reactions":[],"isReport":false}},{"id":"6a768792be02c2153fb53932","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-08T01:34:10.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [ASPIRE: Agentic /Skills Discovery for Robotics](https://huggingface.co/papers/2607.00272) (2026)\n* [A Few Words Go a Long Way: Language Guided Robot Policy Synthesis](https://huggingface.co/papers/2607.23784) (2026)\n* [RHO: Your Coding Agent is Secretly a Roboticist](https://huggingface.co/papers/2606.16458) (2026)\n* [Sequential Planning via Anchored Robotic Keypoints](https://huggingface.co/papers/2606.30613) (2026)\n* [SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation](https://huggingface.co/papers/2607.08354) (2026)\n* [LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories](https://huggingface.co/papers/2606.13578) (2026)\n* [ETA: A New Agentic Paradigm for Embodied Tasks](https://huggingface.co/papers/2608.03924) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.00272\">ASPIRE: Agentic /Skills Discovery for Robotics</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.23784\">A Few Words Go a Long Way: Language Guided Robot Policy Synthesis</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.16458\">RHO: Your Coding Agent is Secretly a Roboticist</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.30613\">Sequential Planning via Anchored Robotic Keypoints</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.08354\">SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.13578\">LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.03924\">ETA: A New Agentic Paradigm for Embodied Tasks</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-08T01:34:10.072Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7353813648223877},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.01851","authors":[{"_id":"6a76484a8e9301703eaa58ad","name":"Gaytri Jena","hidden":false},{"_id":"6a76484a8e9301703eaa58ae","name":"Kapil Wanaskar","hidden":false},{"_id":"6a76484a8e9301703eaa58af","name":"Vinija Jain","hidden":false},{"_id":"6a76484a8e9301703eaa58b0","user":{"_id":"63a4754927f1f64ed7238dac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63a4754927f1f64ed7238dac/aH-eJF-31g4vof9jv2gmI.jpeg","isPro":false,"fullname":"Aman Chadha","user":"amanchadha","type":"user","name":"amanchadha"},"name":"Aman Chadha","status":"claimed_verified","statusLastChangedAt":"2026-08-08T00:45:04.562Z","hidden":false},{"_id":"6a76484a8e9301703eaa58b1","name":"Vasu Sharma","hidden":false},{"_id":"6a76484a8e9301703eaa58b2","name":"Amitava Das","hidden":false}],"publishedAt":"2026-08-03T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills","submittedOnDailyBy":{"_id":"63a4754927f1f64ed7238dac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63a4754927f1f64ed7238dac/aH-eJF-31g4vof9jv2gmI.jpeg","isPro":false,"fullname":"Aman Chadha","user":"amanchadha","type":"user","name":"amanchadha"},"summary":"Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary \"skills\" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word \"skill\" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.","upvotes":5,"discussionId":"6a76484a8e9301703eaa58b3"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a742a683d438d7c49c8eaa2","avatarUrl":"/avatars/f7828c23fad2988df910b983e5f466b2.svg","isPro":false,"fullname":"Benjamin Nathaniel Hathorne","user":"Entmarch","type":"user"},{"_id":"6a6a829fa698558ca76157ee","avatarUrl":"/avatars/98e9a42de130397e7f4efce0039cdacd.svg","isPro":false,"fullname":"Richard Williams","user":"richardwilliams","type":"user"},{"_id":"6a6c94b74804b5ea70e1608e","avatarUrl":"/avatars/7976fe663fc956d24044fee1356f0049.svg","isPro":false,"fullname":"Joseph Thomas","user":"Azure-Joseph","type":"user"},{"_id":"6a6da9316d51d01100d3b255","avatarUrl":"/avatars/64c0fad6c030e74ee6203bf7d2d01776.svg","isPro":false,"fullname":"Jennifer Lopez","user":"PrismJennifer","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
Abstract
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.
Community
This survey reframes modern robot learning around what actually ships—frozen policy weights versus executable skills—and introduces a five-rung taxonomy of code-as-policy systems based on increasingly powerful combinations of execution feedback, persistent memory, and program search, culminating in autonomous self-improving robot skill loops.
➡️ 𝐊𝐞𝐲 𝐇𝐢𝐠𝐡𝐥𝐢𝐠𝐡𝐭𝐬 𝐨𝐟 𝐭𝐡𝐞 𝐖𝐞𝐢𝐠𝐡𝐭𝐬-𝐯𝐬-𝐒𝐤𝐢𝐥𝐥𝐬 𝐅𝐫𝐚𝐦𝐞𝐰𝐨𝐫𝐤:
🧭 𝑾𝒆𝒊𝒈𝒉𝒕𝒔 𝒗𝒔. 𝑺𝒌𝒊𝒍𝒍𝒔 𝑻𝒂𝒙𝒐𝒏𝒐𝒎𝒚: Introduces a unified taxonomy spanning 77 core systems across six robot-learning families—code-as-policy, end-to-end VLA, reward synthesis, skill libraries, sim-to-real/transfer, and benchmarks—plus 225 landscape works. Its key architectural distinction is whether competence is encoded in frozen neural weights (e.g., VLA backbone + action head) or represented as inspectable, executable programs/skills that can be edited and recombined after deployment. The taxonomy on page 3 makes this decomposition explicit.
🔄 𝑭𝒊𝒗𝒆-𝑹𝒖𝒏𝒈 𝑺𝒆𝒍𝒇-𝑰𝒎𝒑𝒓𝒐𝒗𝒆𝒎𝒆𝒏𝒕 𝑳𝒂𝒅𝒅𝒆𝒓 (𝑭 + 𝑴 + 𝑺): The paper's main analytical novelty is decomposing code-as-policy agents by three operational mechanisms—Feedback (F) from execution, Memory (M) persisted across tasks, and Search (S) over multiple candidate programs—and arranging systems from zero-shot synthesis → closed-loop repair → skill-library accumulation → evolutionary search → full F+M+S self-improvement. Crucially, it distinguishes sequential debugging from genuine search and frozen model parameters from runtime memory, making “self-improvement” technically testable rather than a loose label.
🧠 𝑭𝒖𝒍𝒍 𝑺𝒆𝒍𝒇-𝑰𝒎𝒑𝒓𝒐𝒗𝒊𝒏𝒈 𝑹𝒐𝒃𝒐𝒕 𝑳𝒐𝒐𝒑 + 𝑺𝒌𝒊𝒍𝒍 𝑬𝒄𝒐𝒏𝒐𝒎𝒚: Identifies a sparsely populated frontier—represented by ASPIRE, ENPIRE, and RoboClaw—where an agent executes skills, obtains grounded traces/feedback, stores validated skills in persistent memory, and searches/mutates candidate programs, feeding accumulated competence into future tasks. The architecture is summarized in the page-11 diagram as Actor Agent → Execution Engine (F) → Skill Memory (M) → Evolutionary Search (S) with skills recursively returned to future tasks. The survey argues this is the missing adaptation layer between today's static robot-skill marketplaces and genuinely deployable, continually improving robot ecosystems.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.01851 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.01851 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.01851 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.