**Which skill is most worth evaluating next?**\n\nInstead of exhaustively trying and refining candidate skills, COBRA-Skills maintains a population of skills and uses:\n\n- a neural predictor to estimate skill utility,\n- **LinearUCB** to balance exploration and exploitation,\n- target-agent execution feedback to update the bandit,\n- scheduled evolution operators (regeneration, rollout mutation, and crossover) to evolve the skill population.\n\n### Results\n\nWe evaluate COBRA-Skills on **6 diverse agent benchmarks × 3 target models**, covering search QA, spreadsheets, document understanding, mathematical reasoning, social reasoning, and embodied tasks.\n\nCOBRA-Skills achieves:\n\n- 🏆 the best average performance among the compared methods,\n- 💰 **~55–58% lower optimization cost** than SkillOpt,\n- 📊 only **50 optimization examples per benchmark**,\n- 🔧 consistent effectiveness under **Codex and Claude Code** harnesses,\n- 🤖 strong performance even when the **target model itself** generates and refines skills.\n\nAn interesting finding is that the cost reduction does not mainly come from reducing target-agent executions. A large part comes from avoiding repeated LLM-based trajectory analysis and skill rewriting.\n\n📄 Paper: https://arxiv.org/abs/2609.11682 \n💻 Code: https://github.com/Jerry-LuP/COBRA-Skills\n\nFeedback and discussions are very welcome!","html":"<h1 class=\"relative group flex items-baseline\">\n\t<a id=\"cobra-skills-contextual-bandits-for-efficient-agent-skill-optimization-🚀\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#cobra-skills-contextual-bandits-for-efficient-agent-skill-optimization-🚀\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tCOBRA-Skills: Contextual Bandits for Efficient Agent Skill Optimization 🚀\n\t</span>\n</h1>\n<p>How can we optimize Agent Skills without repeatedly spending large amounts of computation on weak candidates and costly LLM-based refinement?</p>\n<p><strong>COBRA-Skills</strong> treats skill optimization as a sequential budget-allocation problem:</p>\n<blockquote>\n<p><strong>Which skill is most worth evaluating next?</strong></p>\n</blockquote>\n<p>Instead of exhaustively trying and refining candidate skills, COBRA-Skills maintains a population of skills and uses:</p>\n<ul>\n<li>a neural predictor to estimate skill utility,</li>\n<li><strong>LinearUCB</strong> to balance exploration and exploitation,</li>\n<li>target-agent execution feedback to update the bandit,</li>\n<li>scheduled evolution operators (regeneration, rollout mutation, and crossover) to evolve the skill population.</li>\n</ul>\n<h3 class=\"relative group flex items-baseline\">\n\t<a id=\"results\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#results\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\tResults\n\t</span>\n</h3>\n<p>We evaluate COBRA-Skills on <strong>6 diverse agent benchmarks × 3 target models</strong>, covering search QA, spreadsheets, document understanding, mathematical reasoning, social reasoning, and embodied tasks.</p>\n<p>COBRA-Skills achieves:</p>\n<ul>\n<li>🏆 the best average performance among the compared methods,</li>\n<li>💰 <strong>~55–58% lower optimization cost</strong> than SkillOpt,</li>\n<li>📊 only <strong>50 optimization examples per benchmark</strong>,</li>\n<li>🔧 consistent effectiveness under <strong>Codex and Claude Code</strong> harnesses,</li>\n<li>🤖 strong performance even when the <strong>target model itself</strong> generates and refines skills.</li>\n</ul>\n<p>An interesting finding is that the cost reduction does not mainly come from reducing target-agent executions. A large part comes from avoiding repeated LLM-based trajectory analysis and skill rewriting.</p>\n<p>📄 Paper: <a href=\"https://arxiv.org/abs/2609.11682\" rel=\"nofollow\">https://arxiv.org/abs/2609.11682</a><br>💻 Code: <a href=\"https://github.com/Jerry-LuP/COBRA-Skills\" rel=\"nofollow\">https://github.com/Jerry-LuP/COBRA-Skills</a></p>\n<p>Feedback and discussions are very welcome!</p>\n","updatedAt":"2026-09-14T02:09:05.673Z","author":{"_id":"67208288846173fe57b31427","avatarUrl":"/avatars/920dfb18cf30d98e17d6a7f3843e663f.svg","fullname":"Jerry smith","name":"Lpc206","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8400620222091675},"editors":["Lpc206"],"editorAvatarUrls":["/avatars/920dfb18cf30d98e17d6a7f3843e663f.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.11682","authors":[{"_id":"6aa3669547a406da7901e6bd","user":{"_id":"67208288846173fe57b31427","avatarUrl":"/avatars/920dfb18cf30d98e17d6a7f3843e663f.svg","isPro":false,"fullname":"Jerry smith","user":"Lpc206","type":"user","name":"Lpc206"},"name":"Pingchen Lu","status":"claimed_verified","statusLastChangedAt":"2026-09-11T09:53:31.646Z","hidden":false},{"_id":"6aa3669547a406da7901e6be","name":"Xiangyi Wang","hidden":false},{"_id":"6aa3669547a406da7901e6bf","name":"Xiang Li","hidden":false},{"_id":"6aa3669547a406da7901e6c0","name":"Jie Mao","hidden":false},{"_id":"6aa3669547a406da7901e6c1","name":"Zikun Qu","hidden":false},{"_id":"6aa3669547a406da7901e6c2","name":"Junfeng Luo","hidden":false},{"_id":"6aa3669547a406da7901e6c3","name":"Yao Shu","hidden":false},{"_id":"6aa3669547a406da7901e6c4","name":"Bryan Kian Hsiang Low","hidden":false},{"_id":"6aa3669547a406da7901e6c5","name":"Zhongxiang Dai","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/67208288846173fe57b31427/Ikw82bmRnjsXzLWN3Wgbs.png","https://cdn-uploads.huggingface.co/production/uploads/67208288846173fe57b31427/VzFHMnI6TKRR5QMXOOZmX.png","https://cdn-uploads.huggingface.co/production/uploads/67208288846173fe57b31427/-XtSd8Kz8u-TX2P3tSd9s.png","https://cdn-uploads.huggingface.co/production/uploads/67208288846173fe57b31427/Sk6VCaoRexRBQ4aFYLUtB.png","https://cdn-uploads.huggingface.co/production/uploads/67208288846173fe57b31427/Ay0CxxJVGIPpxjN8EjG-i.png","https://cdn-uploads.huggingface.co/production/uploads/67208288846173fe57b31427/QisCBkx-rfXrKkdSo_otH.png","https://cdn-uploads.huggingface.co/production/uploads/67208288846173fe57b31427/H-LmlKDksmisVH5kDCz6D.png"],"publishedAt":"2026-09-10T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization","submittedOnDailyBy":{"_id":"67208288846173fe57b31427","avatarUrl":"/avatars/920dfb18cf30d98e17d6a7f3843e663f.svg","isPro":false,"fullname":"Jerry smith","user":"Lpc206","type":"user","name":"Lpc206"},"summary":"Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce COBRA-Skills, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space. COBRA-Skills couples contextual-bandit-guided prioritization with evidence-grounded skill evolution, selectively allocating evaluations to promising or informative candidates while continually refining the skill population from execution feedback. Across six heterogeneous agent benchmarks and three target models, COBRA-Skills consistently achieves the strongest average performance among compared methods, while reducing optimization cost by 55--58\\% relative to SkillOpt and using only 50 unique optimization examples per benchmark. Further analyses show that COBRA-Skills remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.","upvotes":26,"discussionId":"6aa3669547a406da7901e6c6","githubRepo":"https://github.com/Jerry-LuP/COBRA-Skills","githubRepoAddedBy":"user","ai_summary":"COBRA-Skills improves LLM agent skill optimization by using contextual-bandit prioritization and evidence-based evolution to cut evaluation costs while maintaining high performance.","ai_keywords":["COBRA-Skills","contextual-bandit-guided prioritization","evidence-grounded skill evolution","budgeted sequential optimization","skill optimization","LLM agents"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":5,"organization":{"_id":"6223644d0129f2097d69a407","name":"CUHKSZ","fullname":"Chinese University of Hong Kong, Shenzhen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1646486592158-6108ae87823007eaf0c7bd1e.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67208288846173fe57b31427","avatarUrl":"/avatars/920dfb18cf30d98e17d6a7f3843e663f.svg","isPro":false,"fullname":"Jerry smith","user":"Lpc206","type":"user"},{"_id":"6a795012189047f82009f3b3","avatarUrl":"/avatars/3367920f104ca45305944f6bc1d6b8df.svg","isPro":false,"fullname":"Zhongxiang Dai","user":"dzxagent","type":"user"},{"_id":"66440e86bfe15e84d369cb03","avatarUrl":"/avatars/d15b3b3831bc74138206071612169f64.svg","isPro":false,"fullname":"Xinyuan Xie","user":"SatsukiVie","type":"user"},{"_id":"6aa75f1d4a6cd3767abed18a","avatarUrl":"/avatars/3a5334b2873e31bae92f7d6e2fc8241e.svg","isPro":false,"fullname":"Kong Zedong","user":"dongdong666","type":"user"},{"_id":"6aa75e97d26a08997320615f","avatarUrl":"/avatars/549248a87cd5fc71ee2b800d71f7f91b.svg","isPro":false,"fullname":"Yan Lu","user":"Salt1222","type":"user"},{"_id":"6aa75f77c77c725f6846090b","avatarUrl":"/avatars/15aa6e38b0a6c58c1ba81ce43c52416a.svg","isPro":false,"fullname":"Cai ruyin","user":"yymtdykx","type":"user"},{"_id":"6aa76030dd4d0e0172461c23","avatarUrl":"/avatars/802485e594fd0a0cc447003a97853822.svg","isPro":false,"fullname":"Xiaoyu Nie","user":"Nxy-207","type":"user"},{"_id":"69fd548a154c7dfe3c9bea2e","avatarUrl":"/avatars/9f3610aa39d4ab9b01d203f80dbd1ee2.svg","isPro":false,"fullname":"L","user":"lzx666YAYAYA","type":"user"},{"_id":"6aa7640e051d3d79e9910de5","avatarUrl":"/avatars/19187167bdc43616cee60ca4fcd16663.svg","isPro":false,"fullname":"王一川","user":"Wangyichuan","type":"user"},{"_id":"6aa765df83fe5c75d66855a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/tcn_NZiwFhAaM-iZ5oHNN.png","isPro":false,"fullname":"vincent","user":"E23Q","type":"user"},{"_id":"67f369756285f17584396bb0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/J6e1rXahcyUFtmI152nMF.png","isPro":false,"fullname":"李喆","user":"learning-ljj","type":"user"},{"_id":"6aa76f782e4c51b0c95784cc","avatarUrl":"/avatars/43856d09e3cff816405ee582462b633e.svg","isPro":false,"fullname":"mrj369","user":"mrj369","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6223644d0129f2097d69a407","name":"CUHKSZ","fullname":"Chinese University of Hong Kong, Shenzhen","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1646486592158-6108ae87823007eaf0c7bd1e.png"},"query":{}}">
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
Abstract
COBRA-Skills improves LLM agent skill optimization by using contextual-bandit prioritization and evidence-based evolution to cut evaluation costs while maintaining high performance.
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data. We introduce COBRA-Skills, an efficient framework that formulates skill optimization as budgeted sequential optimization over a dynamically evolving candidate space. COBRA-Skills couples contextual-bandit-guided prioritization with evidence-grounded skill evolution, selectively allocating evaluations to promising or informative candidates while continually refining the skill population from execution feedback. Across six heterogeneous agent benchmarks and three target models, COBRA-Skills consistently achieves the strongest average performance among compared methods, while reducing optimization cost by 55--58\% relative to SkillOpt and using only 50 unique optimization examples per benchmark. Further analyses show that COBRA-Skills remains robust to changes in the agent harness and performs effectively when the target model itself is used for skill generation and refinement.
Community
COBRA-Skills: Contextual Bandits for Efficient Agent Skill Optimization 🚀
How can we optimize Agent Skills without repeatedly spending large amounts of computation on weak candidates and costly LLM-based refinement?
COBRA-Skills treats skill optimization as a sequential budget-allocation problem:
Which skill is most worth evaluating next?
Instead of exhaustively trying and refining candidate skills, COBRA-Skills maintains a population of skills and uses:
- a neural predictor to estimate skill utility,
- LinearUCB to balance exploration and exploitation,
- target-agent execution feedback to update the bandit,
- scheduled evolution operators (regeneration, rollout mutation, and crossover) to evolve the skill population.
Results
We evaluate COBRA-Skills on 6 diverse agent benchmarks × 3 target models, covering search QA, spreadsheets, document understanding, mathematical reasoning, social reasoning, and embodied tasks.
COBRA-Skills achieves:
- 🏆 the best average performance among the compared methods,
- 💰 ~55–58% lower optimization cost than SkillOpt,
- 📊 only 50 optimization examples per benchmark,
- 🔧 consistent effectiveness under Codex and Claude Code harnesses,
- 🤖 strong performance even when the target model itself generates and refines skills.
An interesting finding is that the cost reduction does not mainly come from reducing target-agent executions. A large part comes from avoiding repeated LLM-based trajectory analysis and skill rewriting.
📄 Paper: https://arxiv.org/abs/2609.11682
💻 Code: https://github.com/Jerry-LuP/COBRA-Skills
Feedback and discussions are very welcome!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.11682 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.11682 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.11682 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.