TL;DR: An agent is a model <em>and</em> a harness — the code that manages tools,<br>context, and control flow. Optimizing either alone leaves the system<br>bottlenecked by its frozen counterpart. <strong>WHALE</strong> simply alternates: update<br>weights under the current harness, then search for a better harness under<br>the updated model.</p>\n<p>Key findings:</p>\n<ul>\n<li>Beats weight-only, harness-only, and prompt+weight joint adaptation<br>(Fast-Slow Training) by 4.15–24.38 points on SearchQA, Math, and Chess<br>Puzzles with Qwen3.5-2B/4B.</li>\n<li>Bottlenecks are domain-dependent: harness search matches peak weight-only<br>accuracy in SearchQA with ~6% of the rollouts, but is nearly useless in<br>Math until a small weight update unlocks it.</li>\n<li>Small alternating steps beat one-large-pass stagewise optimization —<br>surpassing its final accuracy with only 29%/49% of the rollouts.</li>\n<li>Adaptive WHALE replaces the fixed schedule with a per-phase patience rule<br>on training signals, removing the schedule hyperparameters entirely while<br>matching or beating the best hand-tuned fixed schedule.</li>\n</ul>\n","updatedAt":"2026-09-03T20:38:16.147Z","author":{"_id":"668ff6333bbfdee5f4f14a8a","avatarUrl":"/avatars/83037cfebaf75338296b23bff34c3b19.svg","fullname":"Haechan Kim","name":"HaeChan0305","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8801780343055725},"editors":["HaeChan0305"],"editorAvatarUrls":["/avatars/83037cfebaf75338296b23bff34c3b19.svg"],"reactions":[],"isReport":false}},{"id":"6a9a1dd2476e91ccc190178c","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false},"createdAt":"2026-09-04T01:24:34.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [MemoHarness: Agent Harnesses That Learn from Experience](https://huggingface.co/papers/2607.14159) (2026)\n* [Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories](https://huggingface.co/papers/2608.02276) (2026)\n* [Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents](https://huggingface.co/papers/2607.22688) (2026)\n* [Rethinking the Evaluation of Harness Evolution for Agents](https://huggingface.co/papers/2607.12227) (2026)\n* [HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution](https://huggingface.co/papers/2609.00829) (2026)\n* [JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution](https://huggingface.co/papers/2608.25593) (2026)\n* [Self-Evolving Embodied Agents via Skill-Harness Evolution](https://huggingface.co/papers/2608.11350) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.14159\">MemoHarness: Agent Harnesses That Learn from Experience</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.02276\">Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.22688\">Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.12227\">Rethinking the Evaluation of Harness Evolution for Agents</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2609.00829\">HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.25593\">JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.11350\">Self-Evolving Embodied Agents via Skill-Harness Evolution</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-09-04T01:24:34.845Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7661441564559937},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00196","authors":[{"_id":"6a97bb42fe3c2f89286c3ada","name":"Haechan Kim","hidden":false},{"_id":"6a97bb42fe3c2f89286c3adb","user":{"_id":"64c07361aa57599de1b5e20e","avatarUrl":"/avatars/f67bc6783c14bec55a9c494fa7b89eb2.svg","isPro":true,"fullname":"Yoonho Lee","user":"yoonholee","type":"user","name":"yoonholee"},"name":"Yoonho Lee","status":"claimed_verified","statusLastChangedAt":"2026-09-04T00:45:04.444Z","hidden":false},{"_id":"6a97bb42fe3c2f89286c3adc","name":"Gisang Lee","hidden":false},{"_id":"6a97bb42fe3c2f89286c3add","name":"Chelsea Finn","hidden":false},{"_id":"6a97bb42fe3c2f89286c3ade","name":"Kangwook Lee","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/668ff6333bbfdee5f4f14a8a/g6zHlMA2WJ9MKkGBkVWO9.gif"],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-03T00:00:00.000Z","title":"WHALE: A Simple Recipe for Joint Harness-Weight Optimization","submittedOnDailyBy":{"_id":"668ff6333bbfdee5f4f14a8a","avatarUrl":"/avatars/83037cfebaf75338296b23bff34c3b19.svg","isPro":true,"fullname":"Haechan Kim","user":"HaeChan0305","type":"user","name":"HaeChan0305"},"summary":"Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can change which harness is effective, while harness updates can change which model capabilities are exposed. Existing joint-adaptation methods optimize weights and textual prompts but leave the broader harness fixed. We propose Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model. We instantiate these two phases with online rejection-sampling fine-tuning and Meta-Harness, respectively. When to switch is a key design choice: to separate real improvements from noise without over-optimizing against a changing counterpart, WHALE uses either fixed phase durations or an adaptive patience rule over training signals. Using Qwen3.5-2B/4B agents across three domains (search question answering, mathematical reasoning, and chess puzzles), WHALE outperforms weight-only, harness-only, and Fast-Slow Training by 4.15-24.38 percentage points in best mean@8 accuracy. Either component can be the bottleneck: harness search matches peak weight-only accuracy with far fewer rollouts in SearchQA, but improves math accuracy only after a weight update. Small interleaved updates also outperform stagewise weight-then-harness optimization in accuracy and rollout cost. The code is available at https://github.com/krafton-ai/WHALE.","upvotes":20,"discussionId":"6a97bb42fe3c2f89286c3adf","projectPage":"https://krafton-ai.github.io/blog/whale/","githubRepo":"https://github.com/krafton-ai/WHALE","githubRepoAddedBy":"user","ai_summary":"WHALE alternates model weight updates and harness search to jointly optimize agent performance across reasoning tasks.","ai_keywords":["Weight-Harness Alternating LEarning","online rejection-sampling fine-tuning","Meta-Harness","adaptive patience rule","mean@8 accuracy","Fast-Slow Training"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":15,"organization":{"_id":"6448c201cf9cf3ef36e4f63b","name":"KRAFTON","fullname":"KRAFTON","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6135bc9fa35cb05987acc322/lo1F9RqWgGUSz9n_3CItO.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64c07361aa57599de1b5e20e","avatarUrl":"/avatars/f67bc6783c14bec55a9c494fa7b89eb2.svg","isPro":true,"fullname":"Yoonho Lee","user":"yoonholee","type":"user"},{"_id":"666d506fc0f3d5afc24dd5ca","avatarUrl":"/avatars/eeb98947415d08a26815fd139c76a071.svg","isPro":false,"fullname":"Hyun Ryu","user":"hyun1905","type":"user"},{"_id":"62a4d58e81a4b10e93064ad6","avatarUrl":"/avatars/744d5cbc1745a26b816a458260aba050.svg","isPro":false,"fullname":"hangyulyoon","user":"hangyulmd","type":"user"},{"_id":"67f778ddbb19958f5d96c2a8","avatarUrl":"/avatars/49a3f119b456ff94f28f09b2fe78bb18.svg","isPro":false,"fullname":"Heecheol Yun","user":"yoon6503","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"668ff6333bbfdee5f4f14a8a","avatarUrl":"/avatars/83037cfebaf75338296b23bff34c3b19.svg","isPro":true,"fullname":"Haechan Kim","user":"HaeChan0305","type":"user"},{"_id":"6482959ed9e496b09e06326e","avatarUrl":"/avatars/b49efbcf3ee306bde8fd9f066d3ed16e.svg","isPro":false,"fullname":"Gisang Lee","user":"gisang-lee","type":"user"},{"_id":"61b15ce1a5dd7dc7024406dc","avatarUrl":"/avatars/682ce5ee7d2fec7180dc8e1144cd12ab.svg","isPro":false,"fullname":"Yoonjeon Kim","user":"yjyjyj98","type":"user"},{"_id":"661cc2a9fbadbe6a9d18c8ff","avatarUrl":"/avatars/340e5eea14cff33e080a5f0e11e702dd.svg","isPro":false,"fullname":"Uigyu Kim","user":"Uigyu","type":"user"},{"_id":"657a64c91ccc3c2a5ea5cde4","avatarUrl":"/avatars/aee59f3650cc71424591340ec9f862b2.svg","isPro":false,"fullname":"Hangyeol Jung","user":"Hangyeol","type":"user"},{"_id":"66021a80561089b1a7ebfa01","avatarUrl":"/avatars/a61cf6f95fa0d6750297a00f897436dc.svg","isPro":false,"fullname":"JUNHYEOK CHOI","user":"junhyeokchoi","type":"user"},{"_id":"66303ce3e1c93377db71efd5","avatarUrl":"/avatars/3d444dfb9799c9324c98cba893f4a10f.svg","isPro":false,"fullname":"Yoon Sik Park","user":"nooynoos","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6448c201cf9cf3ef36e4f63b","name":"KRAFTON","fullname":"KRAFTON","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6135bc9fa35cb05987acc322/lo1F9RqWgGUSz9n_3CItO.png"},"query":{}}">
WHALE: A Simple Recipe for Joint Harness-Weight Optimization
Abstract
WHALE alternates model weight updates and harness search to jointly optimize agent performance across reasoning tasks.
Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can change which harness is effective, while harness updates can change which model capabilities are exposed. Existing joint-adaptation methods optimize weights and textual prompts but leave the broader harness fixed. We propose Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model. We instantiate these two phases with online rejection-sampling fine-tuning and Meta-Harness, respectively. When to switch is a key design choice: to separate real improvements from noise without over-optimizing against a changing counterpart, WHALE uses either fixed phase durations or an adaptive patience rule over training signals. Using Qwen3.5-2B/4B agents across three domains (search question answering, mathematical reasoning, and chess puzzles), WHALE outperforms weight-only, harness-only, and Fast-Slow Training by 4.15-24.38 percentage points in best mean@8 accuracy. Either component can be the bottleneck: harness search matches peak weight-only accuracy with far fewer rollouts in SearchQA, but improves math accuracy only after a weight update. Small interleaved updates also outperform stagewise weight-then-harness optimization in accuracy and rollout cost. The code is available at https://github.com/krafton-ai/WHALE.
Community
TL;DR: An agent is a model and a harness — the code that manages tools,
context, and control flow. Optimizing either alone leaves the system
bottlenecked by its frozen counterpart. WHALE simply alternates: update
weights under the current harness, then search for a better harness under
the updated model.
Key findings:
- Beats weight-only, harness-only, and prompt+weight joint adaptation
(Fast-Slow Training) by 4.15–24.38 points on SearchQA, Math, and Chess
Puzzles with Qwen3.5-2B/4B.
- Bottlenecks are domain-dependent: harness search matches peak weight-only
accuracy in SearchQA with ~6% of the rollouts, but is nearly useless in
Math until a small weight update unlocks it.
- Small alternating steps beat one-large-pass stagewise optimization —
surpassing its final accuracy with only 29%/49% of the rollouts.
- Adaptive WHALE replaces the fixed schedule with a per-phase patience rule
on training signals, removing the schedule hyperparameters entirely while
matching or beating the best hand-tuned fixed schedule.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.00196 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.00196 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.00196 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.