We present LLM-as-a-Coach which repurposes LLM-as-a-Judge in RL as a experiential knowledge extractor for non-verifiable tasks. LLM-as-a-Coach extracts transferable knowledge given policy response and rubrics, and internalize it with on-policy context distillation into policy model weights.<br>Code will be available at <a href=\"https://aka.ms/el-code\" rel=\"nofollow\">https://aka.ms/el-code</a><br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/64ac0605091ff88865352b44/0yxev0I9ND0Jntz0SI2Pn.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/64ac0605091ff88865352b44/0yxev0I9ND0Jntz0SI2Pn.png\" alt=\"el_method\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/64ac0605091ff88865352b44/f3dgJUz3Sv4S55M2xujXL.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/64ac0605091ff88865352b44/f3dgJUz3Sv4S55M2xujXL.png\" alt=\"el_intro\"></a></p>\n","updatedAt":"2026-07-21T02:52:20.910Z","author":{"_id":"64ac0605091ff88865352b44","avatarUrl":"/avatars/c28acb08a1fdeab899afd1961d3dd94a.svg","fullname":"ytz","name":"ytz20","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":12,"isUserFollowing":false}},"numEdits":2,"identifiedLanguage":{"language":"en","probability":0.9197226762771606},"editors":["ytz20"],"editorAvatarUrls":["/avatars/c28acb08a1fdeab899afd1961d3dd94a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18110","authors":[{"_id":"6a5edcec4fe5d1d13e84ab3f","name":"Tianzhu Ye","hidden":false},{"_id":"6a5edcec4fe5d1d13e84ab40","name":"Li Dong","hidden":false},{"_id":"6a5edcec4fe5d1d13e84ab41","name":"Guanheng Chen","hidden":false},{"_id":"6a5edcec4fe5d1d13e84ab42","name":"He Zhu","hidden":false},{"_id":"6a5edcec4fe5d1d13e84ab43","name":"Xun Wu","hidden":false},{"_id":"6a5edcec4fe5d1d13e84ab44","name":"Shaohan Huang","hidden":false},{"_id":"6a5edcec4fe5d1d13e84ab45","name":"Furu Wei","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/64ac0605091ff88865352b44/Z4G-kISNJHHP1yqLnUN_e.png"],"publishedAt":"2026-07-20T00:00:00.000Z","submittedOnDailyAt":"2026-07-21T00:00:00.000Z","title":"LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks","submittedOnDailyBy":{"_id":"64ac0605091ff88865352b44","avatarUrl":"/avatars/c28acb08a1fdeab899afd1961d3dd94a.svg","isPro":false,"fullname":"ytz","user":"ytz20","type":"user","name":"ytz20"},"summary":"Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.","upvotes":7,"discussionId":"6a5edcec4fe5d1d13e84ab46","organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64ac0605091ff88865352b44","avatarUrl":"/avatars/c28acb08a1fdeab899afd1961d3dd94a.svg","isPro":false,"fullname":"ytz","user":"ytz20","type":"user"},{"_id":"5df85abada6d0311fd3d5408","avatarUrl":"/avatars/2331cf703c1b5d3a62e2050b1a6eb108.svg","isPro":false,"fullname":"Li Dong","user":"unilm","type":"user"},{"_id":"68ef6325918a07690443c922","avatarUrl":"/avatars/6200c248f89afca2bab5b62428f3df0d.svg","isPro":false,"fullname":"Tianzhu Ye","user":"6eo","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"661230e7941ed67394a873e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661230e7941ed67394a873e7/MoQ-FHlTu1In-UpG6E9j2.jpeg","isPro":false,"fullname":"Li Yu","user":"phxember","type":"user"},{"_id":"63672d7864bcbbd03e39b2f1","avatarUrl":"/avatars/0c100709de78988a9a39dd48cb1be513.svg","isPro":false,"fullname":"Zhu He","user":"chichi56","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18110.md","query":{}}">
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
Published on Jul 20
· Submitted by ytz on Jul 21 Abstract
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.
Community
We present LLM-as-a-Coach which repurposes LLM-as-a-Judge in RL as a experiential knowledge extractor for non-verifiable tasks. LLM-as-a-Coach extracts transferable knowledge given policy response and rubrics, and internalize it with on-policy context distillation into policy model weights.
Code will be available at https://aka.ms/el-code


Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.18110 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.18110 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.18110 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.