Code: <a href=\"https://github.com/liujunzhuo/SMRC-SD\" rel=\"nofollow\">https://github.com/liujunzhuo/SMRC-SD</a></p>\n","updatedAt":"2026-08-10T03:22:30.729Z","author":{"_id":"67eb813842c34766c85a8dc9","avatarUrl":"/avatars/b51a7a90f8564029f0e2873677087416.svg","fullname":"Junzhuo Liu","name":"Junzhuo","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6599797606468201},"editors":["Junzhuo"],"editorAvatarUrls":["/avatars/b51a7a90f8564029f0e2873677087416.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05219","authors":[{"_id":"6a75fcc48e9301703eaa5844","user":{"_id":"67eb813842c34766c85a8dc9","avatarUrl":"/avatars/b51a7a90f8564029f0e2873677087416.svg","isPro":false,"fullname":"Junzhuo Liu","user":"Junzhuo","type":"user","name":"Junzhuo"},"name":"Junzhuo Liu","status":"claimed_verified","statusLastChangedAt":"2026-08-07T16:45:28.224Z","hidden":false},{"_id":"6a75fcc48e9301703eaa5845","name":"Weiwei Li","hidden":false},{"_id":"6a75fcc48e9301703eaa5846","name":"Jun Ling","hidden":false},{"_id":"6a75fcc48e9301703eaa5847","name":"Peng Wang","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents","submittedOnDailyBy":{"_id":"67eb813842c34766c85a8dc9","avatarUrl":"/avatars/b51a7a90f8564029f0e2873677087416.svg","isPro":false,"fullname":"Junzhuo Liu","user":"Junzhuo","type":"user","name":"Junzhuo"},"summary":"Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state--reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student's current execution state. We introduce State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student. At each turn, SMRC-SD verifies whether the student's current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC-SD further constructs state-conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC-SD consistently outperforms unconditional successful full-path distillation. With Qwen3-1.7B, it improves task success from 0.746 to 0.865 on ALFWorld and from 0.574 to 0.693 on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state-compatible teacher context as contributors to these gains. Code is available at https://github.com/liujunzhuo/SMRC-SD.","upvotes":3,"discussionId":"6a75fcc58e9301703eaa5848","githubRepo":"https://github.com/liujunzhuo/SMRC-SD","githubRepoAddedBy":"user","githubStars":1},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"67eb813842c34766c85a8dc9","avatarUrl":"/avatars/b51a7a90f8564029f0e2873677087416.svg","isPro":false,"fullname":"Junzhuo Liu","user":"Junzhuo","type":"user"},{"_id":"6822dff3731ca42bb4c73657","avatarUrl":"/avatars/d95f8fbb8a94f71aa49fbd2bd9b0c602.svg","isPro":false,"fullname":"Dave Lee","user":"davelee-uestc","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05219.md","query":{}}">
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
Abstract
Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state--reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student's current execution state. We introduce State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student. At each turn, SMRC-SD verifies whether the student's current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC-SD further constructs state-conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC-SD consistently outperforms unconditional successful full-path distillation. With Qwen3-1.7B, it improves task success from 0.746 to 0.865 on ALFWorld and from 0.574 to 0.693 on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state-compatible teacher context as contributors to these gains. Code is available at https://github.com/liujunzhuo/SMRC-SD.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.05219 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.05219 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.05219 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.