Hugging Face Daily Papers · · 3 min read

Multi-Turn On-Policy Distillation with Prefix Replay

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

On-policy distillation for agents without the environment: replay teacher prefixes instead of live rollouts.</p>\n","updatedAt":"2026-07-24T12:05:57.248Z","author":{"_id":"62c414354ce7250560a1f67f","avatarUrl":"/avatars/28fd73973d1703c84f4f59644fef8a80.svg","fullname":"Baohao Liao","name":"baohao","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8411811590194702},"editors":["baohao"],"editorAvatarUrls":["/avatars/28fd73973d1703c84f4f59644fef8a80.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.04763","authors":[{"_id":"6a635452501b0a8116753773","name":"Baohao Liao","hidden":false},{"_id":"6a635452501b0a8116753774","name":"Hanze Dong","hidden":false},{"_id":"6a635452501b0a8116753775","name":"Christof Monz","hidden":false},{"_id":"6a635452501b0a8116753776","name":"Xinxing Xu","hidden":false},{"_id":"6a635452501b0a8116753777","name":"Li Dong","hidden":false},{"_id":"6a635452501b0a8116753778","name":"Furu Wei","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/62c414354ce7250560a1f67f/3VEIpYzxuuJ2gJhlIwfzN.png","https://cdn-uploads.huggingface.co/production/uploads/62c414354ce7250560a1f67f/Mag7Omn1xw2rFTKJqQ6cd.png","https://cdn-uploads.huggingface.co/production/uploads/62c414354ce7250560a1f67f/75P3Zs8umXVcrCvjsJ9HT.png"],"publishedAt":"2026-07-16T00:00:00.000Z","submittedOnDailyAt":"2026-07-24T00:00:00.000Z","title":"Multi-Turn On-Policy Distillation with Prefix Replay","submittedOnDailyBy":{"_id":"62c414354ce7250560a1f67f","avatarUrl":"/avatars/28fd73973d1703c84f4f59644fef8a80.svg","isPro":false,"fullname":"Baohao Liao","user":"baohao","type":"user","name":"baohao"},"summary":"We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes: the student acts at selected steps, while the teacher provides dense per-step supervision without executing new environment interactions. We show that multi-turn OPD introduces a prefix trap: making histories more student-on-policy improves relevance to the student, but can query the teacher on histories where its target is unreliable. This creates a two-sided distribution shift between student occupancy and teacher reliability. ReOPD addresses this by treating multi-turn OPD as a reliability-aware prefix distribution design and implements it with a simple step-decaying sampling schedule that emphasizes early, lower-shift prefixes. Across mathematical reasoning with Python and search environments over multiple teacher and student model scales, ReOPD preserves or improves OPD-level accuracy, uses zero tool calls during student training, and is at least 4times faster per rollout than OPD. ReOPD therefore turns expensive agent-environment interaction into a reusable offline resource, enabling scalable distillation across tools, tasks, and environments.","upvotes":7,"discussionId":"6a635453501b0a8116753779","projectPage":"https://baohaoliao.github.io/ReOPD/","githubRepo":"https://github.com/BaohaoLiao/ReOPD","githubRepoAddedBy":"user","githubStars":4,"organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6997ef2f68950cfdb9f81875","avatarUrl":"/avatars/d99a1f9df211f4162b4e177eded49570.svg","isPro":false,"fullname":"Jcdbzh9olj","user":"jcdbzh9olj","type":"user"},{"_id":"62c414354ce7250560a1f67f","avatarUrl":"/avatars/28fd73973d1703c84f4f59644fef8a80.svg","isPro":false,"fullname":"Baohao Liao","user":"baohao","type":"user"},{"_id":"6a5d94251ff6245d782e2c27","avatarUrl":"/avatars/4b654babc31ff99f761a8f5d6a176ee8.svg","isPro":false,"fullname":"TANG Mingzi","user":"amber-sotokn8","type":"user"},{"_id":"69ccaa2c16f1d072b459e85d","avatarUrl":"/avatars/039eaf77f926f511ceaf7fb5b4b55c6c.svg","isPro":false,"fullname":"신다은","user":"julian-lewis1","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68151d0f51add3813f3f7d1b","name":"MicrosoftResearch","fullname":"Microsoft Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6529a4f2f1205983224fa513/PeuVr7jSuJflmDBBGxoDX.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.04763.md","query":{}}">
Papers
arxiv:2607.04763

Multi-Turn On-Policy Distillation with Prefix Replay

Published on Jul 16
· Submitted by
Baohao Liao
on Jul 24
Authors:
,

Abstract

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes: the student acts at selected steps, while the teacher provides dense per-step supervision without executing new environment interactions. We show that multi-turn OPD introduces a prefix trap: making histories more student-on-policy improves relevance to the student, but can query the teacher on histories where its target is unreliable. This creates a two-sided distribution shift between student occupancy and teacher reliability. ReOPD addresses this by treating multi-turn OPD as a reliability-aware prefix distribution design and implements it with a simple step-decaying sampling schedule that emphasizes early, lower-shift prefixes. Across mathematical reasoning with Python and search environments over multiple teacher and student model scales, ReOPD preserves or improves OPD-level accuracy, uses zero tool calls during student training, and is at least 4times faster per rollout than OPD. ReOPD therefore turns expensive agent-environment interaction into a reusable offline resource, enabling scalable distillation across tools, tasks, and environments.

Community

Paper submitter about 8 hours ago

On-policy distillation for agents without the environment: replay teacher prefixes instead of live rollouts.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.04763
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.04763 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.04763 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.04763 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers