We propose token-level off-policy labeling -- a method that reframes off-policy post-training as token correctness classification.</p>\n","updatedAt":"2026-07-21T05:12:24.399Z","author":{"_id":"63c8454e46421a2efe82709d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c8454e46421a2efe82709d/4k_HR5F6cxkVH1b5KEsGi.webp","fullname":"Deqing Fu","name":"deqing","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":14,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7736445665359497},"editors":["deqing"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/63c8454e46421a2efe82709d/4k_HR5F6cxkVH1b5KEsGi.webp"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.17524","authors":[{"_id":"6a5eff614fe5d1d13e84ac0e","name":"Zitong Huang","hidden":false},{"_id":"6a5eff614fe5d1d13e84ac0f","name":"Gustavo Lucas Carvalho","hidden":false},{"_id":"6a5eff614fe5d1d13e84ac10","name":"Deqing Fu","hidden":false},{"_id":"6a5eff614fe5d1d13e84ac11","name":"Robin Jia","hidden":false}],"publishedAt":"2026-07-20T00:00:00.000Z","submittedOnDailyAt":"2026-07-21T00:00:00.000Z","title":"Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift","submittedOnDailyBy":{"_id":"63c8454e46421a2efe82709d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c8454e46421a2efe82709d/4k_HR5F6cxkVH1b5KEsGi.webp","isPro":true,"fullname":"Deqing Fu","user":"deqing","type":"user","name":"deqing"},"summary":"We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-distribution generalization across 11 datasets against a diverse set of sequence-level and token-level baselines. We further demonstrate that TOPL transfers effectively to machine translation, suggesting that its benefits generalize across different faithful generation tasks. Through ablation studies, we confirm that our token-level learning signal is critical to good performance; sequence-level analogues do not confer similar benefits. Finally, we show that TOPL induces interpretable model updates: the LoRA adapters learned through TOPL function as linear classification heads and steering vectors.","upvotes":1,"discussionId":"6a5eff614fe5d1d13e84ac12","organization":{"_id":"66a403d0dcb5bbc6e98bb7d0","name":"UniversityofSouthernCalifornia","fullname":"University of Southern California","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66a403728069e3c30e0d8524/tkYCfeIJfF1FxtYiRZ8bf.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63c8454e46421a2efe82709d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63c8454e46421a2efe82709d/4k_HR5F6cxkVH1b5KEsGi.webp","isPro":true,"fullname":"Deqing Fu","user":"deqing","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"66a403d0dcb5bbc6e98bb7d0","name":"UniversityofSouthernCalifornia","fullname":"University of Southern California","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66a403728069e3c30e0d8524/tkYCfeIJfF1FxtYiRZ8bf.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.17524.md","query":{}}">
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
Abstract
We propose Token-Level Off-Policy Labeling (TOPL), an off-policy training paradigm that reframes post-training as a token-level correctness prediction task. Our key intuition is that by training the model to distinguish good and bad tokens in a response, we naturally guide the model towards generating good tokens, while avoiding the pitfalls that come with directly training the model to generate off-policy tokens. Experiments on document summarization tasks show that TOPL achieves strong out-of-distribution generalization across 11 datasets against a diverse set of sequence-level and token-level baselines. We further demonstrate that TOPL transfers effectively to machine translation, suggesting that its benefits generalize across different faithful generation tasks. Through ablation studies, we confirm that our token-level learning signal is critical to good performance; sequence-level analogues do not confer similar benefits. Finally, we show that TOPL induces interpretable model updates: the LoRA adapters learned through TOPL function as linear classification heads and steering vectors.
Community
We propose token-level off-policy labeling -- a method that reframes off-policy post-training as token correctness classification.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.17524 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.17524 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.17524 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.