Preference datasets often contain incorrect preference directions or weak/ambiguous pairs. PLC-DPO introduces a latent clean, flip, or tie state for each pair and uses the calibrated policy–reference margin to infer posterior-like routing weights. These determine whether training reinforces, reverses, or suppresses a strong directional update.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65eacbe3769b6bad59ee01c9/yoPRM5UAxhvDvare08At0.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65eacbe3769b6bad59ee01c9/yoPRM5UAxhvDvare08At0.png\" alt=\"overview\"></a></p>\n<p>PLC-DPO combines forward, reverse-direction, and tie-regularizing losses using these weights, with EMA calibration, warm-up, and confidence gating for stable routing. Rather than simply filtering suspicious pairs, it attempts to correct their training signal while reusing the log-probabilities already required by DPO, without an auxiliary model or additional supervision. Across 57 dataset–model–benchmark cells, PLC-DPO achieves the best mean win rate against DPO (60.5%).</p>\n","updatedAt":"2026-09-14T02:07:34.892Z","author":{"_id":"65eacbe3769b6bad59ee01c9","avatarUrl":"/avatars/d1b92aebb97408327764840e5a3fdb6e.svg","fullname":"Boryeong Cho","name":"VennTum","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8187707662582397},"editors":["VennTum"],"editorAvatarUrls":["/avatars/d1b92aebb97408327764840e5a3fdb6e.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.30597","authors":[{"_id":"6a9779d3fe3c2f89286c38a0","user":{"_id":"65eacbe3769b6bad59ee01c9","avatarUrl":"/avatars/d1b92aebb97408327764840e5a3fdb6e.svg","isPro":false,"fullname":"Boryeong Cho","user":"VennTum","type":"user","name":"VennTum"},"name":"Boryeong Cho","status":"claimed_verified","statusLastChangedAt":"2026-09-11T09:54:11.869Z","hidden":false},{"_id":"6a9779d3fe3c2f89286c38a1","name":"Sumyeong Ahn","hidden":false},{"_id":"6a9779d3fe3c2f89286c38a2","name":"Se-Young Yun","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization","submittedOnDailyBy":{"_id":"65eacbe3769b6bad59ee01c9","avatarUrl":"/avatars/d1b92aebb97408327764840e5a3fdb6e.svg","isPro":false,"fullname":"Boryeong Cho","user":"VennTum","type":"user","name":"VennTum"},"summary":"Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.","upvotes":22,"discussionId":"6a9779d4fe3c2f89286c38a3","githubRepo":"https://github.com/VennTum99/PLC-DPO","githubRepoAddedBy":"user","ai_summary":"Posterior Label Correction DPO improves preference optimization by routing noisy pairwise labels into clean, flipped, or tied cases using calibrated policy-reference margins.","ai_keywords":["Direct Preference Optimization","Posterior Label Correction DPO","policy-reference margin","preference learning","noisy preference labels"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67d823d86d6baf84fba04392","avatarUrl":"/avatars/8dfe7800aa3f9d9166963a1f5da5657a.svg","isPro":false,"fullname":"segyu lee","user":"segyulee","type":"user"},{"_id":"65f836339e3737dc3040f3be","avatarUrl":"/avatars/70b479f58338b71192a05331dfa1bb15.svg","isPro":true,"fullname":"Hojung Jung","user":"cossmoss","type":"user"},{"_id":"698fbbb49d17ff43291a932c","avatarUrl":"/avatars/46498083e5406d4485281224858534d3.svg","isPro":false,"fullname":"Sujin Kim","user":"niirhv","type":"user"},{"_id":"677e21e781018fcad9872808","avatarUrl":"/avatars/f8dded5e043b93313c71642a7d48654f.svg","isPro":false,"fullname":"Jaehyun Kwak","user":"Jackwaky","type":"user"},{"_id":"64c2954ee8187de0a7b1e5c9","avatarUrl":"/avatars/7391462c36af2892a401d3d77d056b2c.svg","isPro":false,"fullname":"Soowon Oh","user":"thwannbe","type":"user"},{"_id":"63f5f7754b831cc179ba5410","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677064273670-63f5f7754b831cc179ba5410.jpeg","isPro":false,"fullname":"Reiss Koh","user":"Reiss","type":"user"},{"_id":"63f0c2ac9cf89c9ed1bdd25c","avatarUrl":"/avatars/856b2cb482250fb83c6fe793e29dfd74.svg","isPro":false,"fullname":"Sungnyun Kim","user":"sungnyun","type":"user"},{"_id":"6556bb7f66423b57b2dfe749","avatarUrl":"/avatars/c61267ad7d2192380eb30b93d8681421.svg","isPro":false,"fullname":"JihwanOh","user":"ericoh929","type":"user"},{"_id":"62d0f7faad741b94f5d13744","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677826870505-62d0f7faad741b94f5d13744.jpeg","isPro":false,"fullname":"Namgyu Ho","user":"itsnamgyu","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"64049a20ad54665351d7d8e2","avatarUrl":"/avatars/db72824521826bb9f3f4d849be4d36df.svg","isPro":false,"fullname":"Sangmin Hwang","user":"BBang3","type":"user"},{"_id":"64c9cf490d3d1b209d487308","avatarUrl":"/avatars/5db24ec87db0080198bf233f045fb0d4.svg","isPro":false,"fullname":"SehyeoKKang","user":"SH0329","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.30597.md","query":{}}">
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
Abstract
Posterior Label Correction DPO improves preference optimization by routing noisy pairwise labels into clean, flipped, or tied cases using calibrated policy-reference margins.
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.
Community
Preference datasets often contain incorrect preference directions or weak/ambiguous pairs. PLC-DPO introduces a latent clean, flip, or tie state for each pair and uses the calibrated policy–reference margin to infer posterior-like routing weights. These determine whether training reinforces, reverses, or suppresses a strong directional update.

PLC-DPO combines forward, reverse-direction, and tie-regularizing losses using these weights, with EMA calibration, warm-up, and confidence gating for stable routing. Rather than simply filtering suspicious pairs, it attempts to correct their training signal while reusing the log-probabilities already required by DPO, without an auxiliary model or additional supervision. Across 57 dataset–model–benchmark cells, PLC-DPO achieves the best mean win rate against DPO (60.5%).
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.30597 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.30597 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.30597 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.