Hugging Face Daily Papers · · 4 min read

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Preference datasets often contain incorrect preference directions or weak/ambiguous pairs. PLC-DPO introduces a latent clean, flip, or tie state for each pair and uses the calibrated policy–reference margin to infer posterior-like routing weights. These determine whether training reinforces, reverses, or suppresses a strong directional update.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65eacbe3769b6bad59ee01c9/yoPRM5UAxhvDvare08At0.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65eacbe3769b6bad59ee01c9/yoPRM5UAxhvDvare08At0.png\" alt=\"overview\"></a></p>\n<p>PLC-DPO combines forward, reverse-direction, and tie-regularizing losses using these weights, with EMA calibration, warm-up, and confidence gating for stable routing. Rather than simply filtering suspicious pairs, it attempts to correct their training signal while reusing the log-probabilities already required by DPO, without an auxiliary model or additional supervision. Across 57 dataset–model–benchmark cells, PLC-DPO achieves the best mean win rate against DPO (60.5%).</p>\n","updatedAt":"2026-09-14T02:07:34.892Z","author":{"_id":"65eacbe3769b6bad59ee01c9","avatarUrl":"/avatars/d1b92aebb97408327764840e5a3fdb6e.svg","fullname":"Boryeong Cho","name":"VennTum","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8187707662582397},"editors":["VennTum"],"editorAvatarUrls":["/avatars/d1b92aebb97408327764840e5a3fdb6e.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.30597","authors":[{"_id":"6a9779d3fe3c2f89286c38a0","user":{"_id":"65eacbe3769b6bad59ee01c9","avatarUrl":"/avatars/d1b92aebb97408327764840e5a3fdb6e.svg","isPro":false,"fullname":"Boryeong Cho","user":"VennTum","type":"user","name":"VennTum"},"name":"Boryeong Cho","status":"claimed_verified","statusLastChangedAt":"2026-09-11T09:54:11.869Z","hidden":false},{"_id":"6a9779d3fe3c2f89286c38a1","name":"Sumyeong Ahn","hidden":false},{"_id":"6a9779d3fe3c2f89286c38a2","name":"Se-Young Yun","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-14T00:00:00.000Z","title":"PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization","submittedOnDailyBy":{"_id":"65eacbe3769b6bad59ee01c9","avatarUrl":"/avatars/d1b92aebb97408327764840e5a3fdb6e.svg","isPro":false,"fullname":"Boryeong Cho","user":"VennTum","type":"user","name":"VennTum"},"summary":"Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.","upvotes":22,"discussionId":"6a9779d4fe3c2f89286c38a3","githubRepo":"https://github.com/VennTum99/PLC-DPO","githubRepoAddedBy":"user","ai_summary":"Posterior Label Correction DPO improves preference optimization by routing noisy pairwise labels into clean, flipped, or tied cases using calibrated policy-reference margins.","ai_keywords":["Direct Preference Optimization","Posterior Label Correction DPO","policy-reference margin","preference learning","noisy preference labels"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67d823d86d6baf84fba04392","avatarUrl":"/avatars/8dfe7800aa3f9d9166963a1f5da5657a.svg","isPro":false,"fullname":"segyu lee","user":"segyulee","type":"user"},{"_id":"65f836339e3737dc3040f3be","avatarUrl":"/avatars/70b479f58338b71192a05331dfa1bb15.svg","isPro":true,"fullname":"Hojung Jung","user":"cossmoss","type":"user"},{"_id":"698fbbb49d17ff43291a932c","avatarUrl":"/avatars/46498083e5406d4485281224858534d3.svg","isPro":false,"fullname":"Sujin Kim","user":"niirhv","type":"user"},{"_id":"677e21e781018fcad9872808","avatarUrl":"/avatars/f8dded5e043b93313c71642a7d48654f.svg","isPro":false,"fullname":"Jaehyun Kwak","user":"Jackwaky","type":"user"},{"_id":"64c2954ee8187de0a7b1e5c9","avatarUrl":"/avatars/7391462c36af2892a401d3d77d056b2c.svg","isPro":false,"fullname":"Soowon Oh","user":"thwannbe","type":"user"},{"_id":"63f5f7754b831cc179ba5410","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677064273670-63f5f7754b831cc179ba5410.jpeg","isPro":false,"fullname":"Reiss Koh","user":"Reiss","type":"user"},{"_id":"63f0c2ac9cf89c9ed1bdd25c","avatarUrl":"/avatars/856b2cb482250fb83c6fe793e29dfd74.svg","isPro":false,"fullname":"Sungnyun Kim","user":"sungnyun","type":"user"},{"_id":"6556bb7f66423b57b2dfe749","avatarUrl":"/avatars/c61267ad7d2192380eb30b93d8681421.svg","isPro":false,"fullname":"JihwanOh","user":"ericoh929","type":"user"},{"_id":"62d0f7faad741b94f5d13744","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1677826870505-62d0f7faad741b94f5d13744.jpeg","isPro":false,"fullname":"Namgyu Ho","user":"itsnamgyu","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"64049a20ad54665351d7d8e2","avatarUrl":"/avatars/db72824521826bb9f3f4d849be4d36df.svg","isPro":false,"fullname":"Sangmin Hwang","user":"BBang3","type":"user"},{"_id":"64c9cf490d3d1b209d487308","avatarUrl":"/avatars/5db24ec87db0080198bf233f045fb0d4.svg","isPro":false,"fullname":"SehyeoKKang","user":"SH0329","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.30597.md","query":{}}">
Papers
arxiv:2608.30597

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

Published on Aug 31
· Submitted by
Boryeong Cho
on Sep 14
Authors:

Abstract

Posterior Label Correction DPO improves preference optimization by routing noisy pairwise labels into clean, flipped, or tied cases using calibrated policy-reference margins.

Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.

Community

Paper author Paper submitter about 8 hours ago

Preference datasets often contain incorrect preference directions or weak/ambiguous pairs. PLC-DPO introduces a latent clean, flip, or tie state for each pair and uses the calibrated policy–reference margin to infer posterior-like routing weights. These determine whether training reinforces, reverses, or suppresses a strong directional update.

overview

PLC-DPO combines forward, reverse-direction, and tie-regularizing losses using these weights, with EMA calibration, warm-up, and confidence gating for stable routing. Rather than simply filtering suspicious pairs, it attempts to correct their training signal while reusing the log-probabilities already required by DPO, without an auxiliary model or additional supervision. Across 57 dataset–model–benchmark cells, PLC-DPO achieves the best mean win rate against DPO (60.5%).

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.30597
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.30597 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.30597 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.30597 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers