We propose Negative Self-Distillation (NSD), a label-free, fully self-bootstrapped framework that enhances reasoning capabilities by optimizing the model to diverge from self-generated flawed trajectories, eliminating the need for ground-truth solutions or an external teacher.</p>\n","updatedAt":"2026-09-11T11:17:15.223Z","author":{"_id":"67316c6cb9634ac96f65e1a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/uUY9F17MiN7NhGAg01Yom.png","fullname":"PP","name":"PassionPrc","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9067786335945129},"editors":["PassionPrc"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/uUY9F17MiN7NhGAg01Yom.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.11699","authors":[{"_id":"6aa3a26447a406da7901e87d","user":{"_id":"67316c6cb9634ac96f65e1a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/uUY9F17MiN7NhGAg01Yom.png","isPro":false,"fullname":"PP","user":"PassionPrc","type":"user","name":"PassionPrc"},"name":"Rongcan Pei","status":"claimed_verified","statusLastChangedAt":"2026-09-11T09:53:16.674Z","hidden":false},{"_id":"6aa3a26447a406da7901e87e","name":"Zhepei Wei","hidden":false},{"_id":"6aa3a26447a406da7901e87f","name":"Shuyao Xu","hidden":false},{"_id":"6aa3a26447a406da7901e880","name":"Xinyu Zhu","hidden":false},{"_id":"6aa3a26447a406da7901e881","name":"Wei-Lin Chen","hidden":false},{"_id":"6aa3a26447a406da7901e882","name":"Yu Meng","hidden":false}],"publishedAt":"2026-09-10T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"Negative Self-Distillation: Learning to Reason by Avoiding Flaws","submittedOnDailyBy":{"_id":"67316c6cb9634ac96f65e1a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/uUY9F17MiN7NhGAg01Yom.png","isPro":false,"fullname":"PP","user":"PassionPrc","type":"user","name":"PassionPrc"},"summary":"On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.","upvotes":3,"discussionId":"6aa3a26447a406da7901e883","githubRepo":"https://github.com/Prongcan/NSD","githubRepoAddedBy":"user","ai_summary":"Negative Self-Distillation improves large language model reasoning by pushing models away from self-generated flawed reasoning via a dynamic gating mechanism that protects linguistic capabilities.","ai_keywords":["On-Policy Self-Distillation","Negative Self-Distillation","self-distillation","reasoning-critical tokens","dynamic gating mechanism","unlearning","self-bootstrapping reinforcement learning"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":4},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67316c6cb9634ac96f65e1a0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/uUY9F17MiN7NhGAg01Yom.png","isPro":false,"fullname":"PP","user":"PassionPrc","type":"user"},{"_id":"6587e5a4b2177de3967ff434","avatarUrl":"/avatars/f2dfbc44eb2bff8d8d66d26db8539708.svg","isPro":false,"fullname":"Shuyao Xu","user":"Tim-Xu","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.11699.md","query":{}}">
Negative Self-Distillation: Learning to Reason by Avoiding Flaws
Published on Sep 10
· Submitted by PP on Sep 11 Abstract
Negative Self-Distillation improves large language model reasoning by pushing models away from self-generated flawed reasoning via a dynamic gating mechanism that protects linguistic capabilities.
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs on complex reasoning tasks: By forcing the student to imitate an artificially confident reasoning trace conditioned on privileged information, OPSD inadvertently suppresses expressions of uncertainty and penalizes the exploratory, self-corrective behaviors required to solve challenging problems. To address this, we introduce Negative Self-Distillation (NSD), a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions. Instead of relying on ground-truth answers or external supervision, NSD uses the model itself to generate a question-specific negative condition (eg, acting as a ``careless reasoner'') and pushes the student's distribution away from this self-generated negative teacher. Naively applying unlearning objectives to achieve this divergence is problematic, as flawed reasoning tokens are confounded with basic linguistic tokens; indiscriminately penalizing both risks catastrophically degrading the model's foundational language capabilities. We resolve this by designing a dynamic gating mechanism that automatically identifies and isolates reasoning-critical tokens, ensuring gradient updates target only behavioral flaws while preserving the model's linguistic priors. Empirically, NSD consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
Community
We propose Negative Self-Distillation (NSD), a label-free, fully self-bootstrapped framework that enhances reasoning capabilities by optimizing the model to diverge from self-generated flawed trajectories, eliminating the need for ground-truth solutions or an external teacher.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.11699 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.11699 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.11699 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.