\n\t<a id=\"🚀-visual-contrastive-self-distillation-vcsd\" class=\"block pr-1.5 text-lg md:absolute md:p-1.5 md:opacity-0 md:group-hover:opacity-100 md:right-full\" href=\"#🚀-visual-contrastive-self-distillation-vcsd\" rel=\"nofollow\">\n\t\t<span class=\"header-link\"><svg class=\"text-gray-500 hover:text-black dark:hover:text-gray-200 w-4\" xmlns=\"http://www.w3.org/2000/svg\" xmlns:xlink=\"http://www.w3.org/1999/xlink\" aria-hidden=\"true\" role=\"img\" width=\"1em\" height=\"1em\" preserveAspectRatio=\"xMidYMid meet\" viewBox=\"0 0 256 256\"><path d=\"M167.594 88.393a8.001 8.001 0 0 1 0 11.314l-67.882 67.882a8 8 0 1 1-11.314-11.315l67.882-67.881a8.003 8.003 0 0 1 11.314 0zm-28.287 84.86l-28.284 28.284a40 40 0 0 1-56.567-56.567l28.284-28.284a8 8 0 0 0-11.315-11.315l-28.284 28.284a56 56 0 0 0 79.196 79.197l28.285-28.285a8 8 0 1 0-11.315-11.314zM212.852 43.14a56.002 56.002 0 0 0-79.196 0l-28.284 28.284a8 8 0 1 0 11.314 11.314l28.284-28.284a40 40 0 0 1 56.568 56.567l-28.285 28.285a8 8 0 0 0 11.315 11.314l28.284-28.284a56.065 56.065 0 0 0 0-79.196z\" fill=\"currentColor\"></path></svg></span>\n\t</a>\n\t<span>\n\t\t🚀 Visual Contrastive Self-Distillation (VCSD)\n\t</span>\n</h1>\n<p>Can on-policy self-distillation improve Vision-Language Models without privileged answers or auxiliary visual evidence? <strong>Yes.</strong></p>\n<ul>\n<li><p><strong>The Core Idea:</strong> VCSD compares the same EMA teacher under two matched visual conditions, the original image and a content-erased control, while keeping the prompt and student-generated response prefix unchanged.</p>\n</li>\n<li><p><strong>Pure Input Conditioning:</strong> The token-wise prediction difference reveals which candidates are specifically supported by the instance-level image content, turning input contrast into a self-distillation signal.</p>\n</li>\n<li><p><strong>Contrast-Shaped Target:</strong> VCSD uses this visual contrast to sharpen the teacher’s original-image distribution within its plausible support, then distills the resulting full-distribution target into the student.</p>\n</li>\n<li><p><strong>No Auxiliary Supervision:</strong> No external teacher, answer hints, reasoning traces, cropped or annotated visual evidence, or additional inference-time cost.</p>\n</li>\n<li><p><strong>Consistent Gains:</strong> Across <strong>6 Qwen3-VL/Qwen3.5 models</strong> and <strong>7 vision-language benchmarks</strong>, VCSD outperforms matched OPSD on every model by <strong>+1.76 to +5.33 points</strong>. It also improves over the corresponding base models by up to <strong>+4.77 points</strong>, with Qwen3.5-9B reaching <strong>79.24 average accuracy</strong>.</p>\n</li>\n</ul>\n<p>🌐 <strong>Project and Code:</strong> <a href=\"https://joliang17.github.io/VisualCSD/\" rel=\"nofollow\">https://joliang17.github.io/VisualCSD/</a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/66720ab819bebc69b5b93685/A--sk6GFahc5c-lfTPXfl.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/66720ab819bebc69b5b93685/A--sk6GFahc5c-lfTPXfl.png\" alt=\"fig2_method\"></a></p>\n","updatedAt":"2026-07-24T04:27:25.053Z","author":{"_id":"66720ab819bebc69b5b93685","avatarUrl":"/avatars/b2f1314d9a26f6f5eaf6cebdb0d28812.svg","fullname":"Yijun Liang","name":"joliang17","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7741034030914307},"editors":["joliang17"],"editorAvatarUrls":["/avatars/b2f1314d9a26f6f5eaf6cebdb0d28812.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.21556","authors":[{"_id":"6a62cb742ee212ed0e2a1485","name":"Yijun Liang","hidden":false},{"_id":"6a62cb742ee212ed0e2a1486","name":"Yunjie Tian","hidden":false},{"_id":"6a62cb742ee212ed0e2a1487","name":"Yijiang Li","hidden":false},{"_id":"6a62cb742ee212ed0e2a1488","name":"Yuqi Jia","hidden":false},{"_id":"6a62cb742ee212ed0e2a1489","name":"Furong Huang","hidden":false},{"_id":"6a62cb742ee212ed0e2a148a","user":{"_id":"647f5af5b0e96764589f3b2a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/647f5af5b0e96764589f3b2a/DzYsWWmVfcVQ55I-CQT4g.jpeg","isPro":false,"fullname":"Tianyi Zhou","user":"zhoutianyi","type":"user","name":"zhoutianyi"},"name":"Tianyi Zhou","status":"claimed_verified","statusLastChangedAt":"2026-07-24T08:45:04.301Z","hidden":false},{"_id":"6a62cb742ee212ed0e2a148b","name":"Di Fu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66720ab819bebc69b5b93685/DAWO1NIJw1bQwXlnpqvAw.png"],"publishedAt":"2026-07-23T00:00:00.000Z","submittedOnDailyAt":"2026-07-24T00:00:00.000Z","title":"Visual Contrastive Self-Distillation","submittedOnDailyBy":{"_id":"66720ab819bebc69b5b93685","avatarUrl":"/avatars/b2f1314d9a26f6f5eaf6cebdb0d28812.svg","isPro":true,"fullname":"Yijun Liang","user":"joliang17","type":"user","name":"joliang17"},"summary":"On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix -- one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. Using ViRL39K dataset, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. For example, on Qwen3-VL, it improves the seven-benchmark aggregate from 62.27% rightarrow 67.04% at 2B, 71.30% rightarrow 73.16% at 4B, and 72.51% rightarrow 76.26% at 8B. Furthermore, VCSD requires no external teacher, privileged answers, visual evidence signals, reasoning traces, or additional inference-time cost.","upvotes":39,"discussionId":"6a62cb742ee212ed0e2a148c","projectPage":"https://joliang17.github.io/VisualCSD/","githubRepo":"https://github.com/joliang17/VCSD","githubRepoAddedBy":"user","githubStars":4,"organization":{"_id":"68b3c3bbc375e05b059370b2","name":"UMCP","fullname":"University of Maryland College Park","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68b3c2c3a4ea236d1a97871a/bji3nI5ZWm2r4JX_-HLo0.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64a8121e35fab7cd04c30ed0","avatarUrl":"/avatars/48849b84703158772f1022932331b143.svg","isPro":false,"fullname":"Chenrui Fan","user":"Fcr09","type":"user"},{"_id":"66720ab819bebc69b5b93685","avatarUrl":"/avatars/b2f1314d9a26f6f5eaf6cebdb0d28812.svg","isPro":true,"fullname":"Yijun Liang","user":"joliang17","type":"user"},{"_id":"65031d01cccc7b28a388c719","avatarUrl":"/avatars/9d8c94b6ab8ad8b4faba3221b7e76053.svg","isPro":false,"fullname":"Ming Li","user":"MingLiiii","type":"user"},{"_id":"6419309f22270b3ccf177c77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6419309f22270b3ccf177c77/KQa1586iBBKqucUlfpuPp.jpeg","isPro":true,"fullname":"William Li","user":"williamium","type":"user"},{"_id":"668431ccf0236757f43df540","avatarUrl":"/avatars/fcb5394860d92a7e304942df9de5d1e3.svg","isPro":false,"fullname":"Ziyue Li","user":"Litzy0619","type":"user"},{"_id":"6658798231baf30dc75c3dc4","avatarUrl":"/avatars/db6b6afa5eabde9c7780ef50555220d9.svg","isPro":false,"fullname":"QIao","user":"Qiao111111","type":"user"},{"_id":"640729055e6d06cc2cf3bee0","avatarUrl":"/avatars/39fb99188c67ca5409b4e4d458d57cd9.svg","isPro":false,"fullname":"lee","user":"mmmwhy","type":"user"},{"_id":"648d45c429b7d08bf87f36da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/0fBMJN9PNn4ytedZOLtHC.jpeg","isPro":false,"fullname":"BaiYi","user":"DarrenZ","type":"user"},{"_id":"68a8d6593cb048199c2b80fc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/vV2b7HB48f-96YL8WnAWx.png","isPro":false,"fullname":"Bugjudger","user":"Bugjudger","type":"user"},{"_id":"674e2180601d3bcedbff902c","avatarUrl":"/avatars/78a2d5ced5448eb8f22a94b45b2f4f52.svg","isPro":false,"fullname":"Xiaoou","user":"xiao0o0o","type":"user"},{"_id":"63673bb9d0ee6e2662be0ec1","avatarUrl":"/avatars/1b8976785d64bc4e3f7159ccdb7f06c5.svg","isPro":false,"fullname":"Qingqiao Hu","user":"WinstonHu","type":"user"},{"_id":"6470f263be66c5bacd3018b1","avatarUrl":"/avatars/34c1da7273115a037104b9dc94f662af.svg","isPro":false,"fullname":"Jingchen Sun","user":"jsun39","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":3,"organization":{"_id":"68b3c3bbc375e05b059370b2","name":"UMCP","fullname":"University of Maryland College Park","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68b3c2c3a4ea236d1a97871a/bji3nI5ZWm2r4JX_-HLo0.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.21556.md","query":{}}">
Visual Contrastive Self-Distillation
Abstract
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we propose Visual Contrastive Self-Distillation, namely VCSD, which converts image-content removal into an on-policy self-distillation signal. At each student-generated response prefix, the EMA teacher produces two next-token distributions under the same prompt and prefix -- one conditioned on the original image and the other on a content-erased control. Their token-wise log-probability difference highlights candidates whose likelihood is specifically increased by the instance-level visual content. We use this contrast to sharpen the teacher's original-image distribution within its plausible support, and distill the resulting full-distribution target into the student. Using ViRL39K dataset, VCSD consistently outperforms matched OPSD across Qwen3-VL and Qwen3.5 models. For example, on Qwen3-VL, it improves the seven-benchmark aggregate from 62.27% rightarrow 67.04% at 2B, 71.30% rightarrow 73.16% at 4B, and 72.51% rightarrow 76.26% at 8B. Furthermore, VCSD requires no external teacher, privileged answers, visual evidence signals, reasoning traces, or additional inference-time cost.
Community
🚀 Visual Contrastive Self-Distillation (VCSD)
Can on-policy self-distillation improve Vision-Language Models without privileged answers or auxiliary visual evidence? Yes.
The Core Idea: VCSD compares the same EMA teacher under two matched visual conditions, the original image and a content-erased control, while keeping the prompt and student-generated response prefix unchanged.
Pure Input Conditioning: The token-wise prediction difference reveals which candidates are specifically supported by the instance-level image content, turning input contrast into a self-distillation signal.
Contrast-Shaped Target: VCSD uses this visual contrast to sharpen the teacher’s original-image distribution within its plausible support, then distills the resulting full-distribution target into the student.
No Auxiliary Supervision: No external teacher, answer hints, reasoning traces, cropped or annotated visual evidence, or additional inference-time cost.
Consistent Gains: Across 6 Qwen3-VL/Qwen3.5 models and 7 vision-language benchmarks, VCSD outperforms matched OPSD on every model by +1.76 to +5.33 points. It also improves over the corresponding base models by up to +4.77 points, with Qwen3.5-9B reaching 79.24 average accuracy.
🌐 Project and Code: https://joliang17.github.io/VisualCSD/

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.21556 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.21556 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.21556 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.