preview</p>\n","updatedAt":"2026-08-04T02:16:31.889Z","author":{"_id":"63b6def76fca9d2a1902fa14","avatarUrl":"/avatars/c7f2487450ea954e2bca4fc5a6db8eb3.svg","fullname":"张康宁","name":"zhangkangning","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5315865874290466},"editors":["zhangkangning"],"editorAvatarUrls":["/avatars/c7f2487450ea954e2bca4fc5a6db8eb3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28590","authors":[{"_id":"6a714b7eec5082b9f872cce2","name":"Kangning Zhang","hidden":false},{"_id":"6a714b7eec5082b9f872cce3","name":"Yixing Li","hidden":false},{"_id":"6a714b7eec5082b9f872cce4","name":"Shuai Shao","hidden":false},{"_id":"6a714b7eec5082b9f872cce5","name":"Qingyao Li","hidden":false},{"_id":"6a714b7eec5082b9f872cce6","name":"Zhengxi Lu","hidden":false},{"_id":"6a714b7eec5082b9f872cce7","name":"Zhiyuan Yao","hidden":false},{"_id":"6a714b7eec5082b9f872cce8","name":"Jianghao Lin","hidden":false},{"_id":"6a714b7eec5082b9f872cce9","name":"Wenxiang Jiao","hidden":false},{"_id":"6a714b7eec5082b9f872ccea","name":"Yuan Lu","hidden":false},{"_id":"6a714b7eec5082b9f872cceb","name":"Weiwen Liu","hidden":false},{"_id":"6a714b7eec5082b9f872ccec","name":"Weinan Zhang","hidden":false},{"_id":"6a714b7eec5082b9f872cced","name":"Yong Yu","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-08-04T00:00:00.000Z","title":"VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation","submittedOnDailyBy":{"_id":"63b6def76fca9d2a1902fa14","avatarUrl":"/avatars/c7f2487450ea954e2bca4fc5a6db8eb3.svg","isPro":false,"fullname":"张康宁","user":"zhangkangning","type":"user","name":"zhangkangning"},"summary":"Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.","upvotes":26,"discussionId":"6a714b7fec5082b9f872ccee","githubRepo":"https://github.com/DeepExperience/VAD_Multimodal_OPD","githubRepoAddedBy":"user","githubStars":16},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63b6def76fca9d2a1902fa14","avatarUrl":"/avatars/c7f2487450ea954e2bca4fc5a6db8eb3.svg","isPro":false,"fullname":"张康宁","user":"zhangkangning","type":"user"},{"_id":"64c4ab0388373ea6200e1cf3","avatarUrl":"/avatars/8ad27c35d3def048bc4ff96c0510bba6.svg","isPro":false,"fullname":"qingyao li","user":"simonlqy","type":"user"},{"_id":"67f5454e1f54f65efc9ce06b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/QTh1c-okqsl-NLHatc8Gu.png","isPro":false,"fullname":"Shijian Wang","user":"ShijianW01","type":"user"},{"_id":"63db16330cc3bc12bc0b6f8f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63db16330cc3bc12bc0b6f8f/ld0JQIfX1SBlDVDOmw9VT.jpeg","isPro":false,"fullname":"Wenxiang Jiao","user":"wxjiao","type":"user"},{"_id":"656800a7456d7733de5f1d89","avatarUrl":"/avatars/910ace29b232523159f1c6bbd6e7e8ad.svg","isPro":false,"fullname":"Rong Shan","user":"CyberDancer","type":"user"},{"_id":"66d7bd78b9e69dfa9baeaebb","avatarUrl":"/avatars/3e58fc51d892f5f89be55c6446766c1c.svg","isPro":false,"fullname":"jiachen zhu","user":"gebro13","type":"user"},{"_id":"6a01ea8de92e139640ca67e5","avatarUrl":"/avatars/05f44f57f2a076957b36a0f81b10c865.svg","isPro":false,"fullname":"Lee qc","user":"BierLee","type":"user"},{"_id":"64e184d3e3d040e495ba41d3","avatarUrl":"/avatars/830a84cee4f3e277221f573155bb97eb.svg","isPro":false,"fullname":"Weiming Zhang","user":"Yevzh","type":"user"},{"_id":"6a7155ce1d662199cb6e1d83","avatarUrl":"/avatars/ba5536fdba1d5bd22330f815c462b5e3.svg","isPro":false,"fullname":"qin wan","user":"qinwan03520","type":"user"},{"_id":"6495aecbae7d6475a87d64ac","avatarUrl":"/avatars/b3f6d04bd9d00e5b1236e519dc1eb397.svg","isPro":false,"fullname":"MarkF","user":"MarkFht","type":"user"},{"_id":"63d2775d5c52bbd72cac2e86","avatarUrl":"/avatars/88a1391352be2c760ee033fe456dd738.svg","isPro":false,"fullname":"KaiwenZhu","user":"Kaiwen-Zhu","type":"user"},{"_id":"6a43a04c1bb0afac11e45d8a","avatarUrl":"/avatars/6facd16bd179df2723f6fa015242dc60.svg","isPro":false,"fullname":"WillieZhou","user":"WillieZhou111","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.28590.md","query":{}}">
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
Published on Jul 30
· Submitted by 张康宁 on Aug 4 Abstract
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.28590 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.28590 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.28590 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.