Hugging Face Daily Papers · · 4 min read

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.</p>\n","updatedAt":"2026-08-11T01:52:46.717Z","author":{"_id":"69c8d4631d39906ca9ded454","avatarUrl":"/avatars/38bf2d0a4b84334940b701bd17333af2.svg","fullname":"Haojie Huang","name":"hhj-ai","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8376783132553101},"editors":["hhj-ai"],"editorAvatarUrls":["/avatars/38bf2d0a4b84334940b701bd17333af2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.08021","authors":[{"_id":"6a7a8060019ce76dc7b3a983","user":{"_id":"69c8d4631d39906ca9ded454","avatarUrl":"/avatars/38bf2d0a4b84334940b701bd17333af2.svg","isPro":false,"fullname":"Haojie Huang","user":"hhj-ai","type":"user","name":"hhj-ai"},"name":"Haojie Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-11T08:45:04.414Z","hidden":false},{"_id":"6a7a8060019ce76dc7b3a984","name":"Xinlei Yu","hidden":false},{"_id":"6a7a8060019ce76dc7b3a985","name":"Chengming Xu","hidden":false},{"_id":"6a7a8060019ce76dc7b3a986","name":"Zhangquan Chen","hidden":false},{"_id":"6a7a8060019ce76dc7b3a987","name":"Cheng Yang","hidden":false},{"_id":"6a7a8060019ce76dc7b3a988","name":"Qingdong He","hidden":false},{"_id":"6a7a8060019ce76dc7b3a989","name":"Yu Yang","hidden":false},{"_id":"6a7a8060019ce76dc7b3a98a","name":"Jiangning Zhang","hidden":false},{"_id":"6a7a8060019ce76dc7b3a98b","name":"Xiaobin Hu","hidden":false}],"publishedAt":"2026-08-08T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Evidence-RL: Towards Evidence-intensive Visual Reasoning","submittedOnDailyBy":{"_id":"69c8d4631d39906ca9ded454","avatarUrl":"/avatars/38bf2d0a4b84334940b701bd17333af2.svg","isPro":false,"fullname":"Haojie Huang","user":"hhj-ai","type":"user","name":"hhj-ai"},"summary":"Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.","upvotes":9,"discussionId":"6a7a8061019ce76dc7b3a98c","projectPage":"https://evidencerl.github.io/","ai_summary":"Counterfactual Evidence Disentanglement improves vision-language model grounding by auditing whether answers causally depend on local visual evidence during reinforcement learning post-training.","ai_keywords":["Vision-Language Models","Counterfactual Evidence Disentanglement","evidence region","GRPO","object-centric","reinforcement learning","grounding"],"ai_summary_model":"thinkingmachines/Inkling-Small"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69c8d4631d39906ca9ded454","avatarUrl":"/avatars/38bf2d0a4b84334940b701bd17333af2.svg","isPro":false,"fullname":"Haojie Huang","user":"hhj-ai","type":"user"},{"_id":"6a42148f876ee2b4164c40d6","avatarUrl":"/avatars/be0c1f94d373fd19ee2167af27c822e5.svg","isPro":false,"fullname":"Peiling Zhu","user":"Zhuninu","type":"user"},{"_id":"67d63e228d5c7a132cbcf39b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ynwA3Sya5irwMRCmSeLiC.png","isPro":false,"fullname":"neil yu","user":"yxl66666","type":"user"},{"_id":"688dffa7bc2095468e8740d8","avatarUrl":"/avatars/8842daefbd12fb3386370351ba1597a6.svg","isPro":false,"fullname":"peter","user":"EAIer","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"69d8e87662baef56eb6c1927","avatarUrl":"/avatars/8f5b735764c2c15aee7fef15825eb9ab.svg","isPro":false,"fullname":"Charles Zheng","user":"Ethmolks","type":"user"},{"_id":"671b5ce59e5016396edcc78a","avatarUrl":"/avatars/d4af23e312f1e90b9419b1cd8e908b87.svg","isPro":false,"fullname":"ZhengQi Wan","user":"Vanqi","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"6a2c346a6ac1d030d3432f17","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/arvfhlqU3h74nR_ldfeIn.jpeg","isPro":false,"fullname":"Jonathan Boilard","user":"joboilard","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.08021.md","query":{}}">
Papers
arxiv:2608.08021

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Published on Aug 8
· Submitted by
Haojie Huang
on Aug 11
Authors:

Abstract

Counterfactual Evidence Disentanglement improves vision-language model grounding by auditing whether answers causally depend on local visual evidence during reinforcement learning post-training.

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.

Community

Paper author Paper submitter about 17 hours ago

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.08021
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.08021 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.08021 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.08021 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers