Multimodal agents for visual question answering increasingly operate as multi-step<br>trajectories that interleave perception, retrieval, and reasoning, yet evaluation still<br>largely reduces to final-answer accuracy. This aggregate signal cannot tell whether<br>a correct answer was reached through grounded evidence, language priors, or<br>accidental error cancellation. We propose to treat a multimodal agent trajectory as<br>a provenance-constrained state machine: tool outputs are normalized into a Struc-<br>tured Evidence Ledger that serves as the trajectory state, downstream reasoning<br>and decision claims may cite only active ledger entries, grounding is checked at<br>the entity and numeric level, and repair is realized as typed state transitions that<br>cannot introduce content without tool-produced provenance. We instantiate this<br>design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning<br>with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Pro-<br>tocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question<br>complexity, and an Event-Triggered Verification-and-Repair engine with a formal<br>provenance non-amplification guarantee. We use LedgerMind to target four re-<br>curring failure patterns that final-answer accuracy tends to obscure: unsupported<br>intermediate reasoning, citation-backed entity hallucination (Phantom Grounding),<br>over-reasoning on simple queries, and repair-time amplification. Experiments<br>across multiple multimodal reasoning benchmarks and backbone MLLMs show<br>that LedgerMind improves both answer accuracy and trajectory-level faithfulness.</p>\n","updatedAt":"2026-07-31T03:58:44.188Z","author":{"_id":"66decf61f9971122eec44dc8","avatarUrl":"/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg","fullname":"Enjun Du","name":"EnjunDu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9042862057685852},"editors":["EnjunDu"],"editorAvatarUrls":["/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.28374","authors":[{"_id":"6a6c1c98202e2d9e3ffdb7c0","name":"Enjun Du","hidden":false},{"_id":"6a6c1c98202e2d9e3ffdb7c1","name":"Hange Zhou","hidden":false},{"_id":"6a6c1c98202e2d9e3ffdb7c2","name":"Chenxu Du","hidden":false},{"_id":"6a6c1c98202e2d9e3ffdb7c3","name":"Siyi Liu","hidden":false},{"_id":"6a6c1c98202e2d9e3ffdb7c4","name":"Zirong Chen","hidden":false},{"_id":"6a6c1c98202e2d9e3ffdb7c5","name":"Ziyu Zheng","hidden":false},{"_id":"6a6c1c98202e2d9e3ffdb7c6","name":"Yongqi Zhang","hidden":false}],"publishedAt":"2026-07-30T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger","submittedOnDailyBy":{"_id":"66decf61f9971122eec44dc8","avatarUrl":"/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg","isPro":false,"fullname":"Enjun Du","user":"EnjunDu","type":"user","name":"EnjunDu"},"summary":"Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.","upvotes":9,"discussionId":"6a6c1c99202e2d9e3ffdb7c7","projectPage":"https://enjundu.com/blog/ledgermind/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"668fcd97495674543600db0b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/JfBeJ6gSXe379G9ejjrLR.jpeg","isPro":false,"fullname":"Gao Mingqi","user":"EchoMinkki","type":"user"},{"_id":"66decf61f9971122eec44dc8","avatarUrl":"/avatars/ffd1bf114f2fffc1f9a2ffe5964543f3.svg","isPro":false,"fullname":"Enjun Du","user":"EnjunDu","type":"user"},{"_id":"65afbba21edab235a1323ad0","avatarUrl":"/avatars/bd047a559821e2bc802d66a073b994df.svg","isPro":true,"fullname":"3","user":"xuzishan","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"66d05337e62d6bbf50186c2f","avatarUrl":"/avatars/f5be15e754f0fbbb37d2cc5ea417f729.svg","isPro":false,"fullname":"Yijie Jin","user":"DrewJin0827","type":"user"},{"_id":"687f853bb39262ba84f3eeff","avatarUrl":"/avatars/cdfc44fde8237f08f10192553fe5a075.svg","isPro":false,"fullname":"Junhao Shen","user":"shenjunhao","type":"user"},{"_id":"6858f7393794a7a08a254364","avatarUrl":"/avatars/aa2870bed47f7ef25356a4e9ff54ce38.svg","isPro":false,"fullname":"kexu Cheng*","user":"sediment1024","type":"user"},{"_id":"681efa4f2ed93571115986c9","avatarUrl":"/avatars/3af882b69da8459c6091433db83d5054.svg","isPro":false,"fullname":"Tower","user":"HideOnTower","type":"user"},{"_id":"67d62da597767f49259a55d6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/3bsqzf_J0Kv9tZ3qvZsE8.png","isPro":false,"fullname":"Xiangnan Wu","user":"XiangnanW","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.28374.md","query":{}}">
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
Abstract
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.
Community
Multimodal agents for visual question answering increasingly operate as multi-step
trajectories that interleave perception, retrieval, and reasoning, yet evaluation still
largely reduces to final-answer accuracy. This aggregate signal cannot tell whether
a correct answer was reached through grounded evidence, language priors, or
accidental error cancellation. We propose to treat a multimodal agent trajectory as
a provenance-constrained state machine: tool outputs are normalized into a Struc-
tured Evidence Ledger that serves as the trajectory state, downstream reasoning
and decision claims may cite only active ledger entries, grounding is checked at
the entity and numeric level, and repair is realized as typed state transitions that
cannot introduce content without tool-produced provenance. We instantiate this
design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning
with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Pro-
tocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question
complexity, and an Event-Triggered Verification-and-Repair engine with a formal
provenance non-amplification guarantee. We use LedgerMind to target four re-
curring failure patterns that final-answer accuracy tends to obscure: unsupported
intermediate reasoning, citation-backed entity hallucination (Phantom Grounding),
over-reasoning on simple queries, and repair-time amplification. Experiments
across multiple multimodal reasoning benchmarks and backbone MLLMs show
that LedgerMind improves both answer accuracy and trajectory-level faithfulness.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.28374 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.28374 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.28374 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.