Hugging Face Daily Papers · · 3 min read

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

AgentDebugX is a opensource framework that help you analysis your agent traces and improve its performance</p>\n","updatedAt":"2026-07-22T05:22:41.533Z","author":{"_id":"66554507e6ea63012f35824c","avatarUrl":"/avatars/b82de75bd60890e7bb524fc3754b131c.svg","fullname":"Kunlun_Zhu","name":"Leozkl","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9142158627510071},"editors":["Leozkl"],"editorAvatarUrls":["/avatars/b82de75bd60890e7bb524fc3754b131c.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18754","authors":[{"_id":"6a60530f7e7f152167e4721a","name":"Kunlun Zhu","hidden":false},{"_id":"6a60530f7e7f152167e4721b","user":{"_id":"68905a353cf91a8e828fd8a1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68905a353cf91a8e828fd8a1/lyXMELqDGOHNlXeYlTL5X.jpeg","isPro":false,"fullname":"Xuyan Ye","user":"LulaCola","type":"user","name":"LulaCola"},"name":"Xuyan Ye","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.371Z","hidden":false},{"_id":"6a60530f7e7f152167e4721c","name":"Zhiguang Han","hidden":false},{"_id":"6a60530f7e7f152167e4721d","name":"Yuchen Zhao","hidden":false},{"_id":"6a60530f7e7f152167e4721e","name":"Bingxuan Li","hidden":false},{"_id":"6a60530f7e7f152167e4721f","name":"Weijia Zhang","hidden":false},{"_id":"6a60530f7e7f152167e47220","name":"Muxin Tian","hidden":false},{"_id":"6a60530f7e7f152167e47221","name":"Xiangru Tang","hidden":false},{"_id":"6a60530f7e7f152167e47222","name":"Pan Lu","hidden":false},{"_id":"6a60530f7e7f152167e47223","name":"James Zou","hidden":false},{"_id":"6a60530f7e7f152167e47224","name":"Jiaxuan You","hidden":false},{"_id":"6a60530f7e7f152167e47225","name":"Heng Ji","hidden":false}],"publishedAt":"2026-07-21T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents","submittedOnDailyBy":{"_id":"66554507e6ea63012f35824c","avatarUrl":"/avatars/b82de75bd60890e7bb524fc3754b131c.svg","isPro":false,"fullname":"Kunlun_Zhu","user":"Leozkl","type":"user","name":"Leozkl"},"summary":"LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global trajectory understanding, structure-guided investigation, and cross-examination. On the Who and When benchmark, DeepDebug achieves the best strict attribution accuracy among the evaluated methods on both tested open-weight backbones, reaching 28.8 percent exact agent-and-step accuracy on qwen3.5-9b versus 21.7 percent for the strongest single-pass baseline. On GAIA, DeepDebug repairs 13 of 73 failed tasks in a single rerun, compared with 4 to 6 for three decoupled self-correction baselines, improving overall accuracy from 55.8 percent to 63.6 percent. AgentDebugX exposes this workflow through a Python library, CLI, web console, and installable agentic skill, and provides an opt-in Error Hub for sharing scrubbed failure-diagnosis-repair bundles and reusing them as debugging memory.","upvotes":9,"discussionId":"6a60530f7e7f152167e47226","projectPage":"https://www.agentdebugx.com/","githubRepo":"https://github.com/AgentDebugX/AgentDebugX","githubRepoAddedBy":"user","githubStars":5,"organization":{"_id":"65448bef5b5d9185ba3202b9","name":"UIUC-CS","fullname":"University of Illinois at Urbana-Champaign","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65448b21fcb96b8b48733729/ycqcXFayMTTD_KpE37067.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66554507e6ea63012f35824c","avatarUrl":"/avatars/b82de75bd60890e7bb524fc3754b131c.svg","isPro":false,"fullname":"Kunlun_Zhu","user":"Leozkl","type":"user"},{"_id":"66783baec3f824dde8f783ac","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66783baec3f824dde8f783ac/oqFYUrgs2vnGRhAMSrQpC.jpeg","isPro":false,"fullname":"Jeff","user":"JiayuJeff","type":"user"},{"_id":"64a4f6d23a08801966760665","avatarUrl":"/avatars/6d1661bf9d623fff6446e62aa6e9bfb6.svg","isPro":false,"fullname":"Zhiguang Han","user":"ZhiguangHan","type":"user"},{"_id":"68e70bd43f652a3945d8f4f4","avatarUrl":"/avatars/dea0f1d078e8739d6c727973ba396dcf.svg","isPro":false,"fullname":"Yuchen Zhao","user":"YZ0100","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"63a96be8c847db253f3cab95","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63a96be8c847db253f3cab95/PRK9_qzCvVlI-LpvSaz6s.png","isPro":false,"fullname":"CharlieDreemur","user":"CharlieDreemur","type":"user"},{"_id":"68905a353cf91a8e828fd8a1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68905a353cf91a8e828fd8a1/lyXMELqDGOHNlXeYlTL5X.jpeg","isPro":false,"fullname":"Xuyan Ye","user":"LulaCola","type":"user"},{"_id":"65621fd68631d43d2baf33b2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/1JemNCnPkS1mE3SNygsE2.png","isPro":false,"fullname":"siqi zhu","user":"zsqzz","type":"user"},{"_id":"64e624c34e0203df3b80663c","avatarUrl":"/avatars/22cd52b78918e77016f1192cb06fda27.svg","isPro":false,"fullname":"TianMuxin","user":"realtmxi","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65448bef5b5d9185ba3202b9","name":"UIUC-CS","fullname":"University of Illinois at Urbana-Champaign","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65448b21fcb96b8b48733729/ycqcXFayMTTD_KpE37067.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18754.md","query":{}}">
Papers
arxiv:2607.18754

AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

Published on Jul 21
· Submitted by
Kunlun_Zhu
on Jul 22
Authors:
,

Abstract

LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global trajectory understanding, structure-guided investigation, and cross-examination. On the Who and When benchmark, DeepDebug achieves the best strict attribution accuracy among the evaluated methods on both tested open-weight backbones, reaching 28.8 percent exact agent-and-step accuracy on qwen3.5-9b versus 21.7 percent for the strongest single-pass baseline. On GAIA, DeepDebug repairs 13 of 73 failed tasks in a single rerun, compared with 4 to 6 for three decoupled self-correction baselines, improving overall accuracy from 55.8 percent to 63.6 percent. AgentDebugX exposes this workflow through a Python library, CLI, web console, and installable agentic skill, and provides an opt-in Error Hub for sharing scrubbed failure-diagnosis-repair bundles and reusing them as debugging memory.

Community

Paper submitter about 4 hours ago

AgentDebugX is a opensource framework that help you analysis your agent traces and improve its performance

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.18754
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.18754 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.18754 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.18754 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers