How reliable are LLM judges when evaluating agentic tool-calling systems? 🤔</p>\n<p>We introduce AgentJudgeBench, a benchmark of 3,808 dependency-driven agentic workflows across multiple difficulty levels, generators, and LLM judges. We find that judge reliability degrades sharply with task difficulty, with all judges converging to a 77–82% accuracy ceiling on hard queries without ground truth—regardless of model scale.</p>\n<p>We also find that providing ground truth can sometimes hurt judge alignment, while structured evaluation rubrics can improve it by up to 6.5 points.</p>\n","updatedAt":"2026-09-02T18:40:19.793Z","author":{"_id":"64a3bbc378edd17f9652bf3b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a3bbc378edd17f9652bf3b/XJM8hvtNUxRSMrWvznLHT.jpeg","fullname":"Amit Kumar Saha","name":"amitsaha","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9446277618408203},"editors":["amitsaha"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64a3bbc378edd17f9652bf3b/XJM8hvtNUxRSMrWvznLHT.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.26623","authors":[{"_id":"6a963e90cd6ebc484732ec79","name":"Abhigya Verma","hidden":false},{"_id":"6a963e90cd6ebc484732ec7a","name":"Amit Kumar Saha","hidden":false},{"_id":"6a963e90cd6ebc484732ec7b","name":"Seganrasan Subramanian","hidden":false},{"_id":"6a963e90cd6ebc484732ec7c","name":"Sai Harshitha Aluru","hidden":false}],"publishedAt":"2026-08-27T00:00:00.000Z","submittedOnDailyAt":"2026-09-02T00:00:00.000Z","title":"AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling","submittedOnDailyBy":{"_id":"64a3bbc378edd17f9652bf3b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a3bbc378edd17f9652bf3b/XJM8hvtNUxRSMrWvznLHT.jpeg","isPro":false,"fullname":"Amit Kumar Saha","user":"amitsaha","type":"user","name":"amitsaha"},"summary":"LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators (3B-70B open-weight models and GPT-5.4) and six judges (20B to frontier scale) under paired with- and without-ground-truth conditions. Judge alignment degrades monotonically with task difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a narrow 77-82% band regardless of scale, revealing a structural ceiling driven primarily by task difficulty, though its height is partly prompt-dependent for weaker generators, that model capacity alone cannot overcome. Ground-truth exposure is not uniformly beneficial: it reduces alignment for GPT-5.4 (1.5 pp) and Gemini-2.5-Pro (3.9 pp), consistent with over-anchoring. Among mitigation strategies, chain-of-thought reasoning and judge temperature both have negligible effect, while structured evaluation rubrics improve alignment by up to 6.5 pp but do not generalize uniformly across judge-generator pairs. With ground truth, QwQ-32B best matches the programmatic reference, while a human validation study identifies GPT-OSS-120B as the most human-aligned judge; without it, frontier judges lead only marginally within the shared ceiling. These results expose fundamental limitations of current LLM judges and yield practical guidelines for reliable evaluation in agentic systems.","upvotes":14,"discussionId":"6a963e90cd6ebc484732ec7d","ai_summary":"AgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects.","ai_keywords":["LLM-as-a-judge","agentic tool-calling","workflow DAGs","AgentJudgeBench","chain-of-thought reasoning","structured evaluation rubrics","over-anchoring"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"65f4df5de83b55da5d79fbb6","name":"ServiceNow-AI","fullname":"ServiceNow-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63d3095c2727d7888cbb54e2/Uv-Lx8PVGviqokfOyYlCN.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6697bc94c15a4f7037afbddb","avatarUrl":"/avatars/7258dfa45781065f8a0dad71ef1059fe.svg","isPro":false,"fullname":"Verma","user":"Abhigy","type":"user"},{"_id":"64895ed451b69a8c82f493a4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64895ed451b69a8c82f493a4/0G3lDWoVWiFuFmPfXPKfQ.jpeg","isPro":false,"fullname":"Seganrasan Subramanian","user":"Seganrasan","type":"user"},{"_id":"645c26d423ed9b7788d5e24b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/cZMUluWpYUlSLcn6yoC7c.jpeg","isPro":false,"fullname":"Rishabh Maheshwary","user":"rmahesh","type":"user"},{"_id":"663158da50c203b5a90a82d7","avatarUrl":"/avatars/674b5123ac8aaec5316ed7da29ed601e.svg","isPro":false,"fullname":"Shashank V Maiya","user":"shashankvmaiya","type":"user"},{"_id":"68642da9885da181fde1ad70","avatarUrl":"/avatars/d7399cec3c077e2f1b898b69a49b3b98.svg","isPro":false,"fullname":"Esakkivel Esakkiraja","user":"esakkivel","type":"user"},{"_id":"65ea6558b985ae912c57b294","avatarUrl":"/avatars/54eeba9ce8cb5bf35edfa8fad76a813d.svg","isPro":false,"fullname":"Denis Akhiyarov","user":"dtanow","type":"user"},{"_id":"66d0b470cc4d59dba5c70879","avatarUrl":"/avatars/bed6792ded21bf81ea054da43049e0fa.svg","isPro":false,"fullname":"Tara Bogavelli","user":"tarabogavelli","type":"user"},{"_id":"607f060442beb4da0f990182","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/607f060442beb4da0f990182/j5W2tLyU6JqkaTf3kv66s.jpeg","isPro":false,"fullname":"Patrice Bechard","user":"patricebechard","type":"user"},{"_id":"64820d2bd8662b0714a2a3cd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64820d2bd8662b0714a2a3cd/AHp4bGT05PNlaIGue5gDw.png","isPro":false,"fullname":"Orlando Marquez","user":"marquezo","type":"user"},{"_id":"64a3bbc378edd17f9652bf3b","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a3bbc378edd17f9652bf3b/XJM8hvtNUxRSMrWvznLHT.jpeg","isPro":false,"fullname":"Amit Kumar Saha","user":"amitsaha","type":"user"},{"_id":"66855306fe857bb0701b57e3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66855306fe857bb0701b57e3/rWPqlhLSZoKHY2KEFjLz3.png","isPro":false,"fullname":"Gabrielle Gauthier Melancon","user":"gabegma","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65f4df5de83b55da5d79fbb6","name":"ServiceNow-AI","fullname":"ServiceNow-AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/63d3095c2727d7888cbb54e2/Uv-Lx8PVGviqokfOyYlCN.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.26623.md","query":{}}">
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
Abstract
AgentJudgeBench reveals that LLM judges face structural reliability limits on dependency-driven agentic tool-calling workflows, with alignment degrading by difficulty and ground-truth exposure yielding mixed effects.
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators (3B-70B open-weight models and GPT-5.4) and six judges (20B to frontier scale) under paired with- and without-ground-truth conditions. Judge alignment degrades monotonically with task difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a narrow 77-82% band regardless of scale, revealing a structural ceiling driven primarily by task difficulty, though its height is partly prompt-dependent for weaker generators, that model capacity alone cannot overcome. Ground-truth exposure is not uniformly beneficial: it reduces alignment for GPT-5.4 (1.5 pp) and Gemini-2.5-Pro (3.9 pp), consistent with over-anchoring. Among mitigation strategies, chain-of-thought reasoning and judge temperature both have negligible effect, while structured evaluation rubrics improve alignment by up to 6.5 pp but do not generalize uniformly across judge-generator pairs. With ground truth, QwQ-32B best matches the programmatic reference, while a human validation study identifies GPT-OSS-120B as the most human-aligned judge; without it, frontier judges lead only marginally within the shared ceiling. These results expose fundamental limitations of current LLM judges and yield practical guidelines for reliable evaluation in agentic systems.
Community
How reliable are LLM judges when evaluating agentic tool-calling systems? 🤔
We introduce AgentJudgeBench, a benchmark of 3,808 dependency-driven agentic workflows across multiple difficulty levels, generators, and LLM judges. We find that judge reliability degrades sharply with task difficulty, with all judges converging to a 77–82% accuracy ceiling on hard queries without ground truth—regardless of model scale.
We also find that providing ground truth can sometimes hurt judge alignment, while structured evaluation rubrics can improve it by up to 6.5 points.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.26623 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.26623 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.