LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale.</p>\n","updatedAt":"2026-07-07T02:17:13.709Z","author":{"_id":"69d47558a9acd1eb26637fe9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69d47558a9acd1eb26637fe9/ZWSP5HU9uU7Bc2a1I54eq.jpeg","fullname":"YunHao-Feng","name":"Yunhao-Feng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9017963409423828},"editors":["Yunhao-Feng"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/69d47558a9acd1eb26637fe9/ZWSP5HU9uU7Bc2a1I54eq.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.01793","authors":[{"_id":"6a4c617a25849b193a83400e","name":"Yunhao Feng","hidden":false},{"_id":"6a4c617a25849b193a83400f","name":"Ruixiao Lin","hidden":false},{"_id":"6a4c617a25849b193a834010","name":"Ming Wen","hidden":false},{"_id":"6a4c617a25849b193a834011","name":"Qinqin He","hidden":false},{"_id":"6a4c617a25849b193a834012","name":"Yanming Guo","hidden":false},{"_id":"6a4c617a25849b193a834013","name":"Yifan Ding","hidden":false},{"_id":"6a4c617a25849b193a834014","name":"Yutao Wu","hidden":false},{"_id":"6a4c617a25849b193a834015","name":"Jialuo Chen","hidden":false},{"_id":"6a4c617a25849b193a834016","name":"Zhuoer Xu","hidden":false},{"_id":"6a4c617a25849b193a834017","name":"Xiaohu Du","hidden":false},{"_id":"6a4c617a25849b193a834018","name":"Jianan Ma","hidden":false},{"_id":"6a4c617a25849b193a834019","name":"Zixing Chen","hidden":false},{"_id":"6a4c617a25849b193a83401a","name":"Xingjun Ma","hidden":false},{"_id":"6a4c617a25849b193a83401b","name":"Yunhao Chen","hidden":false},{"_id":"6a4c617a25849b193a83401c","name":"Xinhao Deng","hidden":false}],"publishedAt":"2026-07-04T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification","submittedOnDailyBy":{"_id":"69d47558a9acd1eb26637fe9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69d47558a9acd1eb26637fe9/ZWSP5HU9uU7Bc2a1I54eq.jpeg","isPro":false,"fullname":"YunHao-Feng","user":"Yunhao-Feng","type":"user","name":"Yunhao-Feng"},"summary":"LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9\\% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale. The code is publicly available at https://github.com/Yunhao-Feng/Vera.","upvotes":6,"discussionId":"6a4c617b25849b193a83401d","githubRepo":"https://github.com/Yunhao-Feng/Vera","githubRepoAddedBy":"user","ai_summary":"Automated safety testing framework Vera uses a three-stage pipeline to identify and test safety risks in LLM agents through structured risk taxonomies, combinatorial case generation, and adaptive sandbox execution with evidence-based verification.","ai_keywords":["LLM agents","safety testing","software engineering testing principles","three-stage pipeline","risk taxonomies","combinatorial composition","executable safety cases","heterogeneous agents","isolated sandboxes","control agent","evidence-grounded verifiers"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":3,"organization":{"_id":"67c1d682826160b28f778510","name":"antgroup","fullname":"Ant Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/7VcPHdLSGlged3ixK1dys.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"66d8512c54209e9101811e8e","avatarUrl":"/avatars/62dfd8e6261108f2508efe678d5a2a57.svg","isPro":false,"fullname":"M Saad Salman","user":"MSS444","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"69ccc455eb9cdf88f2a23965","avatarUrl":"/avatars/72c60a0ff542a917c37c5978a6a04423.svg","isPro":false,"fullname":"Xie Chenxi","user":"benjaminlewis","type":"user"},{"_id":"6991544a2d8293389779f18f","avatarUrl":"/avatars/b28f76a31f0c77f3b4313b53ccdbcf15.svg","isPro":false,"fullname":"Cexoijv65vud0","user":"cexoijv65vud0","type":"user"},{"_id":"69bd5535778f1b4b3c025cdb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/zAvE-7ARipM69vC1-HT4W.png","isPro":false,"fullname":"Feng Linxi","user":"sophiar576","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67c1d682826160b28f778510","name":"antgroup","fullname":"Ant Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/7VcPHdLSGlged3ixK1dys.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.01793.md","query":{}}">
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
Authors: ,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Automated safety testing framework Vera uses a three-stage pipeline to identify and test safety risks in LLM agents through structured risk taxonomies, combinatorial case generation, and adaptive sandbox execution with evidence-based verification.
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9\% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale. The code is publicly available at https://github.com/Yunhao-Feng/Vera.
Community
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and the corresponding outcomes are evaluated by hard-coded rules, making them costly to extend as agents evolve. To this end, we present Vera, an end-to-end automated safety testing framework that instantiates software engineering testing principles for non-deterministic agents through a three-stage, self-reinforcing pipeline. First, a literature-driven exploration continuously discovers and structures emerging risks into taxonomies of safety risks, attack methods, and tool execution environments. Second, combinatorial composition across taxonomy dimensions produces executable safety cases, each specifying a concrete safety goal, a programmatically constructed initial state, and a deterministic verification predicate grounded in observable artifacts. Third, adaptive execution runs heterogeneous agents in isolated sandboxes where a control agent steers multi-turn interaction based on runtime observations, while evidence-grounded verifiers judge outcomes from environment state and tool-call evidence rather than model self-report. We evaluate Vera on four production agent frameworks (OpenClaw, Hermes, Codex, Claude Code), revealing substantial safety weaknesses, with average attack success rates reaching 93.9% under multi-channel attacks; we also release Vera-Bench, comprising 1600 executable safety cases spanning 124 risk categories across three execution settings. These results indicate that modular, executable testing infrastructure is essential for rigorous and maintainable safety evaluation of rapidly evolving agentic systems at scale.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.01793 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.01793 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.01793 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.