In this study, we propose ToolHazard, a scalable adversarial environment synthesis framework that enables agent security evaluation and adversarial alignment across broader application domains.</p>\n","updatedAt":"2026-08-13T01:35:54.584Z","author":{"_id":"66632a3d2dc4dff9a98c38a5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66632a3d2dc4dff9a98c38a5/UYk99yT6M1mZviL8d2sm_.jpeg","fullname":"Yutao Mou","name":"MurrayTom","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8184682130813599},"editors":["MurrayTom"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66632a3d2dc4dff9a98c38a5/UYk99yT6M1mZviL8d2sm_.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.11878","authors":[{"_id":"6a7d1e1b0ac8bee77474edf3","name":"Yutao Mou","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edf4","name":"Pengfei Yang","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edf5","name":"Zhe Yin","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edf6","name":"Zhangchi Xue","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edf7","name":"Xiaotian Luan","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edf8","name":"Dingyao Yu","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edf9","name":"Tong Zhang","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edfa","name":"Shikun Zhang","hidden":false},{"_id":"6a7d1e1b0ac8bee77474edfb","name":"Wei Ye","hidden":false}],"publishedAt":"2026-08-12T00:00:00.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents","submittedOnDailyBy":{"_id":"66632a3d2dc4dff9a98c38a5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66632a3d2dc4dff9a98c38a5/UYk99yT6M1mZviL8d2sm_.jpeg","isPro":false,"fullname":"Yutao Mou","user":"MurrayTom","type":"user","name":"MurrayTom"},"summary":"Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.","upvotes":6,"discussionId":"6a7d1e1b0ac8bee77474edfc","githubRepo":"https://github.com/MurrayTom/ToolHazard","githubRepoAddedBy":"user","ai_summary":"ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment.","ai_keywords":["indirect prompt injection","LLM agents","adversarial environment synthesis","ToolHazard","environment simulator","attacker agent","user simulator","ToolHazard-Bench","alignment data","AgentDojo"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66632a3d2dc4dff9a98c38a5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66632a3d2dc4dff9a98c38a5/UYk99yT6M1mZviL8d2sm_.jpeg","isPro":false,"fullname":"Yutao Mou","user":"MurrayTom","type":"user"},{"_id":"6908762fde5aa00a3d4038ee","avatarUrl":"/avatars/dd092069f948362c92fa60a6b4d24b10.svg","isPro":false,"fullname":"tong zhang","user":"T2shen","type":"user"},{"_id":"6a56e9799828b80161b0d6f6","avatarUrl":"/avatars/24935d6d5d34c9e49330a445314b39e5.svg","isPro":false,"fullname":"WUshengdong","user":"joy202006","type":"user"},{"_id":"65082d1dffc738079cabdbae","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65082d1dffc738079cabdbae/2qo2GdEGpcoGsKGMesKsb.jpeg","isPro":false,"fullname":"XXHStudyHard","user":"XXHStudyHard","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"64e9f5bbf494f8b2a0670eee","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/rM0q9i3DXc7nVY6MDjiCY.jpeg","isPro":false,"fullname":"Rijusmit Biswas","user":"Phantomcloak19","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61dcd8e344f59573371b5cb6","name":"PekingUniversity","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/vavgrBsnkSejriUF4lXDE.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.11878.md","query":{}}">
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
Abstract
ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment.
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.
Community
In this study, we propose ToolHazard, a scalable adversarial environment synthesis framework that enables agent security evaluation and adversarial alignment across broader application domains.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.11878 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.11878 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.11878 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.