Hugging Face Daily Papers · · 5 min read

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition.</p>\n<p>We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at <a href=\"https://stealthbench.com\" rel=\"nofollow\">stealthbench.com</a>.</p>\n<p><em>29 pages, 9 tables, 1 figure, 2 appendices. Code: <a href=\"https://github.com/GangGreenTemperTatum/stealthbench\" rel=\"nofollow\">https://github.com/GangGreenTemperTatum/stealthbench</a>, URL Dataset: <a href=\"https://huggingface.co/datasets/0xmoose/stealthbench\">https://huggingface.co/datasets/0xmoose/stealthbench</a> Website: <a href=\"https://stealthbench.com\" rel=\"nofollow\">https://stealthbench.com</a></em></p>\n","updatedAt":"2026-07-30T02:28:09.263Z","author":{"_id":"660f04586d98d685ae4944ad","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660f04586d98d685ae4944ad/j-0SoUk0VU6RDbr-jwdcq.png","fullname":"Ads Dawson","name":"0xmoose","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":20,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/_i8yeYGDCwT7C8nczgZmC.png","fullname":"dreadnode","name":"dreadnode","type":"org","isHf":false,"plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8760216236114502},"editors":["0xmoose"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/660f04586d98d685ae4944ad/j-0SoUk0VU6RDbr-jwdcq.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.26314","authors":[{"_id":"6a6a9fb04463a8a84bdc3f29","name":"Ads Dawson","hidden":false},{"_id":"6a6a9fb04463a8a84bdc3f2a","name":"Adrian Wood","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/660f04586d98d685ae4944ad/rLlQyru2rfz0ADzOcQ6Ta.png"],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-30T00:00:00.000Z","title":"StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents","submittedOnDailyBy":{"_id":"660f04586d98d685ae4944ad","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660f04586d98d685ae4944ad/j-0SoUk0VU6RDbr-jwdcq.png","isPro":false,"fullname":"Ads Dawson","user":"0xmoose","type":"user","name":"0xmoose"},"summary":"Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition.\n We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at https://stealthbench.com.","upvotes":2,"discussionId":"6a6a9fb04463a8a84bdc3f2b","projectPage":"https://stealthbench.com/","githubRepo":"https://github.com/GangGreenTemperTatum/stealthbench","githubRepoAddedBy":"user","githubStars":0},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6a14573d44d88d90e2d36d50","avatarUrl":"/avatars/59f01f7926865c48a550c572554631e3.svg","isPro":false,"fullname":"葵 渡辺","user":"jacobbaker2023","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.26314.md","query":{}}">
Papers
arxiv:2607.26314

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Published on Jul 28
· Submitted by
Ads Dawson
on Jul 30
Authors:
,

Abstract

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition. We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at https://stealthbench.com.

Community

Paper submitter about 6 hours ago

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition.

We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at stealthbench.com.

29 pages, 9 tables, 1 figure, 2 appendices. Code: https://github.com/GangGreenTemperTatum/stealthbench, URL Dataset: https://huggingface.co/datasets/0xmoose/stealthbench Website: https://stealthbench.com

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.26314
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.26314 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.26314 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers