Hugging Face Daily Papers · · 7 min read

What AI Red-Team Evaluations Can and Cannot Prove

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

The input space for AI models is infinite, but labs can't just test indefinitely. There's a calculable limit to how much a red-team benchmark or safety evaluation can actually tell you, based on your testing budget and the paper works out exactly where that limit sits. Above a certain harm-rate threshold, a modest-sized benchmark can genuinely certify a model as safe in a rigorous sense (and a clean pass becomes stronger evidence than a single failed case is weak evidence). Below that threshold (like rare, catastrophic failure modes) no passive benchmark of realistic size can give you the confidence you'd want, no matter how well it's designed.<br>We apply this to 8 real evaluation suites and find current benchmarks hold up fine for common harms but fall short by orders of magnitude for rare/catastrophic ones.</p>\n","updatedAt":"2026-08-06T21:20:50.891Z","author":{"_id":"6a679020b19e5a918ba72c43","avatarUrl":"/avatars/e2c64c6222e4430c4ef149a2df743538.svg","fullname":"Bandana","name":"hackwither","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9398754835128784},"editors":["hackwither"],"editorAvatarUrls":["/avatars/e2c64c6222e4430c4ef149a2df743538.svg"],"reactions":[],"isReport":false}},{"id":"6a753a5f4aee73813585019e","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-07T01:52:31.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks](https://huggingface.co/papers/2607.28685) (2026)\n* [When benchmark inferences do not compose: Projectibility in AI evaluation](https://huggingface.co/papers/2607.26159) (2026)\n* [When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design](https://huggingface.co/papers/2608.01378) (2026)\n* [Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control](https://huggingface.co/papers/2607.01153) (2026)\n* [Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric](https://huggingface.co/papers/2607.12469) (2026)\n* [Item Response Theory for AI Safety](https://huggingface.co/papers/2608.05086) (2026)\n* [Auditing AI Investment Recommendations as Executable Actions](https://huggingface.co/papers/2606.27570) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.28685\">Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.26159\">When benchmark inferences do not compose: Projectibility in AI evaluation</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.01378\">When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.01153\">Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.12469\">Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.05086\">Item Response Theory for AI Safety</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.27570\">Auditing AI Investment Recommendations as Executable Actions</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-07T01:52:31.238Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.760615348815918},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.21735","authors":[{"_id":"6a67968a4534e38b50e9b198","user":{"_id":"6a679020b19e5a918ba72c43","avatarUrl":"/avatars/e2c64c6222e4430c4ef149a2df743538.svg","isPro":false,"fullname":"Bandana","user":"hackwither","type":"user","name":"hackwither"},"name":"Bandana Kaur","status":"claimed_verified","statusLastChangedAt":"2026-07-28T08:57:52.274Z","hidden":false}],"publishedAt":"2026-07-23T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"What AI Red-Team Evaluations Can and Cannot Prove","submittedOnDailyBy":{"_id":"6a679020b19e5a918ba72c43","avatarUrl":"/avatars/e2c64c6222e4430c4ef149a2df743538.svg","isPro":false,"fullname":"Bandana","user":"hackwither","type":"user","name":"hackwither"},"summary":"Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.","upvotes":0,"discussionId":"6a67968a4534e38b50e9b199","githubRepo":"https://github.com/hackwither/ai-redteam-evidential-limits","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"65643a8ec683336b7604f190","name":"apisec","fullname":"apisec","avatar":"https://www.gravatar.com/avatar/65310b1af081de0136791cb53bf95335?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"organization":{"_id":"65643a8ec683336b7604f190","name":"apisec","fullname":"apisec","avatar":"https://www.gravatar.com/avatar/65310b1af081de0136791cb53bf95335?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.21735.md","query":{}}">
Papers
arxiv:2607.21735

What AI Red-Team Evaluations Can and Cannot Prove

Published on Jul 23
· Submitted by
Bandana
on Aug 6
Authors:

Abstract

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and use it to locate that boundary exactly. We find that above a calculable harm rate, a benchmark of modest size certifies a category to a stated evidentiary standard, and a clean sheet is then the stronger of the two possible observations, outweighing a single reproduced failure. Below that rate, no passive benchmark of feasible size provides the specified evidence of safety under the fixed scoring rule and approximately independent trial structure. The crossing between the two regimes has a closed form. The bound is not specific to benchmarks: written in terms of a procedure's hypothesis conditioned elicitation rates, it covers adaptive and automated red teaming as well, and shows that discrimination between the hypotheses rather than attack success is what determines evidential worth. Auditing eight evaluation suites against the boundary, we find that current benchmarks are adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones. Safety benchmarks are not uninformative. They are informative about a specific and computable set of propositions, and the discipline they need is to state which.

Community

Paper author Paper submitter about 5 hours ago

The input space for AI models is infinite, but labs can't just test indefinitely. There's a calculable limit to how much a red-team benchmark or safety evaluation can actually tell you, based on your testing budget and the paper works out exactly where that limit sits. Above a certain harm-rate threshold, a modest-sized benchmark can genuinely certify a model as safe in a rigorous sense (and a clean pass becomes stronger evidence than a single failed case is weak evidence). Below that threshold (like rare, catastrophic failure modes) no passive benchmark of realistic size can give you the confidence you'd want, no matter how well it's designed.
We apply this to 8 real evaluation suites and find current benchmarks hold up fine for common harms but fall short by orders of magnitude for rare/catastrophic ones.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.21735
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.21735 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.21735 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.21735 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers