AI Security Leaderboard: benchmarking model robustness [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| We developed a leaderboard ranking frontier model security. There's no shortage of model capability rankings, but we didn't find anything comparable for model security. Yet security is becoming increasingly critical to deployment decisions: from the USG making developers pull models for cybersecurity jailbreaks to developers holding back on AI agent deployments due to risks of adversarial attacks. We developed an automated test suite that runs models through 1500 automatically generated jailbreak attempts and measures the number of universal jailbreaks: prompts that elicit compliant, detailed responses to >75% clearly harmful questions within a domain (like offensive cybersecurity). We find a big gap between the most and least robust models in our technical report. This is v1.0 and we'd really appreciate input from this subreddit on next steps, as well as feedback on the metholodogy. Areas we're considering include: We'd also love to hear ways we could make this benchmark more useful in your work. If you're an adversarial robustness researcher, are there artifacts such as datasets or evaluation rubrics you'd like to re-use? [link] [comments] |
More from r/MachineLearning
-
A collision-entropy floor for watermark/retrieval AI-text detection. Looking for a sanity check before I take this further [D]
Aug 14
-
Are supervised and unsupervised learning still relevant today? [D]
Aug 14
-
TMLR Relevance and Prestige [D]
Aug 13
-
Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D]
Aug 13
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.