Hugging Face Daily Papers · · 3 min read

MOLE: Detecting Insider Threats in AI Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Dataset: <a href=\"https://huggingface.co/datasets/forgelab/mole\">https://huggingface.co/datasets/forgelab/mole</a><br>Code: <a href=\"https://github.com/aashiqmuhamed/mole\" rel=\"nofollow\">https://github.com/aashiqmuhamed/mole</a></p>\n","updatedAt":"2026-09-09T02:30:06.934Z","author":{"_id":"64755a83e0b188d3cb2579d8","avatarUrl":"/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg","fullname":"Aashiq Muhamed","name":"aashiqmuhamed","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":2,"identifiedLanguage":{"language":"en","probability":0.8629922866821289},"editors":["aashiqmuhamed"],"editorAvatarUrls":["/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.06966","authors":[{"_id":"6aa0c3f9d0174964227bec37","name":"Aashiq Muhamed","hidden":false},{"_id":"6aa0c3f9d0174964227bec38","name":"Virginia Smith","hidden":false}],"publishedAt":"2026-09-07T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"MOLE: Detecting Insider Threats in AI Agents","submittedOnDailyBy":{"_id":"64755a83e0b188d3cb2579d8","avatarUrl":"/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg","isPro":false,"fullname":"Aashiq Muhamed","user":"aashiqmuhamed","type":"user","name":"aashiqmuhamed"},"summary":"Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.","upvotes":15,"discussionId":"6aa0c3f9d0174964227bec39","githubRepo":"https://github.com/aashiqmuhamed/mole","githubRepoAddedBy":"user","ai_summary":"MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating shared services under limited review budgets.","ai_keywords":["AI agents","model weights","training data","release gates","MOLE benchmark","stateful services","agent refusal","monitors","audit-event","observability","budget-AUC"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64755a83e0b188d3cb2579d8","avatarUrl":"/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg","isPro":false,"fullname":"Aashiq Muhamed","user":"aashiqmuhamed","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6a6aa2b12af00ac1140c2436","avatarUrl":"/avatars/8e4e76a86e39aee88928d7cc7b8537b6.svg","isPro":false,"fullname":"Mary Thomas","user":"Rapid-Wisp","type":"user"},{"_id":"6a6d57a4c1e23c5ff69d1331","avatarUrl":"/avatars/5b3f532a7739b93eb6a2e2c69ec7aa03.svg","isPro":false,"fullname":"William Thompson","user":"ZenithMind","type":"user"},{"_id":"6a6de7bea5a4538841bb00b6","avatarUrl":"/avatars/d504479bbb3b8acdb90b21a090ba1848.svg","isPro":false,"fullname":"Joshua Lee","user":"frostmind","type":"user"},{"_id":"6a8261711e7731b6201bc8b2","avatarUrl":"/avatars/c75f3d4a27976f941770230ed5428751.svg","isPro":false,"fullname":"Tianyu Zhang","user":"lunarbridge","type":"user"},{"_id":"6a9ae09013e59faad1699caa","avatarUrl":"/avatars/52e82c0e37c9e621efa8e6863bcc2ee7.svg","isPro":false,"fullname":"John Foster","user":"Harbor-Xmiller","type":"user"},{"_id":"6aa0e364c5fce68381d3baeb","avatarUrl":"/avatars/7b71d1c8f7b17b536135150782bce8ad.svg","isPro":false,"fullname":"Yvonne Turner MD","user":"kruegerlaura35","type":"user"},{"_id":"6aa0e832bee2dcd2ebf56e84","avatarUrl":"/avatars/1f196e6d24f77798afb06e0eee41bc77.svg","isPro":false,"fullname":"Lori Howell","user":"hayleybrown","type":"user"},{"_id":"6a9ece96bb45c77270327fb2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a9ece96bb45c77270327fb2/jllrZ6ykXKke1RXBWQl03.png","isPro":false,"fullname":"CharlesMarshall","user":"CharlesM1969","type":"user"},{"_id":"6a6c87ef42314c5fe5775d27","avatarUrl":"/avatars/b9aaecc917ec07f5606d1984587dabc3.svg","isPro":false,"fullname":"Linda Gonzalez","user":"kestrelDawn","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"691d9a1012cc4d473e1c862f","name":"CarnegieMellonU","fullname":"Carnegie Mellon University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68e396f2b5bb631e9b2fac9a/6I146aJvxxlRCEbYFFAeQ.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.06966.md","query":{}}">
Papers
arxiv:2609.06966

MOLE: Detecting Insider Threats in AI Agents

Published on Sep 7
· Submitted by
Aashiq Muhamed
on Sep 9
Authors:
,

Abstract

MOLE is a benchmark for evaluating defenses that detect harmful actions by AI agents operating shared services under limited review budgets.

Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.06966
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.06966 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.06966 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers