Hugging Face Daily Papers · · 3 min read

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A comprehensive multi-framework guard model with cot.</p>\n","updatedAt":"2026-09-15T04:54:31.828Z","author":{"_id":"69d47558a9acd1eb26637fe9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69d47558a9acd1eb26637fe9/ZWSP5HU9uU7Bc2a1I54eq.jpeg","fullname":"YunHao-Feng","name":"Yunhao-Feng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8240867257118225},"editors":["Yunhao-Feng"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/69d47558a9acd1eb26637fe9/ZWSP5HU9uU7Bc2a1I54eq.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.15134","authors":[{"_id":"6aa8cf2e5dd4cb9b4cc0291d","name":"Yunhao Feng","hidden":false},{"_id":"6aa8cf2e5dd4cb9b4cc0291e","name":"Ruixiao Lin","hidden":false},{"_id":"6aa8cf2e5dd4cb9b4cc0291f","name":"Ming Wen","hidden":false},{"_id":"6aa8cf2e5dd4cb9b4cc02920","name":"Yanming Guo","hidden":false},{"_id":"6aa8cf2e5dd4cb9b4cc02921","name":"Xingjun Ma","hidden":false},{"_id":"6aa8cf2e5dd4cb9b4cc02922","name":"Yutao Wu","hidden":false},{"_id":"6aa8cf2e5dd4cb9b4cc02923","name":"Xinhao Deng","hidden":false},{"_id":"6aa8cf2e5dd4cb9b4cc02924","name":"Shouling Ji","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"HazardAuditor: From Executable Threats to Safer Computer-Use Agents","submittedOnDailyBy":{"_id":"69d47558a9acd1eb26637fe9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/69d47558a9acd1eb26637fe9/ZWSP5HU9uU7Bc2a1I54eq.jpeg","isPro":false,"fullname":"YunHao-Feng","user":"Yunhao-Feng","type":"user","name":"Yunhao-Feng"},"summary":"Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across heterogeneous agent frameworks. We introduce HazardAuditor, an execution-grounded framework that closes both gaps. Its infrastructure runs heterogeneous agents (Claude Code, Codex, Hermes, and OpenClaw) in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision. We further observe that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates. Guard Policy Optimization (GuardPO) addresses this by converting deterministic safety outcomes into sequence-level advantages and normalizing rationale and verdict regions, making the safety decision the effective unit of optimization. Across multiple benchmarks and heterogeneous computer-use systems, HazardAuditor improves accuracy by up to 16.5 percentage points over the strongest prior guard. Code, models, and evaluation artifacts will be available at https://yunhao-feng.github.io/HazardAuditor/.","upvotes":10,"discussionId":"6aa8cf2e5dd4cb9b4cc02925","projectPage":"https://yunhao-feng.github.io/HazardAuditor/","githubRepo":"https://github.com/Yunhao-Feng/HazardAuditor","githubRepoAddedBy":"user","ai_summary":"HazardAuditor provides execution-grounded safety supervision for computer-use agents and introduces Guard Policy Optimization to align generative guard training with sequence-level safety outcomes.","ai_keywords":["Guard Policy Optimization","generative guards","token-level post-training","sequence-level advantages"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3,"organization":{"_id":"67c1d682826160b28f778510","name":"antgroup","fullname":"Ant Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/7VcPHdLSGlged3ixK1dys.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66b87455d36b8cb3566b667e","avatarUrl":"/avatars/d39917353c28be5a3950c7246238e2de.svg","isPro":false,"fullname":"Alex","user":"Alctrain","type":"user"},{"_id":"6a150bfdc53df1ab1f937710","avatarUrl":"/avatars/f16eed04536a0136a2ca16eef75fc4eb.svg","isPro":false,"fullname":"W","user":"HiccupRL","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"66c6be8e13962d19a84b949a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66c6be8e13962d19a84b949a/INZETOmZoruCZUmTgVI6a.png","isPro":false,"fullname":"Repoan","user":"Repoaner","type":"user"},{"_id":"666193216c2ebb19b0748003","avatarUrl":"/avatars/7687c4b0c75b8fdfb608228500ed5b74.svg","isPro":false,"fullname":"w l","user":"anesyl","type":"user"},{"_id":"634bde123d11eaedd889e277","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1665916392312-noauth.png","isPro":false,"fullname":"Hengyuan Xu","user":"DobyXu","type":"user"},{"_id":"624540022fccf60b62ce474b","avatarUrl":"/avatars/9946463861761edd7c21022721ab61bb.svg","isPro":false,"fullname":"yangyajie","user":"aoooa","type":"user"},{"_id":"66792c6dec54ee1558c3a6ad","avatarUrl":"/avatars/ecfb479c9f4336f5b437ea1a4d1fbf2a.svg","isPro":false,"fullname":"SII-mingwen","user":"SII-fleeeecer","type":"user"},{"_id":"655101623fe6c0b1f8b58987","avatarUrl":"/avatars/4d36a4988e6011fec3ceac2b59938c3a.svg","isPro":false,"fullname":"Jiabin Hua","user":"Ammmob","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67c1d682826160b28f778510","name":"antgroup","fullname":"Ant Group","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/7VcPHdLSGlged3ixK1dys.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.15134.md","query":{}}">
Papers
arxiv:2609.15134

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Published on Sep 14
· Submitted by
YunHao-Feng
on Sep 15
Authors:
,

Abstract

HazardAuditor provides execution-grounded safety supervision for computer-use agents and introduces Guard Policy Optimization to align generative guard training with sequence-level safety outcomes.

Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across heterogeneous agent frameworks. We introduce HazardAuditor, an execution-grounded framework that closes both gaps. Its infrastructure runs heterogeneous agents (Claude Code, Codex, Hermes, and OpenClaw) in controlled environments and normalizes their interactions into a canonical event representation for cross-framework supervision. We further observe that token-level post-training objectives create a structural mismatch for generative guards, causing longer rationales to dominate gradient updates. Guard Policy Optimization (GuardPO) addresses this by converting deterministic safety outcomes into sequence-level advantages and normalizing rationale and verdict regions, making the safety decision the effective unit of optimization. Across multiple benchmarks and heterogeneous computer-use systems, HazardAuditor improves accuracy by up to 16.5 percentage points over the strongest prior guard. Code, models, and evaluation artifacts will be available at https://yunhao-feng.github.io/HazardAuditor/.

Community

Paper submitter about 3 hours ago

A comprehensive multi-framework guard model with cot.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.15134
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.15134 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.15134 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers