Hugging Face Daily Papers · · 3 min read

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A policy-adaptive guardrail model training and evaluation recipe</p>\n<p>data &amp; models: <a href=\"https://huggingface.co/PolicyShiftGuard\">https://huggingface.co/PolicyShiftGuard</a></p>\n","updatedAt":"2026-07-16T02:58:12.644Z","author":{"_id":"66aca01e33f6b27979856f6f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg","fullname":"Mingyang Song","name":"hitsmy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6880612969398499},"editors":["hitsmy"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.05910","authors":[{"_id":"6a55adb8a9d74d6e65bbd33a","name":"Mingyang Song","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33b","name":"Luxin Xu","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33c","name":"Haoyu Sun","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33d","name":"Minzhou Pan","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33e","name":"Yu Cheng","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33f","name":"Bo Li","hidden":false}],"publishedAt":"2026-07-07T07:04:28.000Z","submittedOnDailyAt":"2026-07-16T00:00:00.000Z","title":"PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails","submittedOnDailyBy":{"_id":"66aca01e33f6b27979856f6f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg","isPro":false,"fullname":"Mingyang Song","user":"hitsmy","type":"user","name":"hitsmy"},"summary":"Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed when a policy boundary changes. We study policy-adaptive image guardrailing, where a model must decide whether an image violates the currently supplied policy and generalize to held-out policy definitions. We introduce PolicyShiftBench, a comprehensive benchmark with 2,000 policy-discriminative instances over 265 images, where each image is paired with 7.55 policy-conditioned prompts on average to test whether models adapt to the active policy rather than relying on image-level safety priors. We then propose PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt trains matched prompts for the same image and risk category using standard label supervision and a pairwise comparison loss that separates blocking policies from passing policies. Experiments show that existing VLMs and specialized guardrails remain brittle under policy shifts, while PolicyShiftGuard substantially improves policy-sensitive performance. The 7B model achieves SOTA performance of 76.9 Avg. F1 and 72.1 Avg. PSS on PolicyShiftBench, transfers well to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off with a concise output format. Ablations confirm that matched pass/block boundary pairs are essential for stable policy adaptation.","upvotes":30,"discussionId":"6a55adb8a9d74d6e65bbd340","projectPage":"https://policyshiftguard.github.io/","githubRepo":"https://github.com/ssmisya/PolicyShiftGuard","githubRepoAddedBy":"user","githubStars":20,"organization":{"_id":"643cb0625fcffe09fb6ca688","name":"Fudan-University","fullname":"Fudan University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6437eca0819f3ab20d162e14/kWv0cGlAhAG3iNWVxowkJ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66aca01e33f6b27979856f6f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg","isPro":false,"fullname":"Mingyang Song","user":"hitsmy","type":"user"},{"_id":"6418228b83957c4eaaad4d01","avatarUrl":"/avatars/b6af01d09bba5d5ba7bf4a62914ca468.svg","isPro":false,"fullname":"wang","user":"astrid01052","type":"user"},{"_id":"65697feb9fb2d79a79e14e0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65697feb9fb2d79a79e14e0a/wVGaBjn8pQIJneZWSFIwS.jpeg","isPro":false,"fullname":"haodi lei","user":"bingyang-lei","type":"user"},{"_id":"64ba47b129d10d4185c46af1","avatarUrl":"/avatars/84a776d283b01f0558a28a5625115f83.svg","isPro":false,"fullname":"Zhilin Wang","user":"linzw","type":"user"},{"_id":"6785e04ce63a669874f166e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6785e04ce63a669874f166e9/xSUEAmRm-KadDNOYiOZew.jpeg","isPro":false,"fullname":"dodojorid","user":"yizhuoli","type":"user"},{"_id":"67763550746c53a0fad415ba","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/l09xhBMmqzX7SR7Sf_v0W.png","isPro":false,"fullname":"Rong-Xi Tan","user":"trxcc2002","type":"user"},{"_id":"62495cb96ee7ee6b646db130","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62495cb96ee7ee6b646db130/UwBXmvcMq7LMvBWUw0xo3.jpeg","isPro":false,"fullname":"Runzhe Zhan","user":"rzzhan","type":"user"},{"_id":"65352acb7139c5dd8d9a8590","avatarUrl":"/avatars/e2ff22b596aee45cdfb8f68dc15572f9.svg","isPro":false,"fullname":"JiachengChen","user":"JC-Chen","type":"user"},{"_id":"6565d17cd4ef4fe85fb44794","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/W1NAEzEEyzRNp6I2fGkcH.jpeg","isPro":false,"fullname":"WangWenxuan","user":"Yummytanmo","type":"user"},{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","isPro":false,"fullname":"Jiawei Gu","user":"kuvvi","type":"user"},{"_id":"63f3502a520c14618925825a","avatarUrl":"/avatars/e986a2a6625e7be6890616a417f908d2.svg","isPro":false,"fullname":"Yafu Li","user":"yaful","type":"user"},{"_id":"6355473d525beaee688b7ba1","avatarUrl":"/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg","isPro":false,"fullname":"Wei Tao","user":"itaowe","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"643cb0625fcffe09fb6ca688","name":"Fudan-University","fullname":"Fudan University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6437eca0819f3ab20d162e14/kWv0cGlAhAG3iNWVxowkJ.png"},"query":{}}">
Papers
arxiv:2607.05910

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Published on Jul 7
· Submitted by
Mingyang Song
on Jul 16
Authors:
,

Abstract

Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed when a policy boundary changes. We study policy-adaptive image guardrailing, where a model must decide whether an image violates the currently supplied policy and generalize to held-out policy definitions. We introduce PolicyShiftBench, a comprehensive benchmark with 2,000 policy-discriminative instances over 265 images, where each image is paired with 7.55 policy-conditioned prompts on average to test whether models adapt to the active policy rather than relying on image-level safety priors. We then propose PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt trains matched prompts for the same image and risk category using standard label supervision and a pairwise comparison loss that separates blocking policies from passing policies. Experiments show that existing VLMs and specialized guardrails remain brittle under policy shifts, while PolicyShiftGuard substantially improves policy-sensitive performance. The 7B model achieves SOTA performance of 76.9 Avg. F1 and 72.1 Avg. PSS on PolicyShiftBench, transfers well to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off with a concise output format. Ablations confirm that matched pass/block boundary pairs are essential for stable policy adaptation.

Community

Paper submitter about 17 hours ago

A policy-adaptive guardrail model training and evaluation recipe

data & models: https://huggingface.co/PolicyShiftGuard

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.05910 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.05910 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.05910 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers