A policy-adaptive guardrail model training and evaluation recipe</p>\n<p>data & models: <a href=\"https://huggingface.co/PolicyShiftGuard\">https://huggingface.co/PolicyShiftGuard</a></p>\n","updatedAt":"2026-07-16T02:58:12.644Z","author":{"_id":"66aca01e33f6b27979856f6f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg","fullname":"Mingyang Song","name":"hitsmy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6880612969398499},"editors":["hitsmy"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.05910","authors":[{"_id":"6a55adb8a9d74d6e65bbd33a","name":"Mingyang Song","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33b","name":"Luxin Xu","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33c","name":"Haoyu Sun","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33d","name":"Minzhou Pan","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33e","name":"Yu Cheng","hidden":false},{"_id":"6a55adb8a9d74d6e65bbd33f","name":"Bo Li","hidden":false}],"publishedAt":"2026-07-07T07:04:28.000Z","submittedOnDailyAt":"2026-07-16T00:00:00.000Z","title":"PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails","submittedOnDailyBy":{"_id":"66aca01e33f6b27979856f6f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg","isPro":false,"fullname":"Mingyang Song","user":"hitsmy","type":"user","name":"hitsmy"},"summary":"Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed when a policy boundary changes. We study policy-adaptive image guardrailing, where a model must decide whether an image violates the currently supplied policy and generalize to held-out policy definitions. We introduce PolicyShiftBench, a comprehensive benchmark with 2,000 policy-discriminative instances over 265 images, where each image is paired with 7.55 policy-conditioned prompts on average to test whether models adapt to the active policy rather than relying on image-level safety priors. We then propose PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt trains matched prompts for the same image and risk category using standard label supervision and a pairwise comparison loss that separates blocking policies from passing policies. Experiments show that existing VLMs and specialized guardrails remain brittle under policy shifts, while PolicyShiftGuard substantially improves policy-sensitive performance. The 7B model achieves SOTA performance of 76.9 Avg. F1 and 72.1 Avg. PSS on PolicyShiftBench, transfers well to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off with a concise output format. Ablations confirm that matched pass/block boundary pairs are essential for stable policy adaptation.","upvotes":30,"discussionId":"6a55adb8a9d74d6e65bbd340","projectPage":"https://policyshiftguard.github.io/","githubRepo":"https://github.com/ssmisya/PolicyShiftGuard","githubRepoAddedBy":"user","githubStars":20,"organization":{"_id":"643cb0625fcffe09fb6ca688","name":"Fudan-University","fullname":"Fudan University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6437eca0819f3ab20d162e14/kWv0cGlAhAG3iNWVxowkJ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66aca01e33f6b27979856f6f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66aca01e33f6b27979856f6f/IyOxv89TudwscGH7tdue3.jpeg","isPro":false,"fullname":"Mingyang Song","user":"hitsmy","type":"user"},{"_id":"6418228b83957c4eaaad4d01","avatarUrl":"/avatars/b6af01d09bba5d5ba7bf4a62914ca468.svg","isPro":false,"fullname":"wang","user":"astrid01052","type":"user"},{"_id":"65697feb9fb2d79a79e14e0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65697feb9fb2d79a79e14e0a/wVGaBjn8pQIJneZWSFIwS.jpeg","isPro":false,"fullname":"haodi lei","user":"bingyang-lei","type":"user"},{"_id":"64ba47b129d10d4185c46af1","avatarUrl":"/avatars/84a776d283b01f0558a28a5625115f83.svg","isPro":false,"fullname":"Zhilin Wang","user":"linzw","type":"user"},{"_id":"6785e04ce63a669874f166e9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6785e04ce63a669874f166e9/xSUEAmRm-KadDNOYiOZew.jpeg","isPro":false,"fullname":"dodojorid","user":"yizhuoli","type":"user"},{"_id":"67763550746c53a0fad415ba","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/l09xhBMmqzX7SR7Sf_v0W.png","isPro":false,"fullname":"Rong-Xi Tan","user":"trxcc2002","type":"user"},{"_id":"62495cb96ee7ee6b646db130","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62495cb96ee7ee6b646db130/UwBXmvcMq7LMvBWUw0xo3.jpeg","isPro":false,"fullname":"Runzhe Zhan","user":"rzzhan","type":"user"},{"_id":"65352acb7139c5dd8d9a8590","avatarUrl":"/avatars/e2ff22b596aee45cdfb8f68dc15572f9.svg","isPro":false,"fullname":"JiachengChen","user":"JC-Chen","type":"user"},{"_id":"6565d17cd4ef4fe85fb44794","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/W1NAEzEEyzRNp6I2fGkcH.jpeg","isPro":false,"fullname":"WangWenxuan","user":"Yummytanmo","type":"user"},{"_id":"645b4819f9d4ec91fdd54852","avatarUrl":"/avatars/e12efb8e030688a0afcc72176b453fb3.svg","isPro":false,"fullname":"Jiawei Gu","user":"kuvvi","type":"user"},{"_id":"63f3502a520c14618925825a","avatarUrl":"/avatars/e986a2a6625e7be6890616a417f908d2.svg","isPro":false,"fullname":"Yafu Li","user":"yaful","type":"user"},{"_id":"6355473d525beaee688b7ba1","avatarUrl":"/avatars/1fb0d57ed5f1a9b872a1ada8b2973ffb.svg","isPro":false,"fullname":"Wei Tao","user":"itaowe","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"643cb0625fcffe09fb6ca688","name":"Fudan-University","fullname":"Fudan University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6437eca0819f3ab20d162e14/kWv0cGlAhAG3iNWVxowkJ.png"},"query":{}}">
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Abstract
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed when a policy boundary changes. We study policy-adaptive image guardrailing, where a model must decide whether an image violates the currently supplied policy and generalize to held-out policy definitions. We introduce PolicyShiftBench, a comprehensive benchmark with 2,000 policy-discriminative instances over 265 images, where each image is paired with 7.55 policy-conditioned prompts on average to test whether models adapt to the active policy rather than relying on image-level safety priors. We then propose PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt trains matched prompts for the same image and risk category using standard label supervision and a pairwise comparison loss that separates blocking policies from passing policies. Experiments show that existing VLMs and specialized guardrails remain brittle under policy shifts, while PolicyShiftGuard substantially improves policy-sensitive performance. The 7B model achieves SOTA performance of 76.9 Avg. F1 and 72.1 Avg. PSS on PolicyShiftBench, transfers well to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off with a concise output format. Ablations confirm that matched pass/block boundary pairs are essential for stable policy adaptation.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.05910 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.05910 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.05910 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.