<a href=\"https://cdn-uploads.huggingface.co/production/uploads/64755a83e0b188d3cb2579d8/ZfcM80w4MLUeBlCdKvG4d.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/64755a83e0b188d3cb2579d8/ZfcM80w4MLUeBlCdKvG4d.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-09-15T04:22:38.258Z","author":{"_id":"64755a83e0b188d3cb2579d8","avatarUrl":"/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg","fullname":"Aashiq Muhamed","name":"aashiqmuhamed","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.4681166112422943},"editors":["aashiqmuhamed"],"editorAvatarUrls":["/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.15029","authors":[{"_id":"6aa8c5ea5dd4cb9b4cc028ef","name":"Aashiq Muhamed","hidden":false},{"_id":"6aa8c5ea5dd4cb9b4cc028f0","name":"Mona T. Diab","hidden":false},{"_id":"6aa8c5ea5dd4cb9b4cc028f1","name":"Virginia Smith","hidden":false},{"_id":"6aa8c5ea5dd4cb9b4cc028f2","name":"Andrew Ilyas","hidden":false},{"_id":"6aa8c5ea5dd4cb9b4cc028f3","name":"Matthew Jagielski","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks","submittedOnDailyBy":{"_id":"64755a83e0b188d3cb2579d8","avatarUrl":"/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg","isPro":false,"fullname":"Aashiq Muhamed","user":"aashiqmuhamed","type":"user","name":"aashiqmuhamed"},"summary":"Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen.\n We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.","upvotes":3,"discussionId":"6aa8c5ea5dd4cb9b4cc028f4","githubRepo":"https://github.com/aashiqmuhamed/poison-set-selection","githubRepoAddedBy":"user","ai_summary":"Backdoor vulnerability in fine-tuned language models varies drastically with poison selection, and a learned set-scoring method improves worst-case attack success by identifying high-impact poisoned examples.","ai_keywords":["backdoor poisoning attacks","finetuning","LLaMA-3-8B","oracle-budgeted set optimization","SAILS","set scorer","influence baselines","code-generation","agentic backdoors"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"63924ab637e424786530c90e","name":"Anthropic","fullname":"Anthropic","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1670531762351-6200d0a443eb0913fa2df7cc.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64755a83e0b188d3cb2579d8","avatarUrl":"/avatars/2c50590905f4bd398a4c9991e1b4b5bb.svg","isPro":false,"fullname":"Aashiq Muhamed","user":"aashiqmuhamed","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63924ab637e424786530c90e","name":"Anthropic","fullname":"Anthropic","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1670531762351-6200d0a443eb0913fa2df7cc.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.15029.md","query":{}}">
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
Abstract
Backdoor vulnerability in fine-tuned language models varies drastically with poison selection, and a learned set-scoring method improves worst-case attack success by identifying high-impact poisoned examples.
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen.
We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.15029 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.15029 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.15029 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.