A paper about how to use training data attribution (TDA) more effectively</p>\n","updatedAt":"2026-09-10T12:45:29.377Z","author":{"_id":"639c215c5b8c217e25ddda0b","avatarUrl":"/avatars/87b4847329309dca1a8eaec1f1a49618.svg","fullname":"Jianhui Chen","name":"JianhuiChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8845131397247314},"editors":["JianhuiChen"],"editorAvatarUrls":["/avatars/87b4847329309dca1a8eaec1f1a49618.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.02771","authors":[{"_id":"6aa2a5e1653f6b802b7e21fd","name":"Yuzhang Luo","hidden":false},{"_id":"6aa2a5e1653f6b802b7e21fe","name":"Chenpeng Wang","hidden":false},{"_id":"6aa2a5e1653f6b802b7e21ff","name":"Jianhui Chen","hidden":false},{"_id":"6aa2a5e1653f6b802b7e2200","name":"Liangming Pan","hidden":false}],"publishedAt":"2026-09-02T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution","submittedOnDailyBy":{"_id":"639c215c5b8c217e25ddda0b","avatarUrl":"/avatars/87b4847329309dca1a8eaec1f1a49618.svg","isPro":false,"fullname":"Jianhui Chen","user":"JianhuiChen","type":"user","name":"JianhuiChen"},"summary":"Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.","upvotes":2,"discussionId":"6aa2a5e1653f6b802b7e2201","ai_summary":"Influence-guided response rewriting of selected training examples produces stronger and more persistent behavioral shifts in language models than conventional reweighting, highlighting the broader intervention leverage of influential data.","ai_keywords":["training data attribution","influence functions","response rewriting","reweighting","epistemic abstention","open-weight LLMs","intervention-aware evaluation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"61c2e4b131692679706c0716","name":"PKU","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61c2e44c39245e7bf62def6f/bGOsSh93qDIlsl2XWsEi2.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"639c215c5b8c217e25ddda0b","avatarUrl":"/avatars/87b4847329309dca1a8eaec1f1a49618.svg","isPro":false,"fullname":"Jianhui Chen","user":"JianhuiChen","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61c2e4b131692679706c0716","name":"PKU","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61c2e44c39245e7bf62def6f/bGOsSh93qDIlsl2XWsEi2.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.02771.md","query":{}}">
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
Abstract
Influence-guided response rewriting of selected training examples produces stronger and more persistent behavioral shifts in language models than conventional reweighting, highlighting the broader intervention leverage of influential data.
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.
Community
A paper about how to use training data attribution (TDA) more effectively
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.02771 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.02771 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.02771 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.