Hugging Face Daily Papers · · 3 min read

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

A paper about how to use training data attribution (TDA) more effectively</p>\n","updatedAt":"2026-09-10T12:45:29.377Z","author":{"_id":"639c215c5b8c217e25ddda0b","avatarUrl":"/avatars/87b4847329309dca1a8eaec1f1a49618.svg","fullname":"Jianhui Chen","name":"JianhuiChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8845131397247314},"editors":["JianhuiChen"],"editorAvatarUrls":["/avatars/87b4847329309dca1a8eaec1f1a49618.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.02771","authors":[{"_id":"6aa2a5e1653f6b802b7e21fd","name":"Yuzhang Luo","hidden":false},{"_id":"6aa2a5e1653f6b802b7e21fe","name":"Chenpeng Wang","hidden":false},{"_id":"6aa2a5e1653f6b802b7e21ff","name":"Jianhui Chen","hidden":false},{"_id":"6aa2a5e1653f6b802b7e2200","name":"Liangming Pan","hidden":false}],"publishedAt":"2026-09-02T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution","submittedOnDailyBy":{"_id":"639c215c5b8c217e25ddda0b","avatarUrl":"/avatars/87b4847329309dca1a8eaec1f1a49618.svg","isPro":false,"fullname":"Jianhui Chen","user":"JianhuiChen","type":"user","name":"JianhuiChen"},"summary":"Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.","upvotes":2,"discussionId":"6aa2a5e1653f6b802b7e2201","ai_summary":"Influence-guided response rewriting of selected training examples produces stronger and more persistent behavioral shifts in language models than conventional reweighting, highlighting the broader intervention leverage of influential data.","ai_keywords":["training data attribution","influence functions","response rewriting","reweighting","epistemic abstention","open-weight LLMs","intervention-aware evaluation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"61c2e4b131692679706c0716","name":"PKU","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61c2e44c39245e7bf62def6f/bGOsSh93qDIlsl2XWsEi2.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"639c215c5b8c217e25ddda0b","avatarUrl":"/avatars/87b4847329309dca1a8eaec1f1a49618.svg","isPro":false,"fullname":"Jianhui Chen","user":"JianhuiChen","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"61c2e4b131692679706c0716","name":"PKU","fullname":"Peking University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/61c2e44c39245e7bf62def6f/bGOsSh93qDIlsl2XWsEi2.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.02771.md","query":{}}">
Papers
arxiv:2609.02771

From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution

Published on Sep 2
· Submitted by
Jianhui Chen
on Sep 10
Authors:
,

Abstract

Influence-guided response rewriting of selected training examples produces stronger and more persistent behavioral shifts in language models than conventional reweighting, highlighting the broader intervention leverage of influential data.

Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.

Community

Paper submitter about 4 hours ago

A paper about how to use training data attribution (TDA) more effectively

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.02771
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.02771 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.02771 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.02771 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers