Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.</p>\n","updatedAt":"2026-08-06T02:29:46.960Z","author":{"_id":"656c5b5cfa91c816094cecaf","avatarUrl":"/avatars/25e58c53bc2a14f05307023d45129246.svg","fullname":"Yushi SUN","name":"Yushi98","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9247632622718811},"editors":["Yushi98"],"editorAvatarUrls":["/avatars/25e58c53bc2a14f05307023d45129246.svg"],"reactions":[],"isReport":false}},{"id":"6a7450259f2214ec691ee89c","author":{"_id":"6a744fbf673d3ab6e78d82cc","avatarUrl":"/avatars/608d52b609f9356a286d1878afb20777.svg","fullname":"amee huyen","name":"amee094huyen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-06T09:13:09.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This tackles an important and underexplored problem in personalized LLMs. The concept of over-inference is well defined, and MirageBench looks like a thoughtful benchmark with strong validation. I especially found the Self-Monitoring Inversion result compelling; it's a valuable reminder that models' self-assessments shouldn't be relied upon as the primary measure of faithfulness.","html":"<p>This tackles an important and underexplored problem in personalized LLMs. The concept of over-inference is well defined, and MirageBench looks like a thoughtful benchmark with strong validation. I especially found the Self-Monitoring Inversion result compelling; it's a valuable reminder that models' self-assessments shouldn't be relied upon as the primary measure of faithfulness.</p>\n","updatedAt":"2026-08-06T09:13:09.427Z","author":{"_id":"6a744fbf673d3ab6e78d82cc","avatarUrl":"/avatars/608d52b609f9356a286d1878afb20777.svg","fullname":"amee huyen","name":"amee094huyen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9534780979156494},"editors":["amee094huyen"],"editorAvatarUrls":["/avatars/608d52b609f9356a286d1878afb20777.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04570","authors":[{"_id":"6a73f18cc5e410d076869a8f","name":"Yushi Sun","hidden":false},{"_id":"6a73f18cc5e410d076869a90","name":"Yanjie Zhang","hidden":false},{"_id":"6a73f18cc5e410d076869a91","name":"Rui Sheng","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads","submittedOnDailyBy":{"_id":"656c5b5cfa91c816094cecaf","avatarUrl":"/avatars/25e58c53bc2a14f05307023d45129246.svg","isPro":false,"fullname":"Yushi SUN","user":"Yushi98","type":"user","name":"Yushi98"},"summary":"Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.","upvotes":35,"discussionId":"6a73f18dc5e410d076869a92","organization":{"_id":"63355133edc1a61aecf74b0e","name":"HKUST","fullname":"HKUST","avatar":"https://www.gravatar.com/avatar/4a4318de793d2c187cb6f312e9d0e7bc?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"656c5b5cfa91c816094cecaf","avatarUrl":"/avatars/25e58c53bc2a14f05307023d45129246.svg","isPro":false,"fullname":"Yushi SUN","user":"Yushi98","type":"user"},{"_id":"64bce857796f20daad639795","avatarUrl":"/avatars/36d46ab089bcac562f98fbd0448895db.svg","isPro":false,"fullname":"Dylan","user":"Dylannnnnnnn","type":"user"},{"_id":"65f2b36733725f138b252311","avatarUrl":"/avatars/69897dd235c24a961c3ebf4e961a9dc0.svg","isPro":false,"fullname":"Patricia","user":"Patriciayufish","type":"user"},{"_id":"656084f44e8918182d4f07c8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/akAvCUCi7eR31PWOXrVPw.jpeg","isPro":false,"fullname":"Yihao Meng","user":"Yhmeng1106","type":"user"},{"_id":"67c4d931589a728ad1167d46","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/ljyz7EEfVmetZvCEOjLAS.png","isPro":false,"fullname":"vkvvkv","user":"vvvvvkkk","type":"user"},{"_id":"671b3adbf06d2ebd6d6e4647","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/KLzk0oWz5LK9PauRe3WNi.png","isPro":false,"fullname":"zack zheng","user":"zack4417","type":"user"},{"_id":"662374fcdd416a585a26faff","avatarUrl":"/avatars/62ad033fea6e8db0df650e08f500ea77.svg","isPro":false,"fullname":"Haobo Li","user":"Haobo55654","type":"user"},{"_id":"69f9156d6b3046f53da0f277","avatarUrl":"/avatars/f99df7c24c2f1a9c9ac63b542c6aa323.svg","isPro":false,"fullname":"Anonymous","user":"nips2026-chemcost","type":"user"},{"_id":"68ba57fb47b19a6313be1556","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/GuCIT4NvLFyA5iQHZ-YiY.png","isPro":false,"fullname":"Xinli Zhu","user":"XinliZhu","type":"user"},{"_id":"677e8b330ef2084985c0a4f5","avatarUrl":"/avatars/d34bbc7cdd9a006c73c85e85ad53429a.svg","isPro":false,"fullname":"Yanjie Zhang","user":"doudouwer","type":"user"},{"_id":"68ec8199b6bd40ad442306eb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/4PVVSx4p9ZsgERq4UHoET.png","isPro":false,"fullname":"XU","user":"zijain","type":"user"},{"_id":"68c8d5d97c55656282d98edd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68c8d5d97c55656282d98edd/ELxXzHVwi2kYGcWHvwvf1.jpeg","isPro":false,"fullname":"LI YAFEI","user":"unnalin","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"63355133edc1a61aecf74b0e","name":"HKUST","fullname":"HKUST","avatar":"https://www.gravatar.com/avatar/4a4318de793d2c187cb6f312e9d0e7bc?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04570.md","query":{}}">
The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
Abstract
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.
Community
Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spanning an ``imagination gradient'', a four-way faithfulness taxonomy operationalized by an independent judge (validated against a blind human annotator on 400 claims: Cohen's kappa = 0.863 four-class, kappa = 0.900 binary), and a leaderboard of 12 models across 7 families on 143616 judged claims. We find that over-inference is pervasive: every one of the 12 models over-infers 35%--49% of its claims (cross-model mean 41.6%; claim-weighted 41.8%), with no model in this evaluation escaping it. Most strikingly, we surface a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with their judge-measured OI (rho = -0.60, p = 0.044; exploratory, wide bootstrap CI [-0.90, +0.06], n = 12). The models that report the least over-inference tend to be flagged as fabricating the most, so self-reported confidence is a misleading signal for comparing models, even though within a single model self-audit still ranks that model's own claims moderately well (AUROC 0.58--0.83). We further show that OI is task-dependent (27%--59%) and that, in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. MirageBench positions external verification, rather than model self-report, as a more reliable foundation for trustworthy personalization.
This tackles an important and underexplored problem in personalized LLMs. The concept of over-inference is well defined, and MirageBench looks like a thoughtful benchmark with strong validation. I especially found the Self-Monitoring Inversion result compelling; it's a valuable reminder that models' self-assessments shouldn't be relied upon as the primary measure of faithfulness.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.04570 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.04570 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.04570 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.