We propose XConf, a new paradigm for confidence estimation that is training-free, black-box, and needs only one answer generation, working from reasoning to agents.</p>\n","updatedAt":"2026-09-17T06:17:25.278Z","author":{"_id":"63920dfac47e36ddeb8f1864","avatarUrl":"/avatars/c36cbf7b084d62368312e5c9292e4260.svg","fullname":"Caiqi Zhang","name":"caiqizh","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9385204315185547},"editors":["caiqizh"],"editorAvatarUrls":["/avatars/c36cbf7b084d62368312e5c9292e4260.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.17708","authors":[{"_id":"6aab84f51d9cc4dec796258e","name":"Caiqi Zhang","hidden":false},{"_id":"6aab84f51d9cc4dec796258f","name":"Xiaochen Zhu","hidden":false},{"_id":"6aab84f51d9cc4dec7962590","name":"Chengzu Li","hidden":false},{"_id":"6aab84f51d9cc4dec7962591","name":"Yulong Chen","hidden":false},{"_id":"6aab84f51d9cc4dec7962592","name":"Dharshan Kumaran","hidden":false},{"_id":"6aab84f51d9cc4dec7962593","name":"Nigel Collier","hidden":false}],"publishedAt":"2026-09-15T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents","submittedOnDailyBy":{"_id":"63920dfac47e36ddeb8f1864","avatarUrl":"/avatars/c36cbf7b084d62368312e5c9292e4260.svg","isPro":false,"fullname":"Caiqi Zhang","user":"caiqizh","type":"user","name":"caiqizh"},"summary":"Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators, however, share one design premise: they only read the current inference process, either by introspecting on it, scoring its token probabilities, or resampling it. We argue that the current inference is not a sufficient basis for confidence. We propose XConf (eXperiential Confidence): estimating confidence together with the model's accumulated experience. The experience is stored as a record of the model's own graded past episodes, each holding the task, the model's reflection, its stated confidence, the outcome, and a lesson written once the grade arrived. Given a new task, XConf's Recall stage retrieves past episodes on similar tasks met with a similar stated confidence, and reads off their historical success rate; its Reflect stage shows the model this record, has it name its recurring failure mode, and restate a confidence now informed by its own track records. Our estimator is format-general, requiring no logit access or weight updates, and costs only one answer generation. Across nine benchmarks spanning reasoning, coding, multimodal QA, and interactive agents, and four models from three families, XConf beats or matches ten-sample self-consistency in discrimination (AUROC) on 23 of 24 comparisons, with much lower calibration error (ECE), at a tenth of the generation cost. Used for selective prediction, abstaining on the 10% least-confident episodes raises the delivered success rate by up to 8.7 points on agent tasks. We therefore see experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.","upvotes":46,"discussionId":"6aab84f61d9cc4dec7962594","projectPage":"https://caiqizh.github.io/xconf/","githubRepo":"https://github.com/caiqizh/xconf","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"65f9e02087d1c912d985eebf","name":"CambUni","fullname":"University of Cambridge","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/gdCou-0KoYsLPYqPrz6xn.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63920dfac47e36ddeb8f1864","avatarUrl":"/avatars/c36cbf7b084d62368312e5c9292e4260.svg","isPro":false,"fullname":"Caiqi Zhang","user":"caiqizh","type":"user"},{"_id":"6064a0eeb1703ddba0d458b9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1617207525789-noauth.png","isPro":false,"fullname":"Qiushi","user":"QiushiSun","type":"user"},{"_id":"64fed23f0871bc5930598ab5","avatarUrl":"/avatars/080a4ef3e4634cd978528dfa899a4eb0.svg","isPro":false,"fullname":"ZhiWei LI","user":"Aragonaa","type":"user"},{"_id":"6675344c20ff491d381a3f1e","avatarUrl":"/avatars/e3ebfce52d4029aa241f40051fcc1e0c.svg","isPro":false,"fullname":"Henry Zhang","user":"Henry0709","type":"user"},{"_id":"66276727368ec2a0b933772c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66276727368ec2a0b933772c/ff5kAJuUr8tVlbNP1If_Y.jpeg","isPro":false,"fullname":"ccloud0525","user":"Ccloud0525","type":"user"},{"_id":"61669c456916c52acd5a1aa3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/61669c456916c52acd5a1aa3/HnZTwRaXgTeTG3ljO3ITb.jpeg","isPro":false,"fullname":"jianbo dai","user":"jbd","type":"user"},{"_id":"67f8ccce9301e8cd1592b71f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/pmQUnMY4GqTYmm0K-_7BA.png","isPro":false,"fullname":"WangZilin","user":"terr1ble","type":"user"},{"_id":"658247c592b5a9664de63882","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658247c592b5a9664de63882/Jn03voLQjDlB3YjpSi-PI.jpeg","isPro":true,"fullname":"fan","user":"EasonFan","type":"user"},{"_id":"643a587fe2b979ae6141b193","avatarUrl":"/avatars/1726b6a1629d800795f9bdf6d03ad190.svg","isPro":false,"fullname":"yilong xu","user":"sapphirex","type":"user"},{"_id":"67bc657cb24be20e87c174cd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/XKm5OMx1r2zGF0hrR7yh7.png","isPro":false,"fullname":"zoe zoe","user":"zoezou2015","type":"user"},{"_id":"67b24be6a8f5fdc2fae14ff1","avatarUrl":"/avatars/872eaae5664047de4e3dc7386e5ee79d.svg","isPro":false,"fullname":"Demi Ruohan Wang","user":"demisama","type":"user"},{"_id":"6971d49935dee2291c1611da","avatarUrl":"/avatars/160ddffbd82f7d9839af143649c218ac.svg","isPro":false,"fullname":"Yucheng Ning","user":"ningyucheng","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65f9e02087d1c912d985eebf","name":"CambUni","fullname":"University of Cambridge","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/gdCou-0KoYsLPYqPrz6xn.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.17708.md","query":{}}">
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
Abstract
Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators, however, share one design premise: they only read the current inference process, either by introspecting on it, scoring its token probabilities, or resampling it. We argue that the current inference is not a sufficient basis for confidence. We propose XConf (eXperiential Confidence): estimating confidence together with the model's accumulated experience. The experience is stored as a record of the model's own graded past episodes, each holding the task, the model's reflection, its stated confidence, the outcome, and a lesson written once the grade arrived. Given a new task, XConf's Recall stage retrieves past episodes on similar tasks met with a similar stated confidence, and reads off their historical success rate; its Reflect stage shows the model this record, has it name its recurring failure mode, and restate a confidence now informed by its own track records. Our estimator is format-general, requiring no logit access or weight updates, and costs only one answer generation. Across nine benchmarks spanning reasoning, coding, multimodal QA, and interactive agents, and four models from three families, XConf beats or matches ten-sample self-consistency in discrimination (AUROC) on 23 of 24 comparisons, with much lower calibration error (ECE), at a tenth of the generation cost. Used for selective prediction, abstaining on the 10% least-confident episodes raises the delivered success rate by up to 8.7 points on agent tasks. We therefore see experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.
Community
We propose XConf, a new paradigm for confidence estimation that is training-free, black-box, and needs only one answer generation, working from reasoning to agents.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.17708 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.17708 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.17708 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.