paper: <a href=\"https://arxiv.org/pdf/2608.05102\" rel=\"nofollow\">https://arxiv.org/pdf/2608.05102</a><br>code: <a href=\"https://github.com/PolarSeeker/ABSeeker\" rel=\"nofollow\">https://github.com/PolarSeeker/ABSeeker</a></p>\n","updatedAt":"2026-08-06T08:51:55.538Z","author":{"_id":"63ecd42a60ff4b318ad1ef47","avatarUrl":"/avatars/c05001fa9c41f713d1fe31d11212faae.svg","fullname":"RuiYe","name":"ruiye-sjtu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7634673118591309},"editors":["ruiye-sjtu"],"editorAvatarUrls":["/avatars/c05001fa9c41f713d1fe31d11212faae.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05102","authors":[{"_id":"6a740eacc5e410d076869b72","name":"Yijun Lu","hidden":false},{"_id":"6a740eacc5e410d076869b73","name":"Rui Ye","hidden":false},{"_id":"6a740eacc5e410d076869b74","name":"Jiajun Wang","hidden":false},{"_id":"6a740eacc5e410d076869b75","name":"Yuwen Du","hidden":false},{"_id":"6a740eacc5e410d076869b76","name":"Tian Jin","hidden":false},{"_id":"6a740eacc5e410d076869b77","name":"Songhua Liu","hidden":false},{"_id":"6a740eacc5e410d076869b78","name":"Siheng Chen","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment","submittedOnDailyBy":{"_id":"63ecd42a60ff4b318ad1ef47","avatarUrl":"/avatars/c05001fa9c41f713d1fe31d11212faae.svg","isPro":false,"fullname":"RuiYe","user":"ruiye-sjtu","type":"user","name":"ruiye-sjtu"},"summary":"Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions. Specifically, given a potentially obscure query and its corresponding ground-truth answer, ABC first performs Answer-Backtracked Clue Recovery, which traces back from the answer to recover intermediate clues required to solve the question. It then applies Clue-Anchored Step Scoring to evaluate each search step against these clues, converting sparse binary outcome supervision into dense step-level rewards. Based on these rewards, we develop ABC-SFT, which reweights the loss of each turn, and ABC-GRPO, which uses the step-level scores as rewards in GRPO. Building on this framework, we train ABSeeker based on Qwen3.5-4B with only 8.5k examples. ABSeeker achieves 37.3% on BrowseComp and 39.1% on BrowseComp-ZH. With context management, the scores further improve to 55.3% and 52.9%, respectively, significantly outperforming same-scale (4B) agents and even matching the performance of larger ones (approximately 30B). These results demonstrate the effectiveness of answer-backtracked step-level credit assignment for training long-horizon search agents.","upvotes":42,"discussionId":"6a740eacc5e410d076869b79","organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67b4079145dc598e0f110530","avatarUrl":"/avatars/30647b6d748a4ff2d1cb1c19daacd005.svg","isPro":false,"fullname":"Harry Lu","user":"HelloWorld9724","type":"user"},{"_id":"65257545b017be1fc1915364","avatarUrl":"/avatars/9bffd3fb567d2fa1e5c3546d77560b43.svg","isPro":false,"fullname":"Siheng Chen","user":"sihengchen","type":"user"},{"_id":"6a7416fdd20fba54d2c8ae42","avatarUrl":"/avatars/27933bcbdc8bdeef8b4704dfea754f57.svg","isPro":false,"fullname":"Zirui Song","user":"ziruisong","type":"user"},{"_id":"67934b85c67af4a116b5594b","avatarUrl":"/avatars/6a5a75cdbb8ddcdff16e3a8a1987d214.svg","isPro":false,"fullname":"yuwendu","user":"yuwendu","type":"user"},{"_id":"64dc39d27f749b6e34702b81","avatarUrl":"/avatars/3db6db301831b838dd172937ef7653df.svg","isPro":false,"fullname":"Du","user":"Dorothydu","type":"user"},{"_id":"64102816f52d7eb22e040659","avatarUrl":"/avatars/230dd27f8d942bbfec79a49d6c18177e.svg","isPro":false,"fullname":"zihao He","user":"RedRoman","type":"user"},{"_id":"65d205b5ff101ee25ee74bff","avatarUrl":"/avatars/f3912f7947c67394f9767b3e01c54d31.svg","isPro":false,"fullname":"XiangRui Liu","user":"gokoururi123","type":"user"},{"_id":"650ea589bfb7dd98bbb8c7a7","avatarUrl":"/avatars/9ae4436074ee9adc414f6fd84251aea5.svg","isPro":false,"fullname":"Brian Han","user":"hanbin92381","type":"user"},{"_id":"641086d27a15af878ae7eae1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641086d27a15af878ae7eae1/t-ZMlVNeJ8g01DW4hFc7l.jpeg","isPro":false,"fullname":"Rich Xu","user":"RichXuOvO","type":"user"},{"_id":"646f63f4753be77a8e94f95d","avatarUrl":"/avatars/771b8a10354906ae9d4cf827a54405d6.svg","isPro":false,"fullname":"yangcongge","user":"Bronion","type":"user"},{"_id":"685aa4b945bd540f62fd787b","avatarUrl":"/avatars/6f783be2eb0a382a9c9e62175cd97f8f.svg","isPro":false,"fullname":"yunfeng Wu","user":"yunfengWu","type":"user"},{"_id":"6614ed4809c63bfbddc53ddf","avatarUrl":"/avatars/018bfc168bb4010cf6018e42148e0f51.svg","isPro":false,"fullname":"Yuzhu Cai","user":"Ethical-Lens","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05102.md","query":{}}">
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Published on Aug 5
· Submitted by RuiYe on Aug 6 Abstract
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions. Specifically, given a potentially obscure query and its corresponding ground-truth answer, ABC first performs Answer-Backtracked Clue Recovery, which traces back from the answer to recover intermediate clues required to solve the question. It then applies Clue-Anchored Step Scoring to evaluate each search step against these clues, converting sparse binary outcome supervision into dense step-level rewards. Based on these rewards, we develop ABC-SFT, which reweights the loss of each turn, and ABC-GRPO, which uses the step-level scores as rewards in GRPO. Building on this framework, we train ABSeeker based on Qwen3.5-4B with only 8.5k examples. ABSeeker achieves 37.3% on BrowseComp and 39.1% on BrowseComp-ZH. With context management, the scores further improve to 55.3% and 52.9%, respectively, significantly outperforming same-scale (4B) agents and even matching the performance of larger ones (approximately 30B). These results demonstrate the effectiveness of answer-backtracked step-level credit assignment for training long-horizon search agents.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.05102 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.05102 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.