SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some frontier models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark version for assessing software engineering agents.<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/659817704dbb962b7c724a61/5Je41flzrHaMuiVF8vwUk.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/659817704dbb962b7c724a61/5Je41flzrHaMuiVF8vwUk.png\" alt=\"results\"></a></p>\n","updatedAt":"2026-09-10T09:21:09.854Z","author":{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","fullname":"Shufan Jiang","name":"Tsumugii","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9031277298927307},"editors":["Tsumugii"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png"],"reactions":[],"isReport":false}},{"id":"6aa29e694469f54b614906d5","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-09-10T12:11:21.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Does the verified set actually reshuffle the leaderboard, or does it just shave a few points off everyone? Reward hacking is real — agents that game the grader look smarter than they are — but the interesting number isn't the new score, it's the rank delta. If the same agents stay on top in the same order, the benchmark is cleaner but not more informative. If a couple of previously top-tier agents drop hard, that's the signal worth publishing. I'd want to see per-agent before/after, not just the aggregate pass rate.","html":"<p>Does the verified set actually reshuffle the leaderboard, or does it just shave a few points off everyone? Reward hacking is real — agents that game the grader look smarter than they are — but the interesting number isn't the new score, it's the rank delta. If the same agents stay on top in the same order, the benchmark is cleaner but not more informative. If a couple of previously top-tier agents drop hard, that's the signal worth publishing. I'd want to see per-agent before/after, not just the aggregate pass rate.</p>\n","updatedAt":"2026-09-10T12:11:21.380Z","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9531888961791992},"editors":["O96a"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg"],"reactions":[],"isReport":false},"replies":[{"id":"6aa2a212917a4ba89f9dc3fa","author":{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","fullname":"mzr1996","name":"mzr1996","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"createdAt":"2026-09-10T12:26:58.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Figure 1 is exactly the comparison between the original score and the verified score.","html":"<p>Figure 1 is exactly the comparison between the original score and the verified score.</p>\n","updatedAt":"2026-09-10T12:26:58.840Z","author":{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","fullname":"mzr1996","name":"mzr1996","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.967523992061615},"editors":["mzr1996"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg"],"reactions":[],"isReport":false,"parentCommentId":"6aa29e694469f54b614906d5"}}]}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08149","authors":[{"_id":"6aa0c46cd0174964227bec3c","name":"Pujun Zheng","hidden":false},{"_id":"6aa0c46cd0174964227bec3d","user":{"_id":"63a671eabe523d69b8c1a0be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671852504287-noauth.jpeg","isPro":false,"fullname":"Zixin Shang","user":"thinszx","type":"user","name":"thinszx"},"name":"Zixin Shang","status":"claimed_verified","statusLastChangedAt":"2026-09-10T10:00:20.314Z","hidden":false},{"_id":"6aa0c46cd0174964227bec3e","user":{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","isPro":false,"fullname":"Shufan Jiang","user":"Tsumugii","type":"user","name":"Tsumugii"},"name":"Shufan Jiang","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.146Z","hidden":false},{"_id":"6aa0c46cd0174964227bec3f","name":"Wenhui Tian","hidden":false},{"_id":"6aa0c46cd0174964227bec40","user":{"_id":"630da0fae57da204209411d3","avatarUrl":"/avatars/e79c250cf8031441ffd0e853e653cef6.svg","isPro":false,"fullname":"dongsheng zhu","user":"dongsheng","type":"user","name":"dongsheng"},"name":"Dongsheng Zhu","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.139Z","hidden":false},{"_id":"6aa0c46cd0174964227bec41","user":{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","isPro":false,"fullname":"mzr1996","user":"mzr1996","type":"user","name":"mzr1996"},"name":"Zerun Ma","status":"claimed_verified","statusLastChangedAt":"2026-09-10T10:00:23.459Z","hidden":false},{"_id":"6aa0c46cd0174964227bec42","user":{"_id":"69d61bb3f84d9d31e5bebdc8","avatarUrl":"/avatars/eddae5b1efdde2826211be26529bf80c.svg","isPro":false,"fullname":"Dingbo Yuan","user":"bigwavelet90","type":"user","name":"bigwavelet90"},"name":"Dingbo Yuan","status":"claimed_verified","statusLastChangedAt":"2026-09-10T16:45:04.658Z","hidden":false},{"_id":"6aa0c46cd0174964227bec43","name":"Qi Zhang","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents","submittedOnDailyBy":{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","isPro":false,"fullname":"Shufan Jiang","user":"Tsumugii","type":"user","name":"Tsumugii"},"summary":"SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.","upvotes":18,"discussionId":"6aa0c46cd0174964227bec44","projectPage":"https://agent-compass.mintlify.app/en/user_guide/modules/benchmarks/swebench_pro_verified","organization":{"_id":"6a4fb75a1c66dbf208e7ddb6","name":"Shanghai-AI-Laboratory","fullname":"Shanghai AI Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65cd955637be1841d0b75397/Rao_Kq6NMtTVfSqLUIR4k.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69d61bb3f84d9d31e5bebdc8","avatarUrl":"/avatars/eddae5b1efdde2826211be26529bf80c.svg","isPro":false,"fullname":"Dingbo Yuan","user":"bigwavelet90","type":"user"},{"_id":"67327d3f870f6ea855f2c23d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/WqIOzN2FvTM8QZhwmV_mX.png","isPro":false,"fullname":"Zhe Sun","user":"ZheSun","type":"user"},{"_id":"630da0fae57da204209411d3","avatarUrl":"/avatars/e79c250cf8031441ffd0e853e653cef6.svg","isPro":false,"fullname":"dongsheng zhu","user":"dongsheng","type":"user"},{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","isPro":false,"fullname":"Shufan Jiang","user":"Tsumugii","type":"user"},{"_id":"64decfc2de3048e81c78f355","avatarUrl":"/avatars/d56ad97160438c173617fdc99eb7b205.svg","isPro":false,"fullname":"Liwei Wu","user":"ssiq-wu","type":"user"},{"_id":"68b507495ceb7eb99e30a7a8","avatarUrl":"/avatars/dab70077494a9c9cedd09bace0a1430c.svg","isPro":false,"fullname":"Pujun Zheng","user":"zpjbtdjm","type":"user"},{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","isPro":false,"fullname":"mzr1996","user":"mzr1996","type":"user"},{"_id":"6703ec213df5fe425086ef73","avatarUrl":"/avatars/e6f9dad6587ee0883ae10f8805ab7ea9.svg","isPro":true,"fullname":"Tianhao Liang","user":"tianhao2k","type":"user"},{"_id":"63a671eabe523d69b8c1a0be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671852504287-noauth.jpeg","isPro":false,"fullname":"Zixin Shang","user":"thinszx","type":"user"},{"_id":"65b92c8d98d8720151544c32","avatarUrl":"/avatars/86a71da795ff0384d292c14dfa93ba57.svg","isPro":false,"fullname":"Zhou","user":"myhs","type":"user"},{"_id":"652112110415e1b734b97d6b","avatarUrl":"/avatars/c13775a9737cab0d03952531e367be87.svg","isPro":false,"fullname":"rrh","user":"RHanhan","type":"user"},{"_id":"64f5964a413ca787f12b8ade","avatarUrl":"/avatars/2b63ace471e08fa0f22c873788424a33.svg","isPro":false,"fullname":"Yang Penghui","user":"ygyjrc","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a4fb75a1c66dbf208e7ddb6","name":"Shanghai-AI-Laboratory","fullname":"Shanghai AI Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65cd955637be1841d0b75397/Rao_Kq6NMtTVfSqLUIR4k.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08149.md","query":{}}">
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Abstract
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.
Community
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some frontier models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark version for assessing software engineering agents.

Does the verified set actually reshuffle the leaderboard, or does it just shave a few points off everyone? Reward hacking is real — agents that game the grader look smarter than they are — but the interesting number isn't the new score, it's the rank delta. If the same agents stay on top in the same order, the benchmark is cleaner but not more informative. If a couple of previously top-tier agents drop hard, that's the signal worth publishing. I'd want to see per-agent before/after, not just the aggregate pass rate.
Figure 1 is exactly the comparison between the original score and the verified score.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.