Hugging Face Daily Papers · · 6 min read

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some frontier models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark version for assessing software engineering agents.<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/659817704dbb962b7c724a61/5Je41flzrHaMuiVF8vwUk.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/659817704dbb962b7c724a61/5Je41flzrHaMuiVF8vwUk.png\" alt=\"results\"></a></p>\n","updatedAt":"2026-09-10T09:21:09.854Z","author":{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","fullname":"Shufan Jiang","name":"Tsumugii","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9031277298927307},"editors":["Tsumugii"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png"],"reactions":[],"isReport":false}},{"id":"6aa29e694469f54b614906d5","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-09-10T12:11:21.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Does the verified set actually reshuffle the leaderboard, or does it just shave a few points off everyone? Reward hacking is real — agents that game the grader look smarter than they are — but the interesting number isn't the new score, it's the rank delta. If the same agents stay on top in the same order, the benchmark is cleaner but not more informative. If a couple of previously top-tier agents drop hard, that's the signal worth publishing. I'd want to see per-agent before/after, not just the aggregate pass rate.","html":"<p>Does the verified set actually reshuffle the leaderboard, or does it just shave a few points off everyone? Reward hacking is real — agents that game the grader look smarter than they are — but the interesting number isn't the new score, it's the rank delta. If the same agents stay on top in the same order, the benchmark is cleaner but not more informative. If a couple of previously top-tier agents drop hard, that's the signal worth publishing. I'd want to see per-agent before/after, not just the aggregate pass rate.</p>\n","updatedAt":"2026-09-10T12:11:21.380Z","author":{"_id":"658412f93a84a40185adaf37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg","fullname":"Aamer Mihaysi","name":"O96a","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9531888961791992},"editors":["O96a"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/658412f93a84a40185adaf37/FKXH7e1jj09KO1v-B5sER.jpeg"],"reactions":[],"isReport":false},"replies":[{"id":"6aa2a212917a4ba89f9dc3fa","author":{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","fullname":"mzr1996","name":"mzr1996","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false},"createdAt":"2026-09-10T12:26:58.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Figure 1 is exactly the comparison between the original score and the verified score.","html":"<p>Figure 1 is exactly the comparison between the original score and the verified score.</p>\n","updatedAt":"2026-09-10T12:26:58.840Z","author":{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","fullname":"mzr1996","name":"mzr1996","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.967523992061615},"editors":["mzr1996"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg"],"reactions":[],"isReport":false,"parentCommentId":"6aa29e694469f54b614906d5"}}]}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08149","authors":[{"_id":"6aa0c46cd0174964227bec3c","name":"Pujun Zheng","hidden":false},{"_id":"6aa0c46cd0174964227bec3d","user":{"_id":"63a671eabe523d69b8c1a0be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671852504287-noauth.jpeg","isPro":false,"fullname":"Zixin Shang","user":"thinszx","type":"user","name":"thinszx"},"name":"Zixin Shang","status":"claimed_verified","statusLastChangedAt":"2026-09-10T10:00:20.314Z","hidden":false},{"_id":"6aa0c46cd0174964227bec3e","user":{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","isPro":false,"fullname":"Shufan Jiang","user":"Tsumugii","type":"user","name":"Tsumugii"},"name":"Shufan Jiang","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.146Z","hidden":false},{"_id":"6aa0c46cd0174964227bec3f","name":"Wenhui Tian","hidden":false},{"_id":"6aa0c46cd0174964227bec40","user":{"_id":"630da0fae57da204209411d3","avatarUrl":"/avatars/e79c250cf8031441ffd0e853e653cef6.svg","isPro":false,"fullname":"dongsheng zhu","user":"dongsheng","type":"user","name":"dongsheng"},"name":"Dongsheng Zhu","status":"claimed_verified","statusLastChangedAt":"2026-09-10T08:45:05.139Z","hidden":false},{"_id":"6aa0c46cd0174964227bec41","user":{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","isPro":false,"fullname":"mzr1996","user":"mzr1996","type":"user","name":"mzr1996"},"name":"Zerun Ma","status":"claimed_verified","statusLastChangedAt":"2026-09-10T10:00:23.459Z","hidden":false},{"_id":"6aa0c46cd0174964227bec42","user":{"_id":"69d61bb3f84d9d31e5bebdc8","avatarUrl":"/avatars/eddae5b1efdde2826211be26529bf80c.svg","isPro":false,"fullname":"Dingbo Yuan","user":"bigwavelet90","type":"user","name":"bigwavelet90"},"name":"Dingbo Yuan","status":"claimed_verified","statusLastChangedAt":"2026-09-10T16:45:04.658Z","hidden":false},{"_id":"6aa0c46cd0174964227bec43","name":"Qi Zhang","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents","submittedOnDailyBy":{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","isPro":false,"fullname":"Shufan Jiang","user":"Tsumugii","type":"user","name":"Tsumugii"},"summary":"SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.","upvotes":18,"discussionId":"6aa0c46cd0174964227bec44","projectPage":"https://agent-compass.mintlify.app/en/user_guide/modules/benchmarks/swebench_pro_verified","organization":{"_id":"6a4fb75a1c66dbf208e7ddb6","name":"Shanghai-AI-Laboratory","fullname":"Shanghai AI Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65cd955637be1841d0b75397/Rao_Kq6NMtTVfSqLUIR4k.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69d61bb3f84d9d31e5bebdc8","avatarUrl":"/avatars/eddae5b1efdde2826211be26529bf80c.svg","isPro":false,"fullname":"Dingbo Yuan","user":"bigwavelet90","type":"user"},{"_id":"67327d3f870f6ea855f2c23d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/WqIOzN2FvTM8QZhwmV_mX.png","isPro":false,"fullname":"Zhe Sun","user":"ZheSun","type":"user"},{"_id":"630da0fae57da204209411d3","avatarUrl":"/avatars/e79c250cf8031441ffd0e853e653cef6.svg","isPro":false,"fullname":"dongsheng zhu","user":"dongsheng","type":"user"},{"_id":"659817704dbb962b7c724a61","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/Pe3eKSx-hkkgYAvl_mGx0.png","isPro":false,"fullname":"Shufan Jiang","user":"Tsumugii","type":"user"},{"_id":"64decfc2de3048e81c78f355","avatarUrl":"/avatars/d56ad97160438c173617fdc99eb7b205.svg","isPro":false,"fullname":"Liwei Wu","user":"ssiq-wu","type":"user"},{"_id":"68b507495ceb7eb99e30a7a8","avatarUrl":"/avatars/dab70077494a9c9cedd09bace0a1430c.svg","isPro":false,"fullname":"Pujun Zheng","user":"zpjbtdjm","type":"user"},{"_id":"634fb993685e43e02d2544a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/634fb993685e43e02d2544a2/L9MuXYlXMBuQgnwXd4lCU.jpeg","isPro":false,"fullname":"mzr1996","user":"mzr1996","type":"user"},{"_id":"6703ec213df5fe425086ef73","avatarUrl":"/avatars/e6f9dad6587ee0883ae10f8805ab7ea9.svg","isPro":true,"fullname":"Tianhao Liang","user":"tianhao2k","type":"user"},{"_id":"63a671eabe523d69b8c1a0be","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671852504287-noauth.jpeg","isPro":false,"fullname":"Zixin Shang","user":"thinszx","type":"user"},{"_id":"65b92c8d98d8720151544c32","avatarUrl":"/avatars/86a71da795ff0384d292c14dfa93ba57.svg","isPro":false,"fullname":"Zhou","user":"myhs","type":"user"},{"_id":"652112110415e1b734b97d6b","avatarUrl":"/avatars/c13775a9737cab0d03952531e367be87.svg","isPro":false,"fullname":"rrh","user":"RHanhan","type":"user"},{"_id":"64f5964a413ca787f12b8ade","avatarUrl":"/avatars/2b63ace471e08fa0f22c873788424a33.svg","isPro":false,"fullname":"Yang Penghui","user":"ygyjrc","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a4fb75a1c66dbf208e7ddb6","name":"Shanghai-AI-Laboratory","fullname":"Shanghai AI Laboratory","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65cd955637be1841d0b75397/Rao_Kq6NMtTVfSqLUIR4k.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08149.md","query":{}}">
Papers
arxiv:2609.08149

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

Published on Sep 8
· Submitted by
Shufan Jiang
on Sep 10

Abstract

SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents.

Community

Paper author Paper submitter about 8 hours ago

SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evaluation information, and task quality issues, including misleading problem statements and improperly scoped tests. These issues can inflate benchmark performance and obscure agents' true coding ability. We present SWE-Bench Pro Verified, a verified version of SWE-Bench Pro that addresses both problems. Our approach combines anti-hacking safeguards that eliminate major leakage channels without disrupting normal agent functionality, with task refinement that minimally corrects inconsistencies within flawed instances. Evaluations on SWE-Bench Pro Verified reveal that some frontier models perform substantially worse than previously reported, suggesting that existing results on SWE-Bench Pro may overestimate real software engineering capability. SWE-Bench Pro Verified offers a more trustworthy benchmark version for assessing software engineering agents.
results

Does the verified set actually reshuffle the leaderboard, or does it just shave a few points off everyone? Reward hacking is real — agents that game the grader look smarter than they are — but the interesting number isn't the new score, it's the rank delta. If the same agents stay on top in the same order, the benchmark is cleaner but not more informative. If a couple of previously top-tier agents drop hard, that's the signal worth publishing. I'd want to see per-agent before/after, not just the aggregate pass rate.

Paper author about 5 hours ago

Figure 1 is exactly the comparison between the original score and the verified score.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Upvote
18

Get this paper in your agent:

hf papers read 2609.08149
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.08149 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.08149 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers