Hugging Face Daily Papers · · 3 min read

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

We introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities.</p>\n","updatedAt":"2026-07-14T03:22:58.311Z","author":{"_id":"6601196cc91ba4c08ad6e270","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6601196cc91ba4c08ad6e270/venywO3WPi2fNi5WUJTH0.jpeg","fullname":"Yuzhe Gu","name":"vanilla1116","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":10,"isUserFollowing":false}},"numEdits":0,"editors":["vanilla1116"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6601196cc91ba4c08ad6e270/venywO3WPi2fNi5WUJTH0.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.11849","authors":[{"_id":"6a55a5a2a9d74d6e65bbd1bc","name":"Lingkai Kong","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1bd","name":"Zijian Wu","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1be","name":"Yuzhe Gu","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1bf","name":"Haiteng Zhao","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c0","name":"Wenyong Huang","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c1","name":"Shuang Sun","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c2","name":"Zhicheng Xiong","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c3","name":"Xiaotian Zhang","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c4","name":"Shuya Zhao","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c5","name":"Yan Wang","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c6","name":"Disheng Xu","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c7","name":"Wenwei Zhang","hidden":false},{"_id":"6a55a5a2a9d74d6e65bbd1c8","name":"Kai Chen","hidden":false}],"publishedAt":"2026-07-13T17:38:22.000Z","submittedOnDailyAt":"2026-07-14T00:00:00.000Z","title":"AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification","submittedOnDailyBy":{"_id":"6601196cc91ba4c08ad6e270","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6601196cc91ba4c08ad6e270/venywO3WPi2fNi5WUJTH0.jpeg","isPro":false,"fullname":"Yuzhe Gu","user":"vanilla1116","type":"user","name":"vanilla1116"},"summary":"Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-generation benchmark, ProverBench, contains 296 problems spanning undergraduate and doctoral qualifying-exam levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 75.8 and 66.1 on the UGD and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model attains a Balanced F1 of only 65.1, and models generally exhibit low true negative rates, suggesting that critical error detection remains a major bottleneck.","upvotes":19,"discussionId":"6a55a5a2a9d74d6e65bbd1c9","organization":{"_id":"64a2d5fa81252883206f24c9","name":"internlm","fullname":"Intern Large Models","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6432683407bad11484a68457/Q3Y0dL79GcsnaBCGRMooZ.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6601196cc91ba4c08ad6e270","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6601196cc91ba4c08ad6e270/venywO3WPi2fNi5WUJTH0.jpeg","isPro":false,"fullname":"Yuzhe Gu","user":"vanilla1116","type":"user"},{"_id":"67652fc11fde77e3bb8b017d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/dVuf5R3vGZuK8uacp1d2f.png","isPro":false,"fullname":"Yinhao Tang","user":"tangyinhao","type":"user"},{"_id":"6431a7c31e22a07ceb8b7f1e","avatarUrl":"/avatars/772167de9f529a959612e1cf3b2b908c.svg","isPro":false,"fullname":"Yanan Sun","user":"nowsyn","type":"user"},{"_id":"6413e0c350358a805205f540","avatarUrl":"/avatars/13a82cdd7c0b9f12206cb8a3e2d3809b.svg","isPro":false,"fullname":"shuo shen","user":"hyperion-shuo","type":"user"},{"_id":"64e8505321540e1da3226b54","avatarUrl":"/avatars/18958b8406d1ce492b54c1c839f18c54.svg","isPro":false,"fullname":"Wenwei Zhang","user":"ZwwWayne","type":"user"},{"_id":"6600f734997ede4f9b1bb33e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6600f734997ede4f9b1bb33e/vvBsz6c1WUzAl9ifGSl_m.jpeg","isPro":false,"fullname":"Fang","user":"Youqing","type":"user"},{"_id":"64a7c6e223622f7f189bcbe1","avatarUrl":"/avatars/4f13a7ed0d2b8d8dfff7dc650e46450a.svg","isPro":false,"fullname":"haiteng zhao","user":"haitengzhao","type":"user"},{"_id":"6385f8598b5acae8d24caf16","avatarUrl":"/avatars/9d261f95d24e882157b987b8827098be.svg","isPro":false,"fullname":"liuwenran","user":"lwrshi1965","type":"user"},{"_id":"63fd691794cc8f815d50c112","avatarUrl":"/avatars/87305d1cbfcc717e910ccdfaf0568f80.svg","isPro":false,"fullname":"liu","user":"Harold-lkk","type":"user"},{"_id":"64ccaa4687ec96aa4752e754","avatarUrl":"/avatars/d2dd2040a521de4f55c7335cb7771c75.svg","isPro":false,"fullname":"Yiming Zhang","user":"ymzhang319","type":"user"},{"_id":"65a7646581cc3017643a217d","avatarUrl":"/avatars/ffc62e6f774e7d76b6a38aff432e01e4.svg","isPro":false,"fullname":"Jiangning Liu","user":"liujiangning","type":"user"},{"_id":"64dee100e437d02ce6aca5ee","avatarUrl":"/avatars/838720434abad60e8dee8e85b4d402f5.svg","isPro":false,"fullname":"Guangran Cheng","user":"penny123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64a2d5fa81252883206f24c9","name":"internlm","fullname":"Intern Large Models","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6432683407bad11484a68457/Q3Y0dL79GcsnaBCGRMooZ.png"},"query":{}}">
Papers
arxiv:2607.11849

AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification

Published on Jul 13
· Submitted by
Yuzhe Gu
on Jul 14
Authors:
,

Abstract

Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-generation benchmark, ProverBench, contains 296 problems spanning undergraduate and doctoral qualifying-exam levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 75.8 and 66.1 on the UGD and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model attains a Balanced F1 of only 65.1, and models generally exhibit low true negative rates, suggesting that critical error detection remains a major bottleneck.

Community

We introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.11849 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.11849 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.11849 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers