Hugging Face Daily Papers · · 5 min read

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages.</p>\n<p>To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional underrepresented languages spanning 6 language families — ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we benchmark 27 reasoning LLMs across four model scales — small, mid-size, large, and closed-source ensembles — probing multilingual mathematical reasoning under diverse linguistic conditions.</p>\n<p>Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.</p>\n","updatedAt":"2026-07-08T07:34:56.289Z","author":{"_id":"615c283104fe2f8312e110d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/615c283104fe2f8312e110d1/OYlFtEipvIAGCxwjTdZTk.jpeg","fullname":"Daryna Dementieva","name":"dardem","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":13,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8982733488082886},"editors":["dardem"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/615c283104fe2f8312e110d1/OYlFtEipvIAGCxwjTdZTk.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.05992","authors":[{"_id":"6a4dfc9125849b193a834c75","name":"Daryna Dementieva","hidden":false},{"_id":"6a4dfc9125849b193a834c76","name":"Nikolay Babakov","hidden":false},{"_id":"6a4dfc9125849b193a834c77","name":"Kathy Hämmerl","hidden":false},{"_id":"6a4dfc9125849b193a834c78","name":"Ilseyar Alimova","hidden":false},{"_id":"6a4dfc9125849b193a834c79","name":"Jindřich Libovický","hidden":false},{"_id":"6a4dfc9125849b193a834c7a","name":"Shu Okabe","hidden":false},{"_id":"6a4dfc9125849b193a834c7b","name":"Miras Baisbay","hidden":false},{"_id":"6a4dfc9125849b193a834c7c","name":"Lukas Edman","hidden":false},{"_id":"6a4dfc9125849b193a834c7d","name":"Abrorkhon Inomkhujaev","hidden":false},{"_id":"6a4dfc9125849b193a834c7e","name":"Antonia Karamolegkou","hidden":false},{"_id":"6a4dfc9125849b193a834c7f","name":"Mateusz Lango","hidden":false},{"_id":"6a4dfc9125849b193a834c80","name":"Volkan Özer","hidden":false},{"_id":"6a4dfc9125849b193a834c81","name":"Nikola Selic","hidden":false},{"_id":"6a4dfc9125849b193a834c82","name":"Subhankar Swain","hidden":false},{"_id":"6a4dfc9125849b193a834c83","name":"Tsedeniya Kinfe Temesgen","hidden":false},{"_id":"6a4dfc9125849b193a834c84","name":"Galit Bary Weisberg","hidden":false},{"_id":"6a4dfc9125849b193a834c85","name":"Alexander Fraser","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/615c283104fe2f8312e110d1/1I8hdR9wxMqtx66g50WHI.png","https://cdn-uploads.huggingface.co/production/uploads/615c283104fe2f8312e110d1/wJaoORwxHTDe_qCM02JKv.png","https://cdn-uploads.huggingface.co/production/uploads/615c283104fe2f8312e110d1/mmIrCXuLca3qIGwqQ6Bk_.png"],"publishedAt":"2026-07-07T00:00:00.000Z","submittedOnDailyAt":"2026-07-08T00:00:00.000Z","title":"PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages","submittedOnDailyBy":{"_id":"615c283104fe2f8312e110d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/615c283104fe2f8312e110d1/OYlFtEipvIAGCxwjTdZTk.jpeg","isPro":false,"fullname":"Daryna Dementieva","user":"dardem","type":"user","name":"dardem"},"summary":"Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.","upvotes":3,"discussionId":"6a4dfc9125849b193a834c86","projectPage":"https://tum-nlp.github.io/pluramath/","githubRepo":"https://github.com/TUM-NLP/pluramath","githubRepoAddedBy":"user","ai_summary":"PluraMath extends the PolyMath dataset to 18 underrepresented languages, revealing persistent gaps in multilingual mathematical reasoning performance between high-resource and low-resource languages.","ai_keywords":["Large Language Models","mathematical reasoning","multilingual benchmark","underrepresented languages","PolyMath","instruction-following ability"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":0,"organization":{"_id":"63f3854da096536aeaadcc20","name":"tum-nlp","fullname":"Natural Language Processing @ TUM","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/615c283104fe2f8312e110d1/_-VwOyJ6hwEk4RLzZBVX_.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"615c283104fe2f8312e110d1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/615c283104fe2f8312e110d1/OYlFtEipvIAGCxwjTdZTk.jpeg","isPro":false,"fullname":"Daryna Dementieva","user":"dardem","type":"user"},{"_id":"63d148d1b30415240fd23970","avatarUrl":"/avatars/44acb120ac4b537032c3f439f6211842.svg","isPro":false,"fullname":"Miras Baisbay","user":"mirasai","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63f3854da096536aeaadcc20","name":"tum-nlp","fullname":"Natural Language Processing @ TUM","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/615c283104fe2f8312e110d1/_-VwOyJ6hwEk4RLzZBVX_.png"},"query":{}}">
Papers
arxiv:2607.05992

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Published on Jul 7
· Submitted by
Daryna Dementieva
on Jul 8
Authors:
,

Abstract

PluraMath extends the PolyMath dataset to 18 underrepresented languages, revealing persistent gaps in multilingual mathematical reasoning performance between high-resource and low-resource languages.

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.

Community

Paper submitter about 9 hours ago

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages.

To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional underrepresented languages spanning 6 language families — ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we benchmark 27 reasoning LLMs across four model scales — small, mid-size, large, and closed-source ensembles — probing multilingual mathematical reasoning under diverse linguistic conditions.

Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.05992 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.05992 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers