Hugging Face Daily Papers · · 5 min read

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

How do you know your translation benchmark is contamination-free and that its scores aren't hiding important differences between locales within the same language? You need to check out Cultivar.</p>\n","updatedAt":"2026-08-11T19:34:23.565Z","author":{"_id":"641248d400634c4fe98589a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641248d400634c4fe98589a8/edYgAz2j_3F4B1h-0Ma6o.png","fullname":"Pinzhen Chen","name":"pinzhenchen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9478712677955627},"editors":["pinzhenchen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/641248d400634c4fe98589a8/edYgAz2j_3F4B1h-0Ma6o.png"],"reactions":[],"isReport":false}},{"id":"6a7bcea67c996dc5eb52e55b","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false},"createdAt":"2026-08-12T01:38:46.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia](https://huggingface.co/papers/2606.20212) (2026)\n* [ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs](https://huggingface.co/papers/2607.00171) (2026)\n* [MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages](https://huggingface.co/papers/2607.00890) (2026)\n* [IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages](https://huggingface.co/papers/2606.19157) (2026)\n* [CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages](https://huggingface.co/papers/2607.21016) (2026)\n* [PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction](https://huggingface.co/papers/2606.19096) (2026)\n* [Multilingual Reasoning Cascades Need More Context](https://huggingface.co/papers/2606.27306) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2606.20212\">CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.00171\">ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.00890\">MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.19157\">IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.21016\">CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.19096\">PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2606.27306\">Multilingual Reasoning Cascades Need More Context</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-08-12T01:38:46.588Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":378,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6938313841819763},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09766","authors":[{"_id":"6a7b77d6b7183653340c13e6","user":{"_id":"641248d400634c4fe98589a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641248d400634c4fe98589a8/edYgAz2j_3F4B1h-0Ma6o.png","isPro":false,"fullname":"Pinzhen Chen","user":"pinzhenchen","type":"user","name":"pinzhenchen"},"name":"Pinzhen Chen","status":"claimed_verified","statusLastChangedAt":"2026-08-12T00:45:04.521Z","hidden":false},{"_id":"6a7b77d6b7183653340c13e7","name":"Koel Dutta Chowdhury","hidden":false},{"_id":"6a7b77d6b7183653340c13e8","name":"Xiaoya Xu","hidden":false},{"_id":"6a7b77d6b7183653340c13e9","name":"David Tan","hidden":false},{"_id":"6a7b77d6b7183653340c13ea","name":"Doreen Osmelak","hidden":false},{"_id":"6a7b77d6b7183653340c13eb","name":"Ona de Gibert","hidden":false},{"_id":"6a7b77d6b7183653340c13ec","name":"Ariun-Erdene Tumurchuluun","hidden":false},{"_id":"6a7b77d6b7183653340c13ed","name":"Ashok Urlana","hidden":false},{"_id":"6a7b77d6b7183653340c13ee","name":"Fedor Sizov","hidden":false},{"_id":"6a7b77d6b7183653340c13ef","name":"Hale Sirin","hidden":false},{"_id":"6a7b77d6b7183653340c13f0","name":"Jesujoba Alabi","hidden":false},{"_id":"6a7b77d6b7183653340c13f1","name":"Karrar Talib Abed","hidden":false},{"_id":"6a7b77d6b7183653340c13f2","name":"Mateusz Klimaszewski","hidden":false},{"_id":"6a7b77d6b7183653340c13f3","name":"Nikolay Bogoychev","hidden":false},{"_id":"6a7b77d6b7183653340c13f4","name":"Niyati Bafna","hidden":false},{"_id":"6a7b77d6b7183653340c13f5","name":"Patricia Schmidtova","hidden":false},{"_id":"6a7b77d6b7183653340c13f6","name":"Preksha Manjunath Shanbhag","hidden":false},{"_id":"6a7b77d6b7183653340c13f7","name":"Sherrie Shen","hidden":false},{"_id":"6a7b77d6b7183653340c13f8","name":"Vilem Zouhar","hidden":false},{"_id":"6a7b77d6b7183653340c13f9","name":"Vivek Iyer","hidden":false},{"_id":"6a7b77d6b7183653340c13fa","name":"Yasser Hamidullah","hidden":false},{"_id":"6a7b77d6b7183653340c13fb","name":"Yusser Al Ghussin","hidden":false},{"_id":"6a7b77d6b7183653340c13fc","name":"Zheng Zhao","hidden":false}],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness","submittedOnDailyBy":{"_id":"641248d400634c4fe98589a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641248d400634c4fe98589a8/edYgAz2j_3F4B1h-0Ma6o.png","isPro":false,"fullname":"Pinzhen Chen","user":"pinzhenchen","type":"user","name":"pinzhenchen"},"summary":"Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.","upvotes":2,"discussionId":"6a7b77d6b7183653340c13fd","projectPage":"https://huggingface.co/datasets/pinzhenchen/Cultivar-flores","ai_summary":"Researchers propose source-contrastive evaluation via a localized benchmark to detect data contamination and assess localization robustness in multilingual translation models.","ai_keywords":["source-contrastive evaluation","FLORES","locale-specific translation evaluation","data contamination","localization robustness","MT-specialised models"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"63a3249eb5fc9ab9f63bb43f","name":"QUBelfast","fullname":"Queen's University Belfast","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671636059729-63a323c3a0ed04b823c0c8c2.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"641248d400634c4fe98589a8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/641248d400634c4fe98589a8/edYgAz2j_3F4B1h-0Ma6o.png","isPro":false,"fullname":"Pinzhen Chen","user":"pinzhenchen","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63a3249eb5fc9ab9f63bb43f","name":"QUBelfast","fullname":"Queen's University Belfast","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671636059729-63a323c3a0ed04b823c0c8c2.png"},"query":{}}">
Papers
arxiv:2608.09766

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Published on Aug 10
· Submitted by
Pinzhen Chen
on Aug 11
Authors:

Abstract

Researchers propose source-contrastive evaluation via a localized benchmark to detect data contamination and assess localization robustness in multilingual translation models.

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

Community

Paper author Paper submitter about 6 hours ago

How do you know your translation benchmark is contamination-free and that its scores aren't hiding important differences between locales within the same language? You need to check out Cultivar.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.09766 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.09766 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers