If you've tried building RAG in Greek, you've probably hit the same wall we did: most multilingual retrievers haven't seen enough Greek, and there's no benchmark to tell you how bad it is. So we adapted the Nemotron retrieval stack end to end and built HERA, a benchmark covering legal, energy, financial and medical Greek. The finding that surprised us most was plain BM25 beating several off-the-shelf dense models on specialist corpora.. fine-tuning fixed that. Models and benchmark are released, and feedback is welcome, especially from people working on other low-resource languages.</p>\n","updatedAt":"2026-08-07T08:14:26.913Z","author":{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","fullname":"Ayoub Kirouane","name":"ayoubkirouane","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":68,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.952118992805481},"editors":["ayoubkirouane"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05138","authors":[{"_id":"6a74940fe1228e04b3237ec8","user":{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","isPro":false,"fullname":"Ayoub Kirouane","user":"ayoubkirouane","type":"user","name":"ayoubkirouane"},"name":"Ayoub Kirouane","status":"claimed_verified","statusLastChangedAt":"2026-08-06T16:45:04.625Z","hidden":false},{"_id":"6a74940fe1228e04b3237ec9","user":{"_id":"6799d2770a1d27acbcc635bb","avatarUrl":"/avatars/66af80ba8bba84f9105551fc41314b42.svg","isPro":false,"fullname":"Christos Petrocheilos","user":"cpetos","type":"user","name":"cpetos"},"name":"Christos Petrocheilos","status":"claimed_verified","statusLastChangedAt":"2026-08-07T08:45:04.470Z","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains","submittedOnDailyBy":{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","isPro":false,"fullname":"Ayoub Kirouane","user":"ayoubkirouane","type":"user","name":"ayoubkirouane"},"summary":"Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.","upvotes":16,"discussionId":"6a749410e1228e04b3237eca","organization":{"_id":"68f14d43f5266b706cf99e17","name":"KIEFERSA","fullname":"KIEFER","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68f14c3c6ede47be458ff051/7iTmA80GKtq3CNQDhvUoc.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6338c06c107c4835a05699f9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6338c06c107c4835a05699f9/uQbrMgQySY2UW7z3R9gT5.jpeg","isPro":false,"fullname":"Ayoub Kirouane","user":"ayoubkirouane","type":"user"},{"_id":"698d85c8a82ef9f4893bc3b0","avatarUrl":"/avatars/6664ce50983cb96346b94e750e8440ef.svg","isPro":false,"fullname":"Christos Porikis","user":"chrispor96","type":"user"},{"_id":"67ccb4e91bacab3ae8c5ebc1","avatarUrl":"/avatars/3632fbe99ea8f7b36d90e10c4029b24d.svg","isPro":false,"fullname":"Giaples Georgios","user":"gigigiapl","type":"user"},{"_id":"6799d2770a1d27acbcc635bb","avatarUrl":"/avatars/66af80ba8bba84f9105551fc41314b42.svg","isPro":false,"fullname":"Christos Petrocheilos","user":"cpetos","type":"user"},{"_id":"698d86a50eca9dd7d1179d76","avatarUrl":"/avatars/ee6f2050f992e2852492e041ff617aed.svg","isPro":false,"fullname":"Hasan Degismez","user":"hasankiefer","type":"user"},{"_id":"63cc5c5180ba2ca41528499a","avatarUrl":"/avatars/3b0a2c4813b1f8e7c6e14d3b7829ed9c.svg","isPro":false,"fullname":"Hs Dz","user":"pupdz","type":"user"},{"_id":"68a6dd8ecccf20993295dafe","avatarUrl":"/avatars/64651415a91e34ef688aee4c49d1afb9.svg","isPro":false,"fullname":"John Karvounas","user":"jkarvounas","type":"user"},{"_id":"698d8f2d7cd86ed5aac6b612","avatarUrl":"/avatars/dc7cca7a251a2f28b0c149aa12f3f759.svg","isPro":false,"fullname":"Themistoklis Nikolis","user":"kiefer-dev-1","type":"user"},{"_id":"6a75a625e7c0b16367422744","avatarUrl":"/avatars/595c601db74ddcb885203d5f32eb1416.svg","isPro":false,"fullname":"Lampis Papakostas","user":"lampispap","type":"user"},{"_id":"6a39091e977556eb25bc44fa","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/eJAo8xa1nlfD-EYp7WJZ_.png","isPro":false,"fullname":"Petr Levtonov","user":"TheLevti","type":"user"},{"_id":"6a34f40fc88b1fc59a2e5d99","avatarUrl":"/avatars/87d5267bcab507c417b172b167d1ab57.svg","isPro":false,"fullname":"Kimis Perros","user":"kperros-kiefer","type":"user"},{"_id":"6a75aaef22e97c5cc2df05ae","avatarUrl":"/avatars/2da12b8af7b5a23106e08f25f44c2748.svg","isPro":false,"fullname":"Apsotolos Stavridis","user":"Apostolos82","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68f14d43f5266b706cf99e17","name":"KIEFERSA","fullname":"KIEFER","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68f14c3c6ede47be458ff051/7iTmA80GKtq3CNQDhvUoc.png"},"query":{}}">
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Abstract
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.
Community
If you've tried building RAG in Greek, you've probably hit the same wall we did: most multilingual retrievers haven't seen enough Greek, and there's no benchmark to tell you how bad it is. So we adapted the Nemotron retrieval stack end to end and built HERA, a benchmark covering legal, energy, financial and medical Greek. The finding that surprised us most was plain BM25 beating several off-the-shelf dense models on specialist corpora.. fine-tuning fixed that. Models and benchmark are released, and feedback is welcome, especially from people working on other low-resource languages.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.05138 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.05138 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.05138 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.