Hugging Face Daily Papers · · 4 min read

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Yiddish carries centuries of literature, humor, and culture—yet it has almost no modern NLP infrastructure.</p>\n<p>With <strong>MameloshnLM</strong>, we introduce the first dedicated Yiddish language model and the first Yiddish LLM benchmark. We show that:</p>\n<ul>\n<li><p>Multilingual web corpora contain substantially more noise and machine-translated text than native sources.</p>\n</li>\n<li><p>Models trained on this data flatten the language, weakening its idioms, style, and cultural voice.</p>\n</li>\n<li><p>Training on authentic Yiddish enables MameloshnLM to better preserve the language’s distinctive character.</p>\n</li>\n</ul>\n<p>Supporting a language means more than generating its words—it means ensuring it still sounds like itself.</p>\n","updatedAt":"2026-08-07T10:30:04.982Z","author":{"_id":"604e1c120fe8ff3ec13d71e8","avatarUrl":"/avatars/22f6463216904fb0ec8306e704432ab7.svg","fullname":"Uri Katz","name":"Uri-ka","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8817066550254822},"editors":["Uri-ka"],"editorAvatarUrls":["/avatars/22f6463216904fb0ec8306e704432ab7.svg"],"reactions":[],"isReport":false}},{"id":"6a7602f1512a99216012848f","author":{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","fullname":"Urro","name":"urroxyz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":17,"isUserFollowing":false},"createdAt":"2026-08-07T16:08:17.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Great work!","html":"<p>Great work!</p>\n","updatedAt":"2026-08-07T16:08:17.119Z","author":{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","fullname":"Urro","name":"urroxyz","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":17,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8486101031303406},"editors":["urroxyz"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png"],"reactions":[{"reaction":"❤️","users":["Uri-ka"],"count":1}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05850","authors":[{"_id":"6a75aeb2e1228e04b3238457","user":{"_id":"604e1c120fe8ff3ec13d71e8","avatarUrl":"/avatars/22f6463216904fb0ec8306e704432ab7.svg","isPro":false,"fullname":"Uri Katz","user":"Uri-ka","type":"user","name":"Uri-ka"},"name":"Uri Katz","status":"claimed_verified","statusLastChangedAt":"2026-08-07T16:45:28.190Z","hidden":false},{"_id":"6a75aeb2e1228e04b3238458","user":{"_id":"668418b8e2e0c3b2150efeed","avatarUrl":"/avatars/8eaa0f3d112e5850e2ae383f269c5734.svg","isPro":false,"fullname":"omer goldman","user":"omergoldman","type":"user","name":"omergoldman"},"name":"Omer Goldman","status":"claimed_verified","statusLastChangedAt":"2026-08-07T16:45:28.197Z","hidden":false},{"_id":"6a75aeb2e1228e04b3238459","name":"Tomasz Limisiewicz","hidden":false},{"_id":"6a75aeb2e1228e04b323845a","name":"Reut Tsarfaty","hidden":false},{"_id":"6a75aeb2e1228e04b323845b","name":"Noah A. Smith","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"MameLoshnLM: Yiddish Language Model and Evaluation Benchmark","submittedOnDailyBy":{"_id":"604e1c120fe8ff3ec13d71e8","avatarUrl":"/avatars/22f6463216904fb0ec8306e704432ab7.svg","isPro":false,"fullname":"Uri Katz","user":"Uri-ka","type":"user","name":"Uri-ka"},"summary":"We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.","upvotes":9,"discussionId":"6a75aeb3e1228e04b323845c","githubRepo":"https://github.com/katzurik/MameLoshnLM","githubRepoAddedBy":"user","githubStars":0,"organization":{"_id":"6a74e99f1ae610f78578687f","name":"Yiddish-NLP","fullname":"Yiddish-NLP","avatar":"https://www.gravatar.com/avatar/4d41ae32f4ecd6f52ee2ffbe7a72b864?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"609eb1fc1172dedeac2200db","avatarUrl":"/avatars/a7384a8d3389610b38388c100a28c86d.svg","isPro":false,"fullname":"H","user":"Eran","type":"user"},{"_id":"6033d0681f993496bc14d9eb","avatarUrl":"/avatars/1c0dfc1fe7b62a78b76ec4f9e6b40b24.svg","isPro":false,"fullname":"Avi Caciularu","user":"codevan","type":"user"},{"_id":"622f35a2bc2a392eaf21b3e7","avatarUrl":"/avatars/383409ebd912ba90d8e7966e61a3910d.svg","isPro":false,"fullname":"Mosh Levy","user":"Mosh","type":"user"},{"_id":"62a7581cf049be35252a2e7c","avatarUrl":"/avatars/91de4eb48f51bfd6e028c08ccfa98f8c.svg","isPro":false,"fullname":"Royi Rassin","user":"Royir","type":"user"},{"_id":"668418b8e2e0c3b2150efeed","avatarUrl":"/avatars/8eaa0f3d112e5850e2ae383f269c5734.svg","isPro":false,"fullname":"omer goldman","user":"omergoldman","type":"user"},{"_id":"675b1c35c7d5fefd93391a31","avatarUrl":"/avatars/9b3ce9acb3ba134b9ca3d65931ef9e51.svg","isPro":false,"fullname":"Uriel Dolev","user":"udolev","type":"user"},{"_id":"604e1c120fe8ff3ec13d71e8","avatarUrl":"/avatars/22f6463216904fb0ec8306e704432ab7.svg","isPro":false,"fullname":"Uri Katz","user":"Uri-ka","type":"user"},{"_id":"661ab1f1fa3b144a381fa454","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661ab1f1fa3b144a381fa454/IlpZBb9NCjo7ntFwMIH53.png","isPro":false,"fullname":"Urro","user":"urroxyz","type":"user"},{"_id":"61e022d076edd24da0a07788","avatarUrl":"/avatars/584a9db003775823e6b7f280645cbd6d.svg","isPro":false,"fullname":"Pavel Larionov","user":"Pavelrst","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a74e99f1ae610f78578687f","name":"Yiddish-NLP","fullname":"Yiddish-NLP","avatar":"https://www.gravatar.com/avatar/4d41ae32f4ecd6f52ee2ffbe7a72b864?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05850.md","query":{}}">
Papers
arxiv:2608.05850

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

Published on Aug 6
· Submitted by
Uri Katz
on Aug 7
Authors:

Abstract

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

Community

Paper author Paper submitter about 7 hours ago

Yiddish carries centuries of literature, humor, and culture—yet it has almost no modern NLP infrastructure.

With MameloshnLM, we introduce the first dedicated Yiddish language model and the first Yiddish LLM benchmark. We show that:

  • Multilingual web corpora contain substantially more noise and machine-translated text than native sources.

  • Models trained on this data flatten the language, weakening its idioms, style, and cultural voice.

  • Training on authentic Yiddish enables MameloshnLM to better preserve the language’s distinctive character.

Supporting a language means more than generating its words—it means ensuring it still sounds like itself.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.05850
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.05850 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.05850 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.05850 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers