Multilingual moral reasoning has 2 gaps: 1) evaluation datasets that rely on direct translation instead of cultural adaptation, and 2) no adaptive, theory-grounded reasoning methods. We address these with MCLASH (a new evaluation dataset), MET (a theory-grounded reasoning method), and MET-D (adding self-distillation on top of MET).</p>\n","updatedAt":"2026-07-14T16:20:31.701Z","author":{"_id":"66d079d94b3b38cefaf1dc4e","avatarUrl":"/avatars/f680a19fb7b8e52c811eb6df218a2cea.svg","fullname":"Ayoung Lee","name":"Ayoung01","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9073209762573242},"editors":["Ayoung01"],"editorAvatarUrls":["/avatars/f680a19fb7b8e52c811eb6df218a2cea.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.11736","authors":[{"_id":"6a5661b0a9d74d6e65bbded6","name":"Ayoung Lee","hidden":false},{"_id":"6a5661b0a9d74d6e65bbded7","name":"Ryan Kwon","hidden":false},{"_id":"6a5661b0a9d74d6e65bbded8","name":"Yunxiang Zhang","hidden":false},{"_id":"6a5661b0a9d74d6e65bbded9","name":"Yuxuan Liu","hidden":false},{"_id":"6a5661b0a9d74d6e65bbdeda","name":"Peter Railton","hidden":false},{"_id":"6a5661b0a9d74d6e65bbdedb","name":"Lu Wang","hidden":false}],"publishedAt":"2026-07-13T15:59:25.000Z","submittedOnDailyAt":"2026-07-14T00:00:00.000Z","title":"MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning","submittedOnDailyBy":{"_id":"66d079d94b3b38cefaf1dc4e","avatarUrl":"/avatars/f680a19fb7b8e52c811eb6df218a2cea.svg","isPro":false,"fullname":"Ayoung Lee","user":"Ayoung01","type":"user","name":"Ayoung01"},"summary":"Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.","upvotes":4,"discussionId":"6a5661b0a9d74d6e65bbdedc","projectPage":"https://huggingface.co/collections/launch/met","organization":{"_id":"626b5c0f05fe1cb657258316","name":"launch","fullname":"LAUNCH Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lC3Kp0o9uNrzdgKNX-Wq0.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"626b5b8b822e3d85324cfa61","avatarUrl":"/avatars/338083cbaa6ee7e62019566ba7ec7ff7.svg","isPro":false,"fullname":"Lu Wang","user":"wangluxy","type":"user"},{"_id":"63239d3746247071271d5d5f","avatarUrl":"/avatars/b1ab3db4934868e545e1744e45fbec87.svg","isPro":false,"fullname":"Xin Liu","user":"xinliucs","type":"user"},{"_id":"65a6015a636afd03b238c6ad","avatarUrl":"/avatars/4eb76ec30653b51e3540ce0d9fa23e89.svg","isPro":false,"fullname":"Inderjeet Nair","user":"inderjeetnair1","type":"user"},{"_id":"63ff42bfcd242b2620a022fa","avatarUrl":"/avatars/e4f29ba9ce6243f6ef9a3e391c8af4d1.svg","isPro":false,"fullname":"Wenquan Lu","user":"daviddavidlu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"626b5c0f05fe1cb657258316","name":"launch","fullname":"LAUNCH Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/lC3Kp0o9uNrzdgKNX-Wq0.png"},"query":{}}">
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
Abstract
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.
Community
Multilingual moral reasoning has 2 gaps: 1) evaluation datasets that rely on direct translation instead of cultural adaptation, and 2) no adaptive, theory-grounded reasoning methods. We address these with MCLASH (a new evaluation dataset), MET (a theory-grounded reasoning method), and MET-D (adding self-distillation on top of MET).
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.11736 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.