Hugging Face Daily Papers · · 4 min read

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

NOLLI is a procedurally generated, behaviorally calibrated English–Korean puzzle benchmark that finds small matched presentation-language gaps and localizes candidate bottlenecks in multi-step Hangul-jamo execution and Korean kinship terminology.</p>\n","updatedAt":"2026-08-06T02:21:49.210Z","author":{"_id":"66120647cac232c1507e13da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66120647cac232c1507e13da/iOBrsyr2mVdVLAFF6pZt6.png","fullname":"DasolChoi","name":"Dasool","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8692439198493958},"editors":["Dasool"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66120647cac232c1507e13da/iOBrsyr2mVdVLAFF6pZt6.png"],"reactions":[],"isReport":false}},{"id":"6a741a6cb8aa10a39f2c58e5","author":{"_id":"60d3e619b8448e1785bbda2a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d3e619b8448e1785bbda2a/q2re5u1HNwsCCyIMtid_I.jpeg","fullname":"GUIJIN SON","name":"amphora","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":86,"isUserFollowing":false},"createdAt":"2026-08-06T05:23:56.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"link to dataset > https://huggingface.co/datasets/HAERAE-HUB/NOLLI","html":"<p>link to dataset &gt; <a href=\"https://huggingface.co/datasets/HAERAE-HUB/NOLLI\">https://huggingface.co/datasets/HAERAE-HUB/NOLLI</a></p>\n","updatedAt":"2026-08-06T05:23:56.386Z","author":{"_id":"60d3e619b8448e1785bbda2a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d3e619b8448e1785bbda2a/q2re5u1HNwsCCyIMtid_I.jpeg","fullname":"GUIJIN SON","name":"amphora","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":86,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5584570169448853},"editors":["amphora"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/60d3e619b8448e1785bbda2a/q2re5u1HNwsCCyIMtid_I.jpeg"],"reactions":[],"isReport":false}},{"id":"6a742086cc8a921228bf148c","author":{"_id":"6752cc1a10576e69f9bdc542","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6752cc1a10576e69f9bdc542/WkFgo6vx07H6IVLRZmFO_.jpeg","fullname":"Chanyoung Kim","name":"chanyoungkim","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-08-06T05:49:58.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"great dataset","html":"<p>great dataset</p>\n","updatedAt":"2026-08-06T05:49:58.277Z","author":{"_id":"6752cc1a10576e69f9bdc542","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6752cc1a10576e69f9bdc542/WkFgo6vx07H6IVLRZmFO_.jpeg","fullname":"Chanyoung Kim","name":"chanyoungkim","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3934417963027954},"editors":["chanyoungkim"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6752cc1a10576e69f9bdc542/WkFgo6vx07H6IVLRZmFO_.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.04397","authors":[{"_id":"6a73ef1ec5e410d076869a65","name":"Dasol Choi","hidden":false},{"_id":"6a73ef1ec5e410d076869a66","name":"Joonyong Park","hidden":false},{"_id":"6a73ef1ec5e410d076869a67","name":"Daegon Yu","hidden":false},{"_id":"6a73ef1ec5e410d076869a68","name":"Soo Yong Kim","hidden":false},{"_id":"6a73ef1ec5e410d076869a69","name":"Youngsook Song","hidden":false},{"_id":"6a73ef1ec5e410d076869a6a","name":"Seunghyeok Hong","hidden":false}],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap","submittedOnDailyBy":{"_id":"66120647cac232c1507e13da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66120647cac232c1507e13da/iOBrsyr2mVdVLAFF6pZt6.png","isPro":false,"fullname":"DasolChoi","user":"Dasool","type":"user","name":"Dasool"},"summary":"We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.","upvotes":15,"discussionId":"6a73ef1ec5e410d076869a6b","githubRepo":"https://github.com/HAE-RAE/NOLLI","githubRepoAddedBy":"user","githubStars":2,"organization":{"_id":"645ae5e85e6871b4b2d6bd80","name":"HAERAE-HUB","fullname":"HAE-RAE","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d3e619b8448e1785bbda2a/zasfyk1U_yRBgrluB6ggc.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66120647cac232c1507e13da","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66120647cac232c1507e13da/iOBrsyr2mVdVLAFF6pZt6.png","isPro":false,"fullname":"DasolChoi","user":"Dasool","type":"user"},{"_id":"67aa9e215716d8c0207eab19","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67aa9e215716d8c0207eab19/GwkmPHV4EHikbEIzcRxth.jpeg","isPro":false,"fullname":"Joonyong Park","user":"JoonYong-Park","type":"user"},{"_id":"64fac8a88d50404bc4ec66d4","avatarUrl":"/avatars/c424a0ef2187457e563c697cd9355a9e.svg","isPro":false,"fullname":"kim","user":"ksyint","type":"user"},{"_id":"665ec951b246e4e792ccf89d","avatarUrl":"/avatars/34043b930ab46426de22bce3e3eecbab.svg","isPro":false,"fullname":"kyubeenhan","user":"kyubeen","type":"user"},{"_id":"662c3305982b04e6a0b354b4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/662c3305982b04e6a0b354b4/QIcRfgGIefBVj9COuWEFu.jpeg","isPro":false,"fullname":"Seunghyeok Hong","user":"shongdr","type":"user"},{"_id":"60d3e619b8448e1785bbda2a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d3e619b8448e1785bbda2a/q2re5u1HNwsCCyIMtid_I.jpeg","isPro":true,"fullname":"GUIJIN SON","user":"amphora","type":"user"},{"_id":"65a9a5c450f85a33ed7f9ac8","avatarUrl":"/avatars/dc24a06cf8d89f4449ed82ba9ff518d9.svg","isPro":false,"fullname":"YuJinOh","user":"Yujin567","type":"user"},{"_id":"66616ee3017d224fd1aa6bce","avatarUrl":"/avatars/c5fc56aa49810bfbbe14f51fb69cf289.svg","isPro":false,"fullname":"saysimple","user":"saysimple0828","type":"user"},{"_id":"6a1ad2feb4238bb17fccfd9c","avatarUrl":"/avatars/27a9d02fe0bb8d669d8101c6c43734d2.svg","isPro":false,"fullname":"Sungjeh Yoon","user":"mujuontheaux","type":"user"},{"_id":"631c386bc73939ffc0716a37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662793811119-noauth.jpeg","isPro":false,"fullname":"SeongWan Kim","user":"idgmatrix","type":"user"},{"_id":"6752cc1a10576e69f9bdc542","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6752cc1a10576e69f9bdc542/WkFgo6vx07H6IVLRZmFO_.jpeg","isPro":false,"fullname":"Chanyoung Kim","user":"chanyoungkim","type":"user"},{"_id":"6a6aa0a8b172d8c070b6b478","avatarUrl":"/avatars/a3e77445bdf99a9d00175ac942c99e81.svg","isPro":false,"fullname":"Edward Wilson","user":"Indigo-Flow","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"645ae5e85e6871b4b2d6bd80","name":"HAERAE-HUB","fullname":"HAE-RAE","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d3e619b8448e1785bbda2a/zasfyk1U_yRBgrluB6ggc.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.04397.md","query":{}}">
Papers
arxiv:2608.04397

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

Published on Aug 5
· Submitted by
DasolChoi
on Aug 6
Authors:
,

Abstract

We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.

Community

Paper submitter about 8 hours ago

NOLLI is a procedurally generated, behaviorally calibrated English–Korean puzzle benchmark that finds small matched presentation-language gaps and localizes candidate bottlenecks in multi-step Hangul-jamo execution and Korean kinship terminology.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.04397
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.04397 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.04397 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers