Inspired by cognitive science, we used Chinese idioms (or Chengyu) as a vehicle to design a cross-concept comprehension task as an anchor for assessing the creative ability of MLLMs, and we created a leaderboard.</p>\n","updatedAt":"2026-08-10T05:01:37.756Z","author":{"_id":"644378a96cea0db46dc96b39","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644378a96cea0db46dc96b39/wjGNaXjfhSdWr8kM6Amai.jpeg","fullname":"Ming Wang","name":"sci-m-wang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9539205431938171},"editors":["sci-m-wang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/644378a96cea0db46dc96b39/wjGNaXjfhSdWr8kM6Amai.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06501","authors":[{"_id":"6a7958fa8e9301703eaa5f71","user":{"_id":"644378a96cea0db46dc96b39","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644378a96cea0db46dc96b39/wjGNaXjfhSdWr8kM6Amai.jpeg","isPro":false,"fullname":"Ming Wang","user":"sci-m-wang","type":"user","name":"sci-m-wang"},"name":"Ming Wang","status":"claimed_verified","statusLastChangedAt":"2026-08-10T08:45:04.388Z","hidden":false},{"_id":"6a7958fa8e9301703eaa5f72","name":"Yuqing Zhang","hidden":false},{"_id":"6a7958fa8e9301703eaa5f73","name":"Tingna Xie","hidden":false},{"_id":"6a7958fa8e9301703eaa5f74","name":"Xiangju Li","hidden":false},{"_id":"6a7958fa8e9301703eaa5f75","name":"Xiaocui Yang","hidden":false},{"_id":"6a7958fa8e9301703eaa5f76","name":"Daling Wang","hidden":false},{"_id":"6a7958fa8e9301703eaa5f77","name":"Shi Feng","hidden":false},{"_id":"6a7958fa8e9301703eaa5f78","name":"Yifei Zhang","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding","submittedOnDailyBy":{"_id":"644378a96cea0db46dc96b39","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644378a96cea0db46dc96b39/wjGNaXjfhSdWr8kM6Amai.jpeg","isPro":false,"fullname":"Ming Wang","user":"sci-m-wang","type":"user","name":"sci-m-wang"},"summary":"Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.","upvotes":1,"discussionId":"6a7958fb8e9301703eaa5f79","organization":{"_id":"6a70927c8bfd8c53670018c1","name":"kinamind","fullname":"KinaMind","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/644378a96cea0db46dc96b39/waICbkYoR7UV7uT0q8SDK.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"644378a96cea0db46dc96b39","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644378a96cea0db46dc96b39/wjGNaXjfhSdWr8kM6Amai.jpeg","isPro":false,"fullname":"Ming Wang","user":"sci-m-wang","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a70927c8bfd8c53670018c1","name":"kinamind","fullname":"KinaMind","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/644378a96cea0db46dc96b39/waICbkYoR7UV7uT0q8SDK.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06501.md","query":{}}">
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
Abstract
Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.
Community
Inspired by cognitive science, we used Chinese idioms (or Chengyu) as a vehicle to design a cross-concept comprehension task as an anchor for assessing the creative ability of MLLMs, and we created a leaderboard.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.06501 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.06501 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.