Hugging Face Daily Papers · · 4 min read

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

.</p>\n","updatedAt":"2026-07-08T02:57:28.721Z","author":{"_id":"652066649004117947e46ed6","avatarUrl":"/avatars/972c97df6f26d2c3d6ce71ec579984bb.svg","fullname":"Jaehong Yoon","name":"jaehong31","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"fr","probability":0.32275810837745667},"editors":["jaehong31"],"editorAvatarUrls":["/avatars/972c97df6f26d2c3d6ce71ec579984bb.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2606.30026","authors":[{"_id":"6a43b502c8741d01182cfa9d","name":"Yuxuan Fan","hidden":false},{"_id":"6a43b502c8741d01182cfa9e","name":"Gyusik Seo","hidden":false},{"_id":"6a43b502c8741d01182cfa9f","name":"Jing Hao","hidden":false},{"_id":"6a43b502c8741d01182cfaa0","name":"Jaemin Cho","hidden":false},{"_id":"6a43b502c8741d01182cfaa1","name":"Mohit Bansal","hidden":false},{"_id":"6a43b502c8741d01182cfaa2","name":"Jaehong Yoon","hidden":false}],"publishedAt":"2026-06-29T00:00:00.000Z","submittedOnDailyAt":"2026-07-08T00:00:00.000Z","title":"MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs","submittedOnDailyBy":{"_id":"652066649004117947e46ed6","avatarUrl":"/avatars/972c97df6f26d2c3d6ce71ec579984bb.svg","isPro":false,"fullname":"Jaehong Yoon","user":"jaehong31","type":"user","name":"jaehong31"},"summary":"Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.","upvotes":3,"discussionId":"6a43b502c8741d01182cfaa3","projectPage":"https://musebench.github.io/","githubRepo":"https://github.com/musebench/musebench-code","githubRepoAddedBy":"user","ai_summary":"A comprehensive benchmark called Musebench is introduced to evaluate multimodal large language models on nuanced artistic understanding, revealing a significant gap between current models and human expert performance in creative domain expertise.","ai_keywords":["multimodal large language models","artistic understanding","creative intent","Musebench","visual arts","cinematic arts","stage performing arts","game arts","zero-shot evaluation","expert validation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":1,"organization":{"_id":"6371470aafbe42caa5a76208","name":"nanyang-technological-university-singapore","fullname":"Nanyang Technological University Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637146c5afbe42caa5a75e1b/sZyHSA1AQaAS4nrGan682.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"652066649004117947e46ed6","avatarUrl":"/avatars/972c97df6f26d2c3d6ce71ec579984bb.svg","isPro":false,"fullname":"Jaehong Yoon","user":"jaehong31","type":"user"},{"_id":"68eef67bbbfffc8550ecc524","avatarUrl":"/avatars/9d73ff2af5620d9db7bc77a9b45fdff4.svg","isPro":false,"fullname":"Gyusik Suh","user":"WillSuh","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6371470aafbe42caa5a76208","name":"nanyang-technological-university-singapore","fullname":"Nanyang Technological University Singapore","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/637146c5afbe42caa5a75e1b/sZyHSA1AQaAS4nrGan682.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2606/2606.30026.md","query":{}}">
Papers
arxiv:2606.30026

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

Published on Jun 29
· Submitted by
Jaehong Yoon
on Jul 8
Authors:
,

Abstract

A comprehensive benchmark called Musebench is introduced to evaluate multimodal large language models on nuanced artistic understanding, revealing a significant gap between current models and human expert performance in creative domain expertise.

Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinations of visual, auditory, and narrative elements (e.g., fear amplified through claustrophobic framing, or grief conveyed through silence and lingering close-ups). True artistic understanding extends beyond recognizing what is depicted to reasoning about why it is expressed through particular creative choices. Despite the strong progress of multimodal large language models (MLLMs), this critical aspect of artistic understanding remains underexplored, as existing benchmarks largely measure perceptual recognition while overlooking reasoning about creative intent. To address this gap, we introduce Musebench, a comprehensive benchmark designed to evaluate MLLMs on nuanced artistic understanding. It comprises 4,016 questions spanning cinematic arts, static visual arts, stage performing arts, and game arts, distilled from over 10K candidate video essays that pair professional commentary with visual demonstration. To capture the open-ended nature of artistic analysis at scale, the benchmark combines single-select and variable-option multi-select questions. All questions are generated and refined through a four-phase iterative pipeline combining shortcut filtering, adversarial distractors, and expert validation. Comprehensive zero-shot evaluation of 28 state-of-the-art MLLMs reveals that even the best-performing model achieves only 48.29% accuracy, substantially below human expert performance of 87.18%, exposing a significant gap in current models' creative domain expertise.

Community

Paper submitter about 14 hours ago
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.30026
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2606.30026 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2606.30026 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2606.30026 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers