Hugging Face Daily Papers · · 2 min read

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Open-ended bias evaluation for LALM.</p>\n","updatedAt":"2026-07-08T17:51:17.802Z","author":{"_id":"650b0d66664f7b7d088ca281","avatarUrl":"/avatars/fce475c301f53e166fc3c8f5c5112c4a.svg","fullname":"Yi-Cheng Lin","name":"dlion168","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8418886661529541},"editors":["dlion168"],"editorAvatarUrls":["/avatars/fce475c301f53e166fc3c8f5c5112c4a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2604.17248","authors":[{"_id":"6a4e8dff48d70828b718dd0a","name":"Yi-Cheng Lin","hidden":false},{"_id":"6a4e8dff48d70828b718dd0b","name":"Yusuke Hirota","hidden":false},{"_id":"6a4e8dff48d70828b718dd0c","name":"Sung-Feng Huang","hidden":false},{"_id":"6a4e8dff48d70828b718dd0d","name":"Hung-yi Lee","hidden":false}],"publishedAt":"2026-07-03T00:00:00.000Z","submittedOnDailyAt":"2026-07-08T00:00:00.000Z","title":"VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech","submittedOnDailyBy":{"_id":"650b0d66664f7b7d088ca281","avatarUrl":"/avatars/fce475c301f53e166fc3c8f5c5112c4a.svg","isPro":false,"fullname":"Yi-Cheng Lin","user":"dlion168","type":"user","name":"dlion168"},"summary":"Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.","upvotes":0,"discussionId":"6a4e8dff48d70828b718dd0e","ai_summary":"Large Audio-Language Models exhibit systematic generative biases in realistic scenarios when evaluated through open-ended tasks using human-recorded speech, with bias magnitude varying significantly by task and triggered by gender and accent cues.","ai_keywords":["Large Audio-Language Models","generative bias","open-ended tasks","human-recorded speech","stereotypical associations","distributional shifts","task-dependent bias"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2604/2604.17248.md","query":{}}">
Papers
arxiv:2604.17248

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

Published on Jul 3
· Submitted by
Yi-Cheng Lin
on Jul 8
Authors:
,

Abstract

Large Audio-Language Models exhibit systematic generative biases in realistic scenarios when evaluated through open-ended tasks using human-recorded speech, with bias magnitude varying significantly by task and triggered by gender and accent cues.

Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing speech fairness benchmarks rely on synthetic speech and Multiple-Choice Questions (MCQs), both offering a fragmented view of fairness. We propose VIBE, a framework that evaluates generative bias through open-ended tasks such as personalized recommendations, using human-recorded speech. Unlike MCQs, our method allows stereotypical associations to manifest organically without predefined options, making it easily extensible to new tasks. Evaluating 12 state-of-the-art LALMs reveals systematic biases in realistic scenarios. Both gender and accent cues trigger statistically significant distributional shifts, and bias magnitude is strongly task-dependent.

Community

Paper submitter about 8 hours ago

Open-ended bias evaluation for LALM.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2604.17248
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2604.17248 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2604.17248 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2604.17248 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers