Hugging Face Daily Papers · · 3 min read

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

The first unified multimodal emotion reasoning model capable of handling sentiment analysis, basic emotion recognition, open-vocabulary emotion detection, intent identification, sarcasm understanding, humor comprehension, empathetic response, and emotional support dialogue.</p>\n","updatedAt":"2026-08-10T09:48:14.071Z","author":{"_id":"66953e2f417bbfcd51933cd0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg","fullname":"Jiahao Huang","name":"Jiaha0Hu4ng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7565445303916931},"editors":["Jiaha0Hu4ng"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06013","authors":[{"_id":"6a757a07e1228e04b3238311","user":{"_id":"66953e2f417bbfcd51933cd0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg","isPro":false,"fullname":"Jiahao Huang","user":"Jiaha0Hu4ng","type":"user","name":"Jiaha0Hu4ng"},"name":"Jiahao Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-08T16:45:04.813Z","hidden":false},{"_id":"6a757a07e1228e04b3238312","name":"Zheng Lian","hidden":false},{"_id":"6a757a07e1228e04b3238313","name":"Jingyi Zhang","hidden":false},{"_id":"6a757a07e1228e04b3238314","name":"Zhide Chen","hidden":false},{"_id":"6a757a07e1228e04b3238315","name":"Xiaojiang Peng","hidden":false},{"_id":"6a757a07e1228e04b3238316","name":"Shaonan Wang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66953e2f417bbfcd51933cd0/Vidy8cpHc8EbaZzCBhO1-.png"],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction","submittedOnDailyBy":{"_id":"66953e2f417bbfcd51933cd0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg","isPro":false,"fullname":"Jiahao Huang","user":"Jiaha0Hu4ng","type":"user","name":"Jiaha0Hu4ng"},"summary":"Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.","upvotes":0,"discussionId":"6a757a07e1228e04b3238317","githubRepo":"https://github.com/waHAHJIAHAO/OneEmo","githubRepoAddedBy":"user","githubStars":5},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06013.md","query":{}}">
Papers
arxiv:2608.06013

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

Published on Aug 6
· Submitted by
Jiahao Huang
on Aug 10
Authors:

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.

Community

Paper author Paper submitter about 1 hour ago

The first unified multimodal emotion reasoning model capable of handling sentiment analysis, basic emotion recognition, open-vocabulary emotion detection, intent identification, sarcasm understanding, humor comprehension, empathetic response, and emotional support dialogue.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.06013
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers