The first unified multimodal emotion reasoning model capable of handling sentiment analysis, basic emotion recognition, open-vocabulary emotion detection, intent identification, sarcasm understanding, humor comprehension, empathetic response, and emotional support dialogue.</p>\n","updatedAt":"2026-08-10T09:48:14.071Z","author":{"_id":"66953e2f417bbfcd51933cd0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg","fullname":"Jiahao Huang","name":"Jiaha0Hu4ng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7565445303916931},"editors":["Jiaha0Hu4ng"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.06013","authors":[{"_id":"6a757a07e1228e04b3238311","user":{"_id":"66953e2f417bbfcd51933cd0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg","isPro":false,"fullname":"Jiahao Huang","user":"Jiaha0Hu4ng","type":"user","name":"Jiaha0Hu4ng"},"name":"Jiahao Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-08T16:45:04.813Z","hidden":false},{"_id":"6a757a07e1228e04b3238312","name":"Zheng Lian","hidden":false},{"_id":"6a757a07e1228e04b3238313","name":"Jingyi Zhang","hidden":false},{"_id":"6a757a07e1228e04b3238314","name":"Zhide Chen","hidden":false},{"_id":"6a757a07e1228e04b3238315","name":"Xiaojiang Peng","hidden":false},{"_id":"6a757a07e1228e04b3238316","name":"Shaonan Wang","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/66953e2f417bbfcd51933cd0/Vidy8cpHc8EbaZzCBhO1-.png"],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-10T00:00:00.000Z","title":"OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction","submittedOnDailyBy":{"_id":"66953e2f417bbfcd51933cd0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66953e2f417bbfcd51933cd0/z7mItS2fzqNO4Xi6Dy04_.jpeg","isPro":false,"fullname":"Jiahao Huang","user":"Jiaha0Hu4ng","type":"user","name":"Jiaha0Hu4ng"},"summary":"Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.","upvotes":0,"discussionId":"6a757a07e1228e04b3238317","githubRepo":"https://github.com/waHAHJIAHAO/OneEmo","githubRepoAddedBy":"user","githubStars":5},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.06013.md","query":{}}">
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.
Community
The first unified multimodal emotion reasoning model capable of handling sentiment analysis, basic emotion recognition, open-vocabulary emotion detection, intent identification, sarcasm understanding, humor comprehension, empathetic response, and emotional support dialogue.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.