Hugging Face Daily Papers · · 3 min read

Unified Audio Intelligence Without Regressing on Text Intelligence

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Tweet: <a href=\"https://x.com/HuggingPapers/status/2074384562952749254\" rel=\"nofollow\">https://x.com/HuggingPapers/status/2074384562952749254</a></p>\n","updatedAt":"2026-07-07T13:16:34.708Z","author":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","fullname":"Niels Rogge","name":"nielsr","type":"user","isPro":false,"isHf":true,"isHfAdmin":false,"isMod":false,"followerCount":1249,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.463441401720047},"editors":["nielsr"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.05196","authors":[{"_id":"6a4c700225849b193a8340d8","name":"Zhifeng Kong","hidden":false},{"_id":"6a4c700225849b193a8340d9","user":{"_id":"628dba7ea13c4b25ff786ec6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671164940704-628dba7ea13c4b25ff786ec6.jpeg","isPro":false,"fullname":"Sang-gil Lee","user":"L0SG","type":"user","name":"L0SG"},"name":"Sang-gil Lee","status":"claimed_verified","statusLastChangedAt":"2026-07-07T12:11:42.973Z","hidden":false},{"_id":"6a4c700225849b193a8340da","name":"Jaehyeon Kim","hidden":false},{"_id":"6a4c700225849b193a8340db","name":"Boxin Wang","hidden":false},{"_id":"6a4c700225849b193a8340dc","name":"Zihan Liu","hidden":false},{"_id":"6a4c700225849b193a8340dd","name":"Sungwon Kim","hidden":false},{"_id":"6a4c700225849b193a8340de","name":"Yang Chen","hidden":false},{"_id":"6a4c700225849b193a8340df","name":"Arushi Goel","hidden":false},{"_id":"6a4c700225849b193a8340e0","name":"Rajarshi Roy","hidden":false},{"_id":"6a4c700225849b193a8340e1","name":"Wenliang Dai","hidden":false},{"_id":"6a4c700225849b193a8340e2","name":"Zhuolin Yang","hidden":false},{"_id":"6a4c700225849b193a8340e3","name":"Yangyi Chen","hidden":false},{"_id":"6a4c700225849b193a8340e4","name":"Dongfu Jiang","hidden":false},{"_id":"6a4c700225849b193a8340e5","name":"Sreyan Ghosh","hidden":false},{"_id":"6a4c700225849b193a8340e6","name":"Tuomas Rintamaki","hidden":false},{"_id":"6a4c700225849b193a8340e7","name":"Andrew Tao","hidden":false},{"_id":"6a4c700225849b193a8340e8","name":"Jonathan Raiman","hidden":false},{"_id":"6a4c700225849b193a8340e9","name":"Mohammad Shoeybi","hidden":false},{"_id":"6a4c700225849b193a8340ea","name":"Bryan Catanzaro","hidden":false},{"_id":"6a4c700225849b193a8340eb","name":"Wei Ping","hidden":false}],"publishedAt":"2026-07-06T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"Unified Audio Intelligence Without Regressing on Text Intelligence","submittedOnDailyBy":{"_id":"5f1158120c833276f61f1a84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1608042047613-5f1158120c833276f61f1a84.jpeg","isPro":false,"fullname":"Niels Rogge","user":"nielsr","type":"user","name":"nielsr"},"summary":"Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.","upvotes":7,"discussionId":"6a4c700225849b193a8340ec","ai_summary":"A unified audio-text large language model is presented that integrates audio and text processing through a shared transformer decoder, achieving superior performance across multiple audio and speech tasks while maintaining strong text reasoning capabilities.","ai_keywords":["audio-text LLM","Transformer decoder","multimodal generation","supervised training","Cascade RL","on-policy distillation","text-to-speech","speech-to-speech generation","audio understanding","speech recognition","audio generation"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"663ee43bfeeb49803537da98","avatarUrl":"/avatars/17c3e9c435cc36fb04b4589e6176a243.svg","isPro":false,"fullname":"Wei Ping","user":"wping","type":"user"},{"_id":"628dba7ea13c4b25ff786ec6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1671164940704-628dba7ea13c4b25ff786ec6.jpeg","isPro":false,"fullname":"Sang-gil Lee","user":"L0SG","type":"user"},{"_id":"652a4dfc36f031c5e6f8b8a6","avatarUrl":"/avatars/9fb56b025dc25f91ca6c31136eaf74b2.svg","isPro":false,"fullname":"Zhifeng Kong","user":"ZhifengKong","type":"user"},{"_id":"69d08207e192e1734ff2a335","avatarUrl":"/avatars/6417436c8fcca47d67402a96511b5603.svg","isPro":false,"fullname":"kyosuke yanagisawa","user":"yyykyosuke","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6a159e2e31430965c653cd91","avatarUrl":"/avatars/3e2a5e959cae38e6ab0c90af59aadf37.svg","isPro":false,"fullname":"张千怡","user":"ISAACGRE","type":"user"},{"_id":"62bc9d90e81dfd65cced9316","avatarUrl":"/avatars/05df14cd1fdbc7d6a80d2960a05a94f0.svg","isPro":false,"fullname":"Yang Chen","user":"ychenNLP","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"},"query":{}}">
Papers
arxiv:2607.05196

Unified Audio Intelligence Without Regressing on Text Intelligence

Published on Jul 6
· Submitted by
Niels Rogge
on Jul 7
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

A unified audio-text large language model is presented that integrates audio and text processing through a shared transformer decoder, achieving superior performance across multiple audio and speech tasks while maintaining strong text reasoning capabilities.

Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper 2

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.05196 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 1

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers