Hugging Face Daily Papers · · 3 min read

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Project website: <a href=\"https://xpeng-ai.github.io/x-aut\" rel=\"nofollow\">xpeng-ai.github.io/x-aut</a><br>Paper: <a href=\"https://arxiv.org/abs/2609.11412\" rel=\"nofollow\">arXiv:2609.11412</a><br>Source code: <a href=\"https://github.com/XPENG-AI/X-AuT\" rel=\"nofollow\">XPENG-AI/X-AuT</a><br>Model weights: <a href=\"https://huggingface.co/XPENG-AI/X-AuT\">XPENG-AI/X-AuT</a></p>\n","updatedAt":"2026-09-11T06:28:52.394Z","author":{"_id":"6406db5cd684369027166986","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406db5cd684369027166986/Zl-orrGcbY0RbfjfKszn1.jpeg","fullname":"Shiyu Huang","name":"ShiyuHuang","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":9,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.46474361419677734},"editors":["ShiyuHuang"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6406db5cd684369027166986/Zl-orrGcbY0RbfjfKszn1.jpeg"],"reactions":[{"reaction":"🚀","users":["Elgin-zhj","ShiyuHuang","Mrhamsterleo","Skywalker0410"],"count":4}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.11412","authors":[{"_id":"6aa397d547a406da7901e805","name":"Haojun Zhang","hidden":false},{"_id":"6aa397d547a406da7901e806","name":"Yi Zou","hidden":false},{"_id":"6aa397d547a406da7901e807","name":"Min Chen","hidden":false},{"_id":"6aa397d547a406da7901e808","user":{"_id":"692fa5d17ff1da99eb783dfb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/692fa5d17ff1da99eb783dfb/GZ2TD9ua-pigxz5kEHRYX.jpeg","isPro":false,"fullname":"Qize Yu","user":"Skywalker0410","type":"user","name":"Skywalker0410"},"name":"Qize Yu","status":"claimed_verified","statusLastChangedAt":"2026-09-11T08:45:04.771Z","hidden":false},{"_id":"6aa397d547a406da7901e809","name":"Lianrui Fan","hidden":false},{"_id":"6aa397d547a406da7901e80a","name":"Xini Ding","hidden":false},{"_id":"6aa397d547a406da7901e80b","name":"Hao Li","hidden":false},{"_id":"6aa397d547a406da7901e80c","name":"Shuchang Zhou","hidden":false},{"_id":"6aa397d547a406da7901e80d","name":"Xianming Liu","hidden":false},{"_id":"6aa397d547a406da7901e80e","user":{"_id":"6406db5cd684369027166986","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406db5cd684369027166986/Zl-orrGcbY0RbfjfKszn1.jpeg","isPro":false,"fullname":"Shiyu Huang","user":"ShiyuHuang","type":"user","name":"ShiyuHuang"},"name":"Shiyu Huang","status":"claimed_verified","statusLastChangedAt":"2026-09-11T08:45:04.765Z","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6406db5cd684369027166986/LqtMDn_MJsPpxMH73nkD0.png"],"publishedAt":"2026-09-10T00:00:00.000Z","submittedOnDailyAt":"2026-09-11T00:00:00.000Z","title":"X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation","submittedOnDailyBy":{"_id":"6406db5cd684369027166986","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406db5cd684369027166986/Zl-orrGcbY0RbfjfKszn1.jpeg","isPro":false,"fullname":"Shiyu Huang","user":"ShiyuHuang","type":"user","name":"ShiyuHuang"},"summary":"Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese--English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 18rightarrow14 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: https://xpeng-ai.github.io/x-aut","upvotes":12,"discussionId":"6aa397d647a406da7901e80f","projectPage":"https://xpeng-ai.github.io/x-aut","githubRepo":"https://github.com/XPENG-AI/X-AuT","githubRepoAddedBy":"user","ai_summary":"X-AuT progressively prunes audio-encoder layers in speech large language models and restores accuracy via behavioral probes, representation alignment, cross-scale distillation, and LoRA adaptation.","ai_keywords":["speech large language models","audio-encoder depth","X-AuT","behavioral probes","representation alignment","cross-scale distillation","LoRA finetuning","attention LoRA adapters","transcript-consistency pipeline","source reweighting"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6aa24a85fc727b9188a95572","name":"XPENG-AI","fullname":"XPENG AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406db5cd684369027166986/DAJtDU_HssW_R7WIw0USp.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6406db5cd684369027166986","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406db5cd684369027166986/Zl-orrGcbY0RbfjfKszn1.jpeg","isPro":false,"fullname":"Shiyu Huang","user":"ShiyuHuang","type":"user"},{"_id":"676e697e59181559a79f74f8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/676e697e59181559a79f74f8/XKpitiSDq6OOOiLMW4Vto.png","isPro":false,"fullname":"Elgin","user":"Elgin-zhj","type":"user"},{"_id":"6a8d9b698485c871798e8e9c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a8d9b698485c871798e8e9c/5fDc8SXAKDRb_1OHNHWxv.jpeg","isPro":false,"fullname":"chenmin","user":"chenminupup","type":"user"},{"_id":"69ddb91498acc5881fb04cb9","avatarUrl":"/avatars/c1274814f24432fd1d4458c1bb9e8d84.svg","isPro":false,"fullname":"dd","user":"ashe77","type":"user"},{"_id":"64265104a5ec4a5cbc52f741","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64265104a5ec4a5cbc52f741/M2vj92Jj6ZdC4PQ6D_v-q.jpeg","isPro":false,"fullname":"LR","user":"FanLR","type":"user"},{"_id":"695f7dc97fadebd82c2f168e","avatarUrl":"/avatars/9ae7fe85cec83187ea1798fb986917db.svg","isPro":false,"fullname":"Hao Li","user":"Mrhamsterleo","type":"user"},{"_id":"692fa5d17ff1da99eb783dfb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/692fa5d17ff1da99eb783dfb/GZ2TD9ua-pigxz5kEHRYX.jpeg","isPro":false,"fullname":"Qize Yu","user":"Skywalker0410","type":"user"},{"_id":"68552d0530337926c42c1871","avatarUrl":"/avatars/4f457e02c3333ed1f9d5c59e7a464124.svg","isPro":false,"fullname":"Zetian Song","user":"pkuvcl","type":"user"},{"_id":"67068275300b431e16cd24d7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67068275300b431e16cd24d7/Rv0mMgvDh9TDQp8oLqNLg.jpeg","isPro":false,"fullname":"Yuran Wang","user":"wayrise","type":"user"},{"_id":"682239af00f78bd046f3817b","avatarUrl":"/avatars/8f8b752aaf66986cced927dc960ade7f.svg","isPro":false,"fullname":"Jiaqi Liang","user":"SCZZ","type":"user"},{"_id":"67cc5e0232aeea9209d35033","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/67cc5e0232aeea9209d35033/UFMuOu_GfHecZKJk1s1UL.jpeg","isPro":false,"fullname":"YueChen","user":"YueChen0614","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6aa24a85fc727b9188a95572","name":"XPENG-AI","fullname":"XPENG AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6406db5cd684369027166986/DAJtDU_HssW_R7WIw0USp.jpeg"},"query":{}}">
Papers
arxiv:2609.11412

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

Published on Sep 10
· Submitted by
Shiyu Huang
on Sep 11
Authors:
,

Abstract

X-AuT progressively prunes audio-encoder layers in speech large language models and restores accuracy via behavioral probes, representation alignment, cross-scale distillation, and LoRA adaptation.

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese--English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 18rightarrow14 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: https://xpeng-ai.github.io/x-aut

Community

Paper author Paper submitter about 8 hours ago

Project website: xpeng-ai.github.io/x-aut
Paper: arXiv:2609.11412
Source code: XPENG-AI/X-AuT
Model weights: XPENG-AI/X-AuT

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.11412 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers