On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.</p>\n","updatedAt":"2026-07-31T05:28:21.696Z","author":{"_id":"650b0d66664f7b7d088ca281","avatarUrl":"/avatars/fce475c301f53e166fc3c8f5c5112c4a.svg","fullname":"Yi-Cheng Lin","name":"dlion168","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8938990235328674},"editors":["dlion168"],"editorAvatarUrls":["/avatars/fce475c301f53e166fc3c8f5c5112c4a.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.25289","authors":[{"_id":"6a6c3247202e2d9e3ffdb7fd","name":"Yuqi Li","hidden":false},{"_id":"6a6c3247202e2d9e3ffdb7fe","user":{"_id":"650b0d66664f7b7d088ca281","avatarUrl":"/avatars/fce475c301f53e166fc3c8f5c5112c4a.svg","isPro":false,"fullname":"Yi-Cheng Lin","user":"dlion168","type":"user","name":"dlion168"},"name":"Yi-Cheng Lin","status":"claimed_verified","statusLastChangedAt":"2026-07-31T08:45:05.636Z","hidden":false},{"_id":"6a6c3247202e2d9e3ffdb7ff","name":"Xianglong Wang","hidden":false},{"_id":"6a6c3247202e2d9e3ffdb800","name":"Kuo Yang","hidden":false},{"_id":"6a6c3247202e2d9e3ffdb801","name":"Xiaoqin Feng","hidden":false},{"_id":"6a6c3247202e2d9e3ffdb802","name":"Yixuan Wang","hidden":false},{"_id":"6a6c3247202e2d9e3ffdb803","name":"Huiran Duan","hidden":false},{"_id":"6a6c3247202e2d9e3ffdb804","name":"Yingli Tian","hidden":false}],"publishedAt":"2026-07-28T00:00:00.000Z","submittedOnDailyAt":"2026-07-31T00:00:00.000Z","title":"AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition","submittedOnDailyBy":{"_id":"650b0d66664f7b7d088ca281","avatarUrl":"/avatars/fce475c301f53e166fc3c8f5c5112c4a.svg","isPro":false,"fullname":"Yi-Cheng Lin","user":"dlion168","type":"user","name":"dlion168"},"summary":"On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.","upvotes":0,"discussionId":"6a6c3247202e2d9e3ffdb805"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[],"acceptLanguages":["en"],"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.25289.md","query":{}}">
AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
Abstract
On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.
Community
On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.25289 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.25289 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.25289 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.