<strong>Taste-aware music retrieval from audio embeddings</strong> 🎵🍫</p>\n<p>Can we predict how <em>sweet</em>, <em>bitter</em>, <em>salty</em>, <em>sour</em>, or <em>spicy</em> a piece of music sounds—using only its audio?</p>\n<p>In this work, we introduce a <strong>content-based music retrieval benchmark</strong> for taste prediction from audio. We compare <strong>10 pretrained audio encoders</strong> (HEAR families) under a unified frozen-encoder protocol and show that:</p>\n<p>• 🎯 Best models achieve <strong>0.134 RMSE</strong>, outperforming the previous state of the art (<strong>0.219 RMSE</strong>).<br>• 👥 On real music, predictions are <strong>closer to the consensus than an average human rater</strong> (0.13 vs. 0.28 RMSE).<br>• 🔎 The learned 5D taste space enables <strong>taste-based music retrieval</strong>, substantially outperforming a CLAP-text baseline.<br>• 🧠 We also provide <strong>psychophysics-grounded interpretability</strong>, linking learned representations to known sound–taste correspondences.</p>\n<p>This is, to our knowledge, the first benchmark framing <strong>taste-from-audio prediction as a music information retrieval task</strong>.</p>\n<p>Paper: <em>Taste-aware Music Retrieval from Audio Embeddings</em></p>\n<p>#MusicInformationRetrieval #AudioML #MultimodalAI #MachineLearning #HEAR #RepresentationLearning #Crossmodal #MusicAI</p>\n","updatedAt":"2026-07-07T11:33:01.430Z","author":{"_id":"65b00f6bc9a5a7680f531c1d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b00f6bc9a5a7680f531c1d/z4hzKR5tgdkft6c3Wxt9Y.jpeg","fullname":"Matteo Spanio","name":"matteospanio","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7442927360534668},"editors":["matteospanio"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65b00f6bc9a5a7680f531c1d/z4hzKR5tgdkft6c3Wxt9Y.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.03296","authors":[{"_id":"6a4c98f725849b193a834371","user":{"_id":"65b00f6bc9a5a7680f531c1d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b00f6bc9a5a7680f531c1d/z4hzKR5tgdkft6c3Wxt9Y.jpeg","isPro":false,"fullname":"Matteo Spanio","user":"matteospanio","type":"user","name":"matteospanio"},"name":"Matteo Spanio","status":"claimed_verified","statusLastChangedAt":"2026-07-07T08:30:24.545Z","hidden":false},{"_id":"6a4c98f725849b193a834372","name":"Antonio Rodà","hidden":false}],"publishedAt":"2026-07-03T00:00:00.000Z","submittedOnDailyAt":"2026-07-07T00:00:00.000Z","title":"Taste-aware music retrieval from audio embeddings","submittedOnDailyBy":{"_id":"65b00f6bc9a5a7680f531c1d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b00f6bc9a5a7680f531c1d/z4hzKR5tgdkft6c3Wxt9Y.jpeg","isPro":false,"fullname":"Matteo Spanio","user":"matteospanio","type":"user","name":"matteospanio"},"summary":"Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0.134; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE 0.13 vs. 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0.219). On absolute error the encoders are statistically flat, with a single VGGish matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson r 0.724 vs. 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.","upvotes":2,"discussionId":"6a4c98f725849b193a834373","githubRepo":"https://github.com/CSCPadova/wav2taste","githubRepoAddedBy":"user","ai_summary":"Audio encoders from HEAR families are evaluated for taste prediction, with gated late-fusion showing superior rank correlation and the best models achieving human-level accuracy on held-out music.","ai_keywords":["audio encoders","HEAR families","multi-task regression","gated late-fusion","RMSE","rank correlation","macro Pearson r","content-based retrieval","CLAP-text baseline","ridge probes","audio-bandstop knockout"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":0,"organization":{"_id":"65cbf5256938a6a81f488ba0","name":"csc-unipd","fullname":"Centro di Sonologia Computazionale","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b00f6bc9a5a7680f531c1d/I25ZGtU4dOq44Q2sERhPB.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65b00f6bc9a5a7680f531c1d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b00f6bc9a5a7680f531c1d/z4hzKR5tgdkft6c3Wxt9Y.jpeg","isPro":false,"fullname":"Matteo Spanio","user":"matteospanio","type":"user"},{"_id":"69bcebdbb0b4d685f7c1f22a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/avNCmL2FuQgxjNJQGZd3w.jpeg","isPro":false,"fullname":"한선우","user":"hudsonrodriguez","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65cbf5256938a6a81f488ba0","name":"csc-unipd","fullname":"Centro di Sonologia Computazionale","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b00f6bc9a5a7680f531c1d/I25ZGtU4dOq44Q2sERhPB.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.03296.md","query":{}}">
Taste-aware music retrieval from audio embeddings
Abstract
Audio encoders from HEAR families are evaluated for taste prediction, with gated late-fusion showing superior rank correlation and the best models achieving human-level accuracy on held-out music.
Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0.134; on held-out real music their error is less than half a single rater's deviation from the consensus (RMSE 0.13 vs. 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0.219). On absolute error the encoders are statistically flat, with a single VGGish matching the best fusion, but gated late-fusion's advantage is confined to rank correlation (macro Pearson r 0.724 vs. 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound-taste correspondences.
Community
Taste-aware music retrieval from audio embeddings 🎵🍫
Can we predict how sweet, bitter, salty, sour, or spicy a piece of music sounds—using only its audio?
In this work, we introduce a content-based music retrieval benchmark for taste prediction from audio. We compare 10 pretrained audio encoders (HEAR families) under a unified frozen-encoder protocol and show that:
• 🎯 Best models achieve 0.134 RMSE, outperforming the previous state of the art (0.219 RMSE).
• 👥 On real music, predictions are closer to the consensus than an average human rater (0.13 vs. 0.28 RMSE).
• 🔎 The learned 5D taste space enables taste-based music retrieval, substantially outperforming a CLAP-text baseline.
• 🧠 We also provide psychophysics-grounded interpretability, linking learned representations to known sound–taste correspondences.
This is, to our knowledge, the first benchmark framing taste-from-audio prediction as a music information retrieval task.
Paper: Taste-aware Music Retrieval from Audio Embeddings
#MusicInformationRetrieval #AudioML #MultimodalAI #MachineLearning #HEAR #RepresentationLearning #Crossmodal #MusicAI
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.03296 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.