Hugging Face Daily Papers · · 4 min read

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Most automatic speaker verification (ASV) systems<br>operate on individual utterances, despite real-world interactions<br>typically consisting of multiple utterances. As speech accumu-<br>lates, increasingly rich speaker information becomes available<br>through acoustic, prosodic, and linguistic cues, potentially chal-<br>lenging speaker anonymization methods that primarily target<br>vocal characteristics. We investigate ASV in a multi-utterance,<br>multimodal setting and examine whether aggregating information<br>across anonymized speech impacts privacy. We first study audio-<br>only aggregation across multiple anonymized utterances and<br>observe consistent performance improvements as more speech<br>becomes available. We then incorporate prosodic and linguistic<br>information, showing that multimodal systems outperform uni-<br>modal approaches. Finally, we compare aggregation strategies<br>and find that frame-level aggregation yields the lowest EERs.<br>Even with only five anonymized utterances, combining audio<br>and text reduces EER by over 15% relative to audio-only ag-<br>gregation, demonstrating that substantial speaker-discriminative<br>information remains accessible despite anonymization.</p>\n","updatedAt":"2026-07-27T09:28:53.577Z","author":{"_id":"6446ab9815a27291ef8b7313","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6446ab9815a27291ef8b7313/rw5xK-BV2jyct0G3ft13a.png","fullname":"Ashi Garg","name":"ash56","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8597519993782043},"editors":["ash56"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6446ab9815a27291ef8b7313/rw5xK-BV2jyct0G3ft13a.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.19636","authors":[{"_id":"6a67246dab9cdf9be5794bb1","name":"Ashi Garg","hidden":false},{"_id":"6a67246dab9cdf9be5794bb2","name":"Cristina Aggazzotti","hidden":false},{"_id":"6a67246dab9cdf9be5794bb3","name":"Leibny Paola García-Perera","hidden":false},{"_id":"6a67246dab9cdf9be5794bb4","name":"Nicholas Andrews","hidden":false}],"publishedAt":"2026-07-22T00:00:00.000Z","submittedOnDailyAt":"2026-07-27T00:00:00.000Z","title":"Multimodal Speaker Verification as a Threat to Speaker Anonymization","submittedOnDailyBy":{"_id":"6446ab9815a27291ef8b7313","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6446ab9815a27291ef8b7313/rw5xK-BV2jyct0G3ft13a.png","isPro":false,"fullname":"Ashi Garg","user":"ash56","type":"user","name":"ash56"},"summary":"Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.","upvotes":1,"discussionId":"6a67246dab9cdf9be5794bb5","githubRepo":"https://github.com/Ashigarg123/multimodal-speaker-verification","githubRepoAddedBy":"user","githubStars":1,"organization":{"_id":"653945b47ba797097a7b4eab","name":"JohnsHopkins","fullname":"Johns Hopkins University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/653944e58e687a41625a4694/qqHzBOarppVrUuZbbjqwh.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6446ab9815a27291ef8b7313","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6446ab9815a27291ef8b7313/rw5xK-BV2jyct0G3ft13a.png","isPro":false,"fullname":"Ashi Garg","user":"ash56","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"653945b47ba797097a7b4eab","name":"JohnsHopkins","fullname":"Johns Hopkins University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/653944e58e687a41625a4694/qqHzBOarppVrUuZbbjqwh.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.19636.md","query":{}}">
Papers
arxiv:2607.19636

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Published on Jul 22
· Submitted by
Ashi Garg
on Jul 27
Authors:
,

Abstract

Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.

Community

Paper submitter about 4 hours ago

Most automatic speaker verification (ASV) systems
operate on individual utterances, despite real-world interactions
typically consisting of multiple utterances. As speech accumu-
lates, increasingly rich speaker information becomes available
through acoustic, prosodic, and linguistic cues, potentially chal-
lenging speaker anonymization methods that primarily target
vocal characteristics. We investigate ASV in a multi-utterance,
multimodal setting and examine whether aggregating information
across anonymized speech impacts privacy. We first study audio-
only aggregation across multiple anonymized utterances and
observe consistent performance improvements as more speech
becomes available. We then incorporate prosodic and linguistic
information, showing that multimodal systems outperform uni-
modal approaches. Finally, we compare aggregation strategies
and find that frame-level aggregation yields the lowest EERs.
Even with only five anonymized utterances, combining audio
and text reduces EER by over 15% relative to audio-only ag-
gregation, demonstrating that substantial speaker-discriminative
information remains accessible despite anonymization.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.19636
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.19636 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.19636 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.19636 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers