Hugging Face Daily Papers · · 6 min read

The Attention Triangle in Audio-Video Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage.<br>We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation.<br>Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation.<br>This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes.<br>These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways.<br>Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions.<br>This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment.<br>Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.</p>\n<p>🌐 visit our project page at: <a href=\"https://sagipolaczek.github.io/The-Attention-Triangle/\" rel=\"nofollow\">https://sagipolaczek.github.io/The-Attention-Triangle/</a></p>\n","updatedAt":"2026-09-07T08:35:09.752Z","author":{"_id":"630b37dee67c604e9b79e773","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630b37dee67c604e9b79e773/roSaYAsQ4E9rQrahYZitG.png","fullname":"Sagi Polaczek","name":"SagiPolaczek","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":8,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8828911185264587},"editors":["SagiPolaczek"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/630b37dee67c604e9b79e773/roSaYAsQ4E9rQrahYZitG.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03586","authors":[{"_id":"6a9e7640de5ea82090db66c6","name":"Sagi Polaczek","hidden":false},{"_id":"6a9e7640de5ea82090db66c7","name":"Noa Kraicer","hidden":false},{"_id":"6a9e7640de5ea82090db66c8","name":"Gal Metzer","hidden":false},{"_id":"6a9e7640de5ea82090db66c9","name":"Zhuo Ning","hidden":false},{"_id":"6a9e7640de5ea82090db66ca","name":"Ali Mahdavi-Amiri","hidden":false},{"_id":"6a9e7640de5ea82090db66cb","name":"Daniel Cohen-Or","hidden":false},{"_id":"6a9e7640de5ea82090db66cc","name":"Raja Giryes","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/630b37dee67c604e9b79e773/5hBzxy_SFwwAKlHSJ_RON.mp4"],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-07T00:00:00.000Z","title":"The Attention Triangle in Audio-Video Models","submittedOnDailyBy":{"_id":"630b37dee67c604e9b79e773","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630b37dee67c604e9b79e773/roSaYAsQ4E9rQrahYZitG.png","isPro":false,"fullname":"Sagi Polaczek","user":"SagiPolaczek","type":"user","name":"SagiPolaczek"},"summary":"Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.","upvotes":2,"discussionId":"6a9e7640de5ea82090db66cd","projectPage":"https://sagipolaczek.github.io/The-Attention-Triangle","githubRepo":"https://github.com/SagiPolaczek/The-Attention-Triangle","githubRepoAddedBy":"user","ai_summary":"Audio-video diffusion models exhibit bidirectional semantic leakage through cross-modal attention pathways, which can be diagnosed via attention-derived signals and mitigated through inference-time alignment interventions.","ai_keywords":["audio-video diffusion models","cross-modal attention","attention triangle","semantic leakage","cross-attention edges","inference-time interventions","cross-modal alignment"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6107dfc57602f8e9ed8bb5cb","name":"tau","fullname":"Tel Aviv University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1628143727824-610b729f9da682cd54ad9adf.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"630b37dee67c604e9b79e773","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/630b37dee67c604e9b79e773/roSaYAsQ4E9rQrahYZitG.png","isPro":false,"fullname":"Sagi Polaczek","user":"SagiPolaczek","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6107dfc57602f8e9ed8bb5cb","name":"tau","fullname":"Tel Aviv University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1628143727824-610b729f9da682cd54ad9adf.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03586.md","query":{}}">
Papers
arxiv:2609.03586

The Attention Triangle in Audio-Video Models

Published on Sep 3
· Submitted by
Sagi Polaczek
on Sep 7
Authors:
,

Abstract

Audio-video diffusion models exhibit bidirectional semantic leakage through cross-modal attention pathways, which can be diagnosed via attention-derived signals and mitigated through inference-time alignment interventions.

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.

Community

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage.
We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation.
Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation.
This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes.
These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways.
Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions.
This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment.
Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.

🌐 visit our project page at: https://sagipolaczek.github.io/The-Attention-Triangle/

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.03586
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.03586 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.03586 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.03586 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers