Hugging Face Daily Papers · · 3 min read

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

AVE-Compass provides a comprehensive and realistic benchmark for evaluating whether audio-visual editing systems can follow instructions while preserving unedited content, cross-modal consistency, and perceptual quality.</p>\n","updatedAt":"2026-08-06T05:34:09.742Z","author":{"_id":"660165de9e1cf5eb41fe4b0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660165de9e1cf5eb41fe4b0a/rpNxle6Px04AFTAomec0k.jpeg","fullname":"Qianqian Xie","name":"mistletoe111","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.829277515411377},"editors":["mistletoe111"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/660165de9e1cf5eb41fe4b0a/rpNxle6Px04AFTAomec0k.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.24821","authors":[{"_id":"6a69b2069d3a1231d492b98d","user":{"_id":"68dbf78b9cb88433d58aa386","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9AsQXqWAIwQE9h0VIE-9D.png","isPro":false,"fullname":"Yuqing Wen","user":"yuriceee","type":"user","name":"yuriceee"},"name":"Yuqing Wen","status":"claimed_verified","statusLastChangedAt":"2026-08-06T08:45:05.338Z","hidden":false},{"_id":"6a69b2069d3a1231d492b98e","user":{"_id":"69ba919c9bc6c280cf974324","avatarUrl":"/avatars/7d0945af4dd0a1a19c6ea042901141ec.svg","isPro":false,"fullname":"Yukai Huang","user":"huayuankou","type":"user","name":"huayuankou"},"name":"Yukai Huang","status":"claimed_verified","statusLastChangedAt":"2026-08-06T08:45:05.344Z","hidden":false},{"_id":"6a69b2069d3a1231d492b98f","user":{"_id":"660165de9e1cf5eb41fe4b0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660165de9e1cf5eb41fe4b0a/rpNxle6Px04AFTAomec0k.jpeg","isPro":false,"fullname":"Qianqian Xie","user":"mistletoe111","type":"user","name":"mistletoe111"},"name":"Qianqian Xie","status":"claimed_verified","statusLastChangedAt":"2026-08-06T08:45:05.333Z","hidden":false},{"_id":"6a69b2069d3a1231d492b990","name":"Jiangtao Wu","hidden":false},{"_id":"6a69b2069d3a1231d492b991","name":"Yibin Lin","hidden":false},{"_id":"6a69b2069d3a1231d492b992","name":"Yikai Gu","hidden":false},{"_id":"6a69b2069d3a1231d492b993","name":"Jialu Chen","hidden":false},{"_id":"6a69b2069d3a1231d492b994","name":"Yuanxing Zhang","hidden":false},{"_id":"6a69b2069d3a1231d492b995","name":"Jiaheng Liu","hidden":false}],"publishedAt":"2026-07-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-06T00:00:00.000Z","title":"AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities","submittedOnDailyBy":{"_id":"660165de9e1cf5eb41fe4b0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660165de9e1cf5eb41fe4b0a/rpNxle6Px04AFTAomec0k.jpeg","isPro":false,"fullname":"Qianqian Xie","user":"mistletoe111","type":"user","name":"mistletoe111"},"summary":"While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.","upvotes":14,"discussionId":"6a69b2069d3a1231d492b996","projectPage":"https://ave-compass.github.io/","githubRepo":"https://github.com/NJU-LINK/AVE-Compass","githubRepoAddedBy":"user","githubStars":3,"organization":{"_id":"68edc767abe005ac1b354573","name":"NJU-LINK","fullname":"NJU-LINK Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f9d060395fb1a0d7e4ae21/O3V4UZjcSGnOivcQqTcXW.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"660165de9e1cf5eb41fe4b0a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/660165de9e1cf5eb41fe4b0a/rpNxle6Px04AFTAomec0k.jpeg","isPro":false,"fullname":"Qianqian Xie","user":"mistletoe111","type":"user"},{"_id":"68dbf78b9cb88433d58aa386","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/9AsQXqWAIwQE9h0VIE-9D.png","isPro":false,"fullname":"Yuqing Wen","user":"yuriceee","type":"user"},{"_id":"65377c30e48353201e6fdda0","avatarUrl":"/avatars/a8f803b6f2e598eaee9c52c0d2ddfc16.svg","isPro":false,"fullname":"Jiaheng Liu","user":"CheeryLJH","type":"user"},{"_id":"66a9a55d7cda19fabeedbb89","avatarUrl":"/avatars/8e7acdd3a9c3552fbeff882bf32f245e.svg","isPro":false,"fullname":"lxp","user":"lxpp","type":"user"},{"_id":"6940f7acbc8acecb47fa91cb","avatarUrl":"/avatars/7ca424dfb39296089d15d25ba0e7f06e.svg","isPro":false,"fullname":"Lando","user":"Lando0","type":"user"},{"_id":"68d537ea1d2ee6800f0b57e6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68d537ea1d2ee6800f0b57e6/9D2Dnwz_NIyHhYyRQk4nC.jpeg","isPro":false,"fullname":"vicky","user":"Vickyinmyheart824","type":"user"},{"_id":"6716650a4061ad3cc6c6f56a","avatarUrl":"/avatars/8a3a62c041959a8d1d6e6eacdf4fe0c9.svg","isPro":false,"fullname":"zhangxiaohan","user":"zxhhhhhh","type":"user"},{"_id":"6920035cd48b817fb1297da3","avatarUrl":"/avatars/109a16cb17427cec1248e70ff09ac239.svg","isPro":false,"fullname":"Mingrong Gong","user":"Gmr5233","type":"user"},{"_id":"66efe938ce6e5db9b365b39e","avatarUrl":"/avatars/42b96198fd19414e14add302c48cee8a.svg","isPro":false,"fullname":"Junhao Dong","user":"tomvii","type":"user"},{"_id":"6a6a82f2a698558ca7615c4d","avatarUrl":"/avatars/2840b88a59f45386596ba96cd51fa32a.svg","isPro":false,"fullname":"Mark Rodriguez","user":"mark-r","type":"user"},{"_id":"69ba919c9bc6c280cf974324","avatarUrl":"/avatars/7d0945af4dd0a1a19c6ea042901141ec.svg","isPro":false,"fullname":"Yukai Huang","user":"huayuankou","type":"user"},{"_id":"6a6d3d66ef16fe7cceaf0e2a","avatarUrl":"/avatars/9a0c85f1ad367ca2b3909cd536abf6f9.svg","isPro":false,"fullname":"Michael Brown","user":"cedarglade","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68edc767abe005ac1b354573","name":"NJU-LINK","fullname":"NJU-LINK Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67f9d060395fb1a0d7e4ae21/O3V4UZjcSGnOivcQqTcXW.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.24821.md","query":{}}">
Papers
arxiv:2607.24821

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Published on Jul 17
· Submitted by
Qianqian Xie
on Aug 6

Abstract

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.

Community

Paper author Paper submitter about 4 hours ago

AVE-Compass provides a comprehensive and realistic benchmark for evaluating whether audio-visual editing systems can follow instructions while preserving unedited content, cross-modal consistency, and perceptual quality.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.24821
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.24821 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.24821 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers