Github: <a href=\"https://github.com/NVIDIA/audio-flamingo\" rel=\"nofollow\">https://github.com/NVIDIA/audio-flamingo</a></p>\n","updatedAt":"2026-07-20T02:08:29.223Z","author":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","fullname":"taesiri","name":"taesiri","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":337,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.31712591648101807},"editors":["taesiri"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.16107","authors":[{"_id":"6a5d7fb26a69ce099f4d6d21","name":"Sreyan Ghosh","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d22","name":"Arushi Goel","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d23","name":"Kaousheik Jayakumar","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d24","name":"Lasha Koroshinadze","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d25","name":"Nishit Anand","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d26","name":"Siddharth Gururani","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d27","name":"Hanrong Ye","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d28","name":"Pritam Biswas","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d29","name":"Yuanhang Su","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d2a","name":"Ehsan Hosseini-Asl","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d2b","name":"Sang-gil Lee","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d2c","name":"Zhifeng Kong","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d2d","name":"Jaehyeon Kim","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d2e","name":"Sungwon Kim","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d2f","name":"S Sakshi","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d30","name":"Ramani Duraiswami","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d31","name":"Dinesh Manocha","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d32","name":"Andrew Tao","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d33","name":"Mohammad Shoeybi","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d34","name":"Bryan Catanzaro","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d35","name":"Ming-Yu Liu","hidden":false},{"_id":"6a5d7fb26a69ce099f4d6d36","name":"Wei Ping","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/maON_7TSbjJGrt899QDvd.webp"],"publishedAt":"2026-07-17T00:00:00.000Z","submittedOnDailyAt":"2026-07-20T00:00:00.000Z","title":"Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos","submittedOnDailyBy":{"_id":"6039478ab3ecf716b1a5fd4d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6039478ab3ecf716b1a5fd4d/_Thy4E7taiSYBLKxEKJbT.jpeg","isPro":true,"fullname":"taesiri","user":"taesiri","type":"user","name":"taesiri"},"summary":"We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.","upvotes":4,"discussionId":"6a5d7fb26a69ce099f4d6d37","projectPage":"https://avflamingo.pages.dev/","organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62c9664eb34e600d7eaa4beb","avatarUrl":"/avatars/ca23ecdec2d31c99ecce97d9b180ae0c.svg","isPro":false,"fullname":"Ghosh","user":"Sreyan88","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"69bd1cdb76fe7a5ea0adbd25","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QCmjCtgaXLKwZO_QPLjpb.png","isPro":false,"fullname":"山口陽翔","user":"ladams88","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"60262b67268c201cdc8b7d43","name":"nvidia","fullname":"NVIDIA","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65df9200dc3292a8983e5017/Vs5FPVCH-VZBipV3qKTuy.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.16107.md","query":{}}">
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
Abstract
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.