VideoLLMs are expensive because cost scales with frames and context length, and the<br>efficiency literature is scattered across frame sampling, encoders, connectors and the<br>LLM itself, with no shared way to compare methods.</p>\n<p>We survey 125 papers on inference efficiency for video and audiovisual LLMs, organized<br>by where in the pipeline the cost is cut: frame sampling, modality encoding,<br>connector-level token reduction, and LLM prefilling/decoding. We only include methods<br>reporting concrete reductions (params, FLOPs, latency, memory, or visual/audio tokens).</p>\n<p>Where papers share a host model and input protocol, we assemble accuracy–cost<br>comparisons and keep them separate from heterogeneous cross-paper numbers. Two gaps<br>stand out: audiovisual efficiency is barely studied, and efficiency evaluation is not<br>standardized.</p>\n<p>Living repo of the papers: <a href=\"https://github.com/momentslab/awesome-efficient-videollm\" rel=\"nofollow\">https://github.com/momentslab/awesome-efficient-videollm</a></p>\n","updatedAt":"2026-09-10T09:54:10.055Z","author":{"_id":"6362afc7d3be91534c2ee7c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png","fullname":"Killian Steunou","name":"nelikCode","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd4fc3299c0c2e0e42f249/OETSSIzgwQTYYDBD5b1H4.png","fullname":"Moments Lab","name":"momentslab","type":"org","isHf":false,"details":"video understanding, vision language models, information retrieval","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8956930637359619},"editors":["nelikCode"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.10355","authors":[{"_id":"6aa27d2ba2aeb74440b1e010","name":"Killian Steunou","hidden":false},{"_id":"6aa27d2ba2aeb74440b1e011","name":"Yannis Tevissen","hidden":false},{"_id":"6aa27d2ba2aeb74440b1e012","name":"Mounîm A. El Yacoubi","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6362afc7d3be91534c2ee7c1/rravgJhzs-dHXlrMb8Ngf.png"],"publishedAt":"2026-09-09T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs","submittedOnDailyBy":{"_id":"6362afc7d3be91534c2ee7c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png","isPro":false,"fullname":"Killian Steunou","user":"nelikCode","type":"user","name":"nelikCode"},"summary":"Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count. We analyze bottlenecks across frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding. We organize methods by the pipeline stage at which they act, covering VideoLLMs developed since late 2022 together with earlier frame-sampling and vision-encoder mechanisms that remain components of current pipelines. We assemble literature-reported accuracy--cost comparisons under shared host models and input protocols wherever available, distinguish them from heterogeneous cross-paper evidence, and identify gaps in audiovisual efficiency and standardized evaluation. We maintain a repository at https://github.com/momentslab/awesome-efficient-videollm.","upvotes":4,"discussionId":"6aa27d2ba2aeb74440b1e013","projectPage":"https://www.killian-steunou.com/videollm-survey/","githubRepo":"https://github.com/momentslab/awesome-efficient-videollm","githubRepoAddedBy":"user","ai_summary":"This survey examines inference-efficiency techniques for video large language models, analyzing cost reductions across frame sampling, encoding, token compression, and language model stages while identifying evaluation gaps.","ai_keywords":["VideoLLMs","inference-efficiency","frame sampling","modality encoding","connector-level token reduction","LLM prefilling","decoding","audiovisual VideoLLMs","token reduction","FLOPs"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"65f98c03281c4728d6963e04","name":"momentslab","fullname":"Moments Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd4fc3299c0c2e0e42f249/OETSSIzgwQTYYDBD5b1H4.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6362afc7d3be91534c2ee7c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png","isPro":false,"fullname":"Killian Steunou","user":"nelikCode","type":"user"},{"_id":"6825ec235aa29e7dec262576","avatarUrl":"/avatars/8e2bd75c920ecc13da5f54243ec9afde.svg","isPro":false,"fullname":"Filali Razzouki","user":"filalianas-dev","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65f98c03281c4728d6963e04","name":"momentslab","fullname":"Moments Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd4fc3299c0c2e0e42f249/OETSSIzgwQTYYDBD5b1H4.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.10355.md","query":{}}">
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
Abstract
This survey examines inference-efficiency techniques for video large language models, analyzing cost reductions across frame sampling, encoding, token compression, and language model stages while identifying evaluation gaps.
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count. We analyze bottlenecks across frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding. We organize methods by the pipeline stage at which they act, covering VideoLLMs developed since late 2022 together with earlier frame-sampling and vision-encoder mechanisms that remain components of current pipelines. We assemble literature-reported accuracy--cost comparisons under shared host models and input protocols wherever available, distinguish them from heterogeneous cross-paper evidence, and identify gaps in audiovisual efficiency and standardized evaluation. We maintain a repository at https://github.com/momentslab/awesome-efficient-videollm.
Community
VideoLLMs are expensive because cost scales with frames and context length, and the
efficiency literature is scattered across frame sampling, encoders, connectors and the
LLM itself, with no shared way to compare methods.
We survey 125 papers on inference efficiency for video and audiovisual LLMs, organized
by where in the pipeline the cost is cut: frame sampling, modality encoding,
connector-level token reduction, and LLM prefilling/decoding. We only include methods
reporting concrete reductions (params, FLOPs, latency, memory, or visual/audio tokens).
Where papers share a host model and input protocol, we assemble accuracy–cost
comparisons and keep them separate from heterogeneous cross-paper numbers. Two gaps
stand out: audiovisual efficiency is barely studied, and efficiency evaluation is not
standardized.
Living repo of the papers: https://github.com/momentslab/awesome-efficient-videollm
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.10355 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.10355 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.10355 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.