Hugging Face Daily Papers · · 4 min read

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

VideoLLMs are expensive because cost scales with frames and context length, and the<br>efficiency literature is scattered across frame sampling, encoders, connectors and the<br>LLM itself, with no shared way to compare methods.</p>\n<p>We survey 125 papers on inference efficiency for video and audiovisual LLMs, organized<br>by where in the pipeline the cost is cut: frame sampling, modality encoding,<br>connector-level token reduction, and LLM prefilling/decoding. We only include methods<br>reporting concrete reductions (params, FLOPs, latency, memory, or visual/audio tokens).</p>\n<p>Where papers share a host model and input protocol, we assemble accuracy–cost<br>comparisons and keep them separate from heterogeneous cross-paper numbers. Two gaps<br>stand out: audiovisual efficiency is barely studied, and efficiency evaluation is not<br>standardized.</p>\n<p>Living repo of the papers: <a href=\"https://github.com/momentslab/awesome-efficient-videollm\" rel=\"nofollow\">https://github.com/momentslab/awesome-efficient-videollm</a></p>\n","updatedAt":"2026-09-10T09:54:10.055Z","author":{"_id":"6362afc7d3be91534c2ee7c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png","fullname":"Killian Steunou","name":"nelikCode","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false,"primaryOrg":{"avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd4fc3299c0c2e0e42f249/OETSSIzgwQTYYDBD5b1H4.png","fullname":"Moments Lab","name":"momentslab","type":"org","isHf":false,"details":"video understanding, vision language models, information retrieval","plan":"team"}}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8956930637359619},"editors":["nelikCode"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.10355","authors":[{"_id":"6aa27d2ba2aeb74440b1e010","name":"Killian Steunou","hidden":false},{"_id":"6aa27d2ba2aeb74440b1e011","name":"Yannis Tevissen","hidden":false},{"_id":"6aa27d2ba2aeb74440b1e012","name":"Mounîm A. El Yacoubi","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6362afc7d3be91534c2ee7c1/rravgJhzs-dHXlrMb8Ngf.png"],"publishedAt":"2026-09-09T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs","submittedOnDailyBy":{"_id":"6362afc7d3be91534c2ee7c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png","isPro":false,"fullname":"Killian Steunou","user":"nelikCode","type":"user","name":"nelikCode"},"summary":"Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count. We analyze bottlenecks across frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding. We organize methods by the pipeline stage at which they act, covering VideoLLMs developed since late 2022 together with earlier frame-sampling and vision-encoder mechanisms that remain components of current pipelines. We assemble literature-reported accuracy--cost comparisons under shared host models and input protocols wherever available, distinguish them from heterogeneous cross-paper evidence, and identify gaps in audiovisual efficiency and standardized evaluation. We maintain a repository at https://github.com/momentslab/awesome-efficient-videollm.","upvotes":4,"discussionId":"6aa27d2ba2aeb74440b1e013","projectPage":"https://www.killian-steunou.com/videollm-survey/","githubRepo":"https://github.com/momentslab/awesome-efficient-videollm","githubRepoAddedBy":"user","ai_summary":"This survey examines inference-efficiency techniques for video large language models, analyzing cost reductions across frame sampling, encoding, token compression, and language model stages while identifying evaluation gaps.","ai_keywords":["VideoLLMs","inference-efficiency","frame sampling","modality encoding","connector-level token reduction","LLM prefilling","decoding","audiovisual VideoLLMs","token reduction","FLOPs"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"65f98c03281c4728d6963e04","name":"momentslab","fullname":"Moments Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd4fc3299c0c2e0e42f249/OETSSIzgwQTYYDBD5b1H4.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6362afc7d3be91534c2ee7c1","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1667411893967-noauth.png","isPro":false,"fullname":"Killian Steunou","user":"nelikCode","type":"user"},{"_id":"6825ec235aa29e7dec262576","avatarUrl":"/avatars/8e2bd75c920ecc13da5f54243ec9afde.svg","isPro":false,"fullname":"Filali Razzouki","user":"filalianas-dev","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65f98c03281c4728d6963e04","name":"momentslab","fullname":"Moments Lab","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62cd4fc3299c0c2e0e42f249/OETSSIzgwQTYYDBD5b1H4.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.10355.md","query":{}}">
Papers
arxiv:2609.10355

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Published on Sep 9
· Submitted by
Killian Steunou
on Sep 10
Authors:
,

Abstract

This survey examines inference-efficiency techniques for video large language models, analyzing cost reductions across frame sampling, encoding, token compression, and language model stages while identifying evaluation gaps.

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count. We analyze bottlenecks across frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding. We organize methods by the pipeline stage at which they act, covering VideoLLMs developed since late 2022 together with earlier frame-sampling and vision-encoder mechanisms that remain components of current pipelines. We assemble literature-reported accuracy--cost comparisons under shared host models and input protocols wherever available, distinguish them from heterogeneous cross-paper evidence, and identify gaps in audiovisual efficiency and standardized evaluation. We maintain a repository at https://github.com/momentslab/awesome-efficient-videollm.

Community

Paper submitter about 7 hours ago

VideoLLMs are expensive because cost scales with frames and context length, and the
efficiency literature is scattered across frame sampling, encoders, connectors and the
LLM itself, with no shared way to compare methods.

We survey 125 papers on inference efficiency for video and audiovisual LLMs, organized
by where in the pipeline the cost is cut: frame sampling, modality encoding,
connector-level token reduction, and LLM prefilling/decoding. We only include methods
reporting concrete reductions (params, FLOPs, latency, memory, or visual/audio tokens).

Where papers share a host model and input protocol, we assemble accuracy–cost
comparisons and keep them separate from heterogeneous cross-paper numbers. Two gaps
stand out: audiovisual efficiency is barely studied, and efficiency evaluation is not
standardized.

Living repo of the papers: https://github.com/momentslab/awesome-efficient-videollm

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.10355
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.10355 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.10355 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.10355 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers