Hugging Face Daily Papers · · 5 min read

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

If an agent truly understands a video, can it reconstruct it?</p>\n<p>Introducing BVB: benchmarking agentic video understanding via programmatic reconstruction in Blender.</p>\n<p>288 real videos. 51 agent configurations.</p>\n<p>Paper, demos &amp; leaderboard: <a href=\"https://yoloytang.me/BVB/\" rel=\"nofollow\">https://yoloytang.me/BVB/</a><br>Code: <a href=\"https://github.com/yunlong10/BVB\" rel=\"nofollow\">https://github.com/yunlong10/BVB</a><br>Paper: <a href=\"https://arxiv.org/abs/2609.15478\" rel=\"nofollow\">https://arxiv.org/abs/2609.15478</a></p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/6344c87f0f69ad8aa61dfcf6/bEGK45Jfcai7blzZrJ3Tj.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-09-15T03:47:44.127Z","author":{"_id":"6344c87f0f69ad8aa61dfcf6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6344c87f0f69ad8aa61dfcf6/24TV9sc3rSQzTssTOtW8f.jpeg","fullname":"Yolo Y. Tang","name":"yunlong10","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.5316031575202942},"editors":["yunlong10"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6344c87f0f69ad8aa61dfcf6/24TV9sc3rSQzTssTOtW8f.jpeg"],"reactions":[],"isReport":false}},{"id":"6aa8c080645656b2d4481d8f","author":{"_id":"6344c87f0f69ad8aa61dfcf6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6344c87f0f69ad8aa61dfcf6/24TV9sc3rSQzTssTOtW8f.jpeg","fullname":"Yolo Y. Tang","name":"yunlong10","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-09-15T03:50:24.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"Hi @AdinaY @akhaliq, I’m the author and submitter of BVB:\nhttps://huggingface.co/papers/2609.15478\n\nI uploaded a 30-second comparison video in my comment on the paper page. Could you please use it as the Daily Papers video cover?\n\nPlease preserve the existing upvotes and comments. If resubmission is required, please confirm what will be preserved before removing the entry. Thank you!","html":"<p>Hi <span class=\"SVELTE_PARTIAL_HYDRATER contents\" data-target=\"UserMention\" data-props=\"{&quot;user&quot;:&quot;AdinaY&quot;}\"><span class=\"inline-block\"><span class=\"contents\"><a href=\"/AdinaY\">@<span class=\"underline\">AdinaY</span></a></span> </span></span> <span class=\"SVELTE_PARTIAL_HYDRATER contents\" data-target=\"UserMention\" data-props=\"{&quot;user&quot;:&quot;akhaliq&quot;}\"><span class=\"inline-block\"><span class=\"contents\"><a href=\"/akhaliq\">@<span class=\"underline\">akhaliq</span></a></span> </span></span>, I’m the author and submitter of BVB:<br><a href=\"https://huggingface.co/papers/2609.15478\">https://huggingface.co/papers/2609.15478</a></p>\n<p>I uploaded a 30-second comparison video in my comment on the paper page. Could you please use it as the Daily Papers video cover?</p>\n<p>Please preserve the existing upvotes and comments. If resubmission is required, please confirm what will be preserved before removing the entry. Thank you!</p>\n","updatedAt":"2026-09-15T03:50:24.622Z","author":{"_id":"6344c87f0f69ad8aa61dfcf6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6344c87f0f69ad8aa61dfcf6/24TV9sc3rSQzTssTOtW8f.jpeg","fullname":"Yolo Y. Tang","name":"yunlong10","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8896486759185791},"editors":["yunlong10"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/6344c87f0f69ad8aa61dfcf6/24TV9sc3rSQzTssTOtW8f.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.15478","authors":[{"_id":"6aa8a8905dd4cb9b4cc024be","name":"Yolo Y. Tang","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024bf","name":"Daiki Shimada","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c0","name":"Jiayue Meng","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c1","name":"Jing Bi","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c2","name":"Pinxin Liu","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c3","name":"Yicheng Wang","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c4","name":"Yunzhong Xiao","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c5","name":"Zhangyun Tan","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c6","name":"Zeliang Zhang","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c7","name":"Chao Huang","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c8","name":"Susan Liang","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024c9","name":"Qianxiang Shen","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024ca","name":"Luchuan Song","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024cb","name":"Ali Vosoughi","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024cc","name":"Mingqian Feng","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024cd","name":"Melika Filvantorkaman","hidden":false},{"_id":"6aa8a8905dd4cb9b4cc024ce","name":"Chenliang Xu","hidden":false}],"publishedAt":"2026-09-14T00:00:00.000Z","submittedOnDailyAt":"2026-09-15T00:00:00.000Z","title":"BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender","submittedOnDailyBy":{"_id":"6344c87f0f69ad8aa61dfcf6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6344c87f0f69ad8aa61dfcf6/24TV9sc3rSQzTssTOtW8f.jpeg","isPro":false,"fullname":"Yolo Y. Tang","user":"yunlong10","type":"user","name":"yunlong10"},"summary":"Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introduce BVB, Blender-VideoBench, a benchmark that tests this ability by asking agents to reconstruct real-world videos as animated Blender scenes. To ensure fair comparison, each agent programs the reconstruction through a lightweight harness, Mini-BVB, in an identical sandbox under a shared cost limit. The benchmark renders each reconstruction from its animated camera and evaluates it on two axes: (1) Dual VQA measures how many spatiotemporal facts the reconstruction preserves. (2) Latent Similarity measures how closely the reconstruction matches the source video perceptually. Our overall score, a square-root mean, favors balanced performance. We evaluate 51 configurations from 10 model families and analyze semantic retention, perceptual similarity, reasoning effort, and cost. The best model reaches 88.6 Latent Similarity but retains only 53.7% of the source-correct spatiotemporal answers. Additional reasoning improves visual similarity but does not close this gap in factual accuracy. In a blind study with 15 raters and five configurations, Latent Similarity correlates strongly with human preference. These results show that programmatic reconstruction is a viable test of agentic video understanding, and that semantic retention remains the main challenge.","upvotes":17,"discussionId":"6aa8a8915dd4cb9b4cc024cf","projectPage":"https://yoloytang.me/BVB","githubRepo":"https://github.com/yunlong10/BVB","githubRepoAddedBy":"user","ai_summary":"A benchmark requiring agents to programmatically reconstruct real-world videos in Blender reveals that current models achieve high perceptual similarity but struggle to retain spatiotemporal facts.","ai_keywords":["multimodal agents","diffusion models","Dual VQA","Latent Similarity"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":7},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62eb469dade76f18dd4f0dea","avatarUrl":"/avatars/c558254a0352d115a73febb90bb9370f.svg","isPro":false,"fullname":"Pinxin Liu","user":"pliu23","type":"user"},{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"6344c87f0f69ad8aa61dfcf6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6344c87f0f69ad8aa61dfcf6/24TV9sc3rSQzTssTOtW8f.jpeg","isPro":false,"fullname":"Yolo Y. Tang","user":"yunlong10","type":"user"},{"_id":"66332a98e39731d65a4b7e45","avatarUrl":"/avatars/193e9ee3ac7488960229d6edddb9d1e9.svg","isPro":true,"fullname":"Mingqian Feng","user":"fmmarkmq","type":"user"},{"_id":"675cae2c127c72c5682a4df6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/675cae2c127c72c5682a4df6/-F7uVyRXiWlRHpJ8QxrSV.png","isPro":false,"fullname":"jing bi","user":"jing-bi","type":"user"},{"_id":"67156067e7f22a6aeec76215","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Tq_IpsZJO4LNoR3Plxjc2.png","isPro":false,"fullname":"Yicheng Wang","user":"Nicho123","type":"user"},{"_id":"66e34d5177cdfe475b5fdb07","avatarUrl":"/avatars/92597ede7cb8185e497591c419a067b5.svg","isPro":false,"fullname":"Meng","user":"JoyceMeng","type":"user"},{"_id":"68b84994a820d89d98818f77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/wsW88R_uSsfRGRGWpcoFZ.png","isPro":false,"fullname":"Songlin Yang","user":"songlin1997","type":"user"},{"_id":"65fb02d5bfe25edace8fd22b","avatarUrl":"/avatars/3ab7e364c787b8ba126795ba46799def.svg","isPro":false,"fullname":"Daiki Shimada","user":"dshimada","type":"user"},{"_id":"65763434a4ee9a4fe7cfb156","avatarUrl":"/avatars/2c4a23ff309f750dd9c0d67bd9fd7abc.svg","isPro":true,"fullname":"Susan Liang","user":"susanliang","type":"user"},{"_id":"63f384d40be81bdc5d937852","avatarUrl":"/avatars/8196458794e315291f4e7c986141c311.svg","isPro":false,"fullname":"Shawn Xiao","user":"ShawnXiao","type":"user"},{"_id":"6712c99adda81cf20b59038d","avatarUrl":"/avatars/dffebc5a609fbf846fc282131c34ab95.svg","isPro":false,"fullname":"Leo Huang","user":"masuGC","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.15478.md","query":{}}">
Papers
arxiv:2609.15478

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

Published on Sep 14
· Submitted by
Yolo Y. Tang
on Sep 15
Authors:
,

Abstract

A benchmark requiring agents to programmatically reconstruct real-world videos in Blender reveals that current models achieve high perceptual similarity but struggle to retain spatiotemporal facts.

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introduce BVB, Blender-VideoBench, a benchmark that tests this ability by asking agents to reconstruct real-world videos as animated Blender scenes. To ensure fair comparison, each agent programs the reconstruction through a lightweight harness, Mini-BVB, in an identical sandbox under a shared cost limit. The benchmark renders each reconstruction from its animated camera and evaluates it on two axes: (1) Dual VQA measures how many spatiotemporal facts the reconstruction preserves. (2) Latent Similarity measures how closely the reconstruction matches the source video perceptually. Our overall score, a square-root mean, favors balanced performance. We evaluate 51 configurations from 10 model families and analyze semantic retention, perceptual similarity, reasoning effort, and cost. The best model reaches 88.6 Latent Similarity but retains only 53.7% of the source-correct spatiotemporal answers. Additional reasoning improves visual similarity but does not close this gap in factual accuracy. In a blind study with 15 raters and five configurations, Latent Similarity correlates strongly with human preference. These results show that programmatic reconstruction is a viable test of agentic video understanding, and that semantic retention remains the main challenge.

Community

If an agent truly understands a video, can it reconstruct it?

Introducing BVB: benchmarking agentic video understanding via programmatic reconstruction in Blender.

288 real videos. 51 agent configurations.

Paper, demos & leaderboard: https://yoloytang.me/BVB/
Code: https://github.com/yunlong10/BVB
Paper: https://arxiv.org/abs/2609.15478

Paper submitter about 4 hours ago

Hi @AdinaY @akhaliq , I’m the author and submitter of BVB:
https://huggingface.co/papers/2609.15478

I uploaded a 30-second comparison video in my comment on the paper page. Could you please use it as the Daily Papers video cover?

Please preserve the existing upvotes and comments. If resubmission is required, please confirm what will be preserved before removing the entry. Thank you!

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.15478
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.15478 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.15478 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.15478 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers