We also evaluate the latest models Gemini-Omni-Flash, HappyHorse-1.1, and MiniMax-H3. Welcome any feedback!</p>\n","updatedAt":"2026-08-11T02:39:35.165Z","author":{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","fullname":"Tingyu Song","name":"songtingyu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6781448721885681},"editors":["songtingyu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.09873","authors":[{"_id":"6a7a8a52019ce76dc7b3a9d5","name":"Diandian Zhang","hidden":false},{"_id":"6a7a8a52019ce76dc7b3a9d6","name":"Tingyu Song","hidden":false},{"_id":"6a7a8a52019ce76dc7b3a9d7","name":"Lin Fu","hidden":false},{"_id":"6a7a8a52019ce76dc7b3a9d8","name":"Zheyuan Yang","hidden":false},{"_id":"6a7a8a52019ce76dc7b3a9d9","name":"Yilun Zhao","hidden":false}],"publishedAt":"2026-08-10T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains","submittedOnDailyBy":{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","isPro":false,"fullname":"Tingyu Song","user":"songtingyu","type":"user","name":"songtingyu"},"summary":"We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.","upvotes":22,"discussionId":"6a7a8a53019ce76dc7b3a9da","githubRepo":"https://github.com/sci-vbench/sci-vbench","githubRepoAddedBy":"user","ai_summary":"Sci-VBench evaluates video generation requiring scientific reasoning across disciplines, revealing that visual realism advances have not ensured accurate scientific and causal dynamics.","ai_keywords":["video generation","scientific reasoning","MLLM-as-Judge","prompt grounding","causal correctness","perceptual-quality"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":3},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64dc29d9b5d625e0e9a6ecb9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/QxGBsnk1cNsBEPqSx4ae-.jpeg","isPro":false,"fullname":"Tingyu Song","user":"songtingyu","type":"user"},{"_id":"62f662bcc58915315c4eccea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/62f662bcc58915315c4eccea/zOAQLONfMP88zr70sxHK-.jpeg","isPro":true,"fullname":"Yilun Zhao","user":"yilunzhao","type":"user"},{"_id":"68084d54aca60e6178b3afb5","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/68084d54aca60e6178b3afb5/TshN3Ka3VRFD_I3WJ6Vys.jpeg","isPro":false,"fullname":"Lin Fu","user":"minuzero","type":"user"},{"_id":"65dfeee3d16fb170031df293","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65dfeee3d16fb170031df293/2VbNuqcpN3XrWB18NfzRQ.jpeg","isPro":false,"fullname":"gan","user":"guo9","type":"user"},{"_id":"66af69222f4c59963afc874f","avatarUrl":"/avatars/034ca7688282bdbeddbd4f03e54dead7.svg","isPro":false,"fullname":"Zheyuan Yang","user":"Raywithyou","type":"user"},{"_id":"638f1803c67af472d317a922","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/638f1803c67af472d317a922/9BMVXqHa-AsdZPmBprcbd.jpeg","isPro":false,"fullname":"siyue zhang","user":"siyue","type":"user"},{"_id":"6a2f2e96d5560ff540676f39","avatarUrl":"/avatars/ab8a604f2a74c0917ae50676bbfc9f2d.svg","isPro":false,"fullname":"Zhang Wei","user":"zhangwei-hf","type":"user"},{"_id":"6a2f0c5f63c271161df37e79","avatarUrl":"/avatars/d6598ca54b33fd2895fbfdefdfa0574e.svg","isPro":false,"fullname":"Lijun Tan","user":"lijuntan","type":"user"},{"_id":"6a2f0ca98971f84f78246fb9","avatarUrl":"/avatars/79a6bab6d4f352c23f92ec314ef05e8d.svg","isPro":false,"fullname":"Kevin Li","user":"kaiweili","type":"user"},{"_id":"6a2f32c2a14a4189799755e0","avatarUrl":"/avatars/bf3e4ecdfc251c9280cbeddf9cd65f87.svg","isPro":false,"fullname":"Yifan Gao","user":"gao-yifan","type":"user"},{"_id":"6a2f322b000819df3135c0f2","avatarUrl":"/avatars/4206347a362c47e0ff7a22a1ac252c44.svg","isPro":false,"fullname":"Jiaqi Gao","user":"jgao20","type":"user"},{"_id":"6a2f32007e3480ed6543904b","avatarUrl":"/avatars/01760cf6f41ba95ace395d09c7a7b703.svg","isPro":false,"fullname":"Zihan Liang","user":"liang9553","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.09873.md","query":{}}">
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Abstract
Sci-VBench evaluates video generation requiring scientific reasoning across disciplines, revealing that visual realism advances have not ensured accurate scientific and causal dynamics.
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
Community
We also evaluate the latest models Gemini-Omni-Flash, HappyHorse-1.1, and MiniMax-H3. Welcome any feedback!
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.09873 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.09873 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.09873 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.