Most existing evaluations of short‑drama generation only assess the video‑generation stage in isolation. Their test samples are offline‑produced, rather than outputs from upstream modules of real‑world production pipelines. Consequently, it is difficult to observe the cross‑stage propagation effect of defects throughout the full production chain. To address this issue, we present DramaChain‑Bench. We build an end‑to‑end generation pipeline that emulates the industrial workflow of commercial short dramas, together with a professional human‑annotation system and an automated evaluation framework. This creates the industry’s first full‑chain benchmark for short‑drama generation covering storyboard design, key‑frame images, storyboard‑level videos, and final finished episodes.</p>\n<p>Project page: <a href=\"https://dramachain-bench.github.io/\" rel=\"nofollow\">https://dramachain-bench.github.io/</a></p>\n","updatedAt":"2026-09-02T16:03:08.726Z","author":{"_id":"652fb8bcc9dd2692a25ef2e3","avatarUrl":"/avatars/461e6cc1c3441cde18192b080b0b8576.svg","fullname":"Haoyuan Shi","name":"MrSunshy","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.834307074546814},"editors":["MrSunshy"],"editorAvatarUrls":["/avatars/461e6cc1c3441cde18192b080b0b8576.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00646","authors":[{"_id":"6a9846f8fea818274321fcea","user":{"_id":"652fb8bcc9dd2692a25ef2e3","avatarUrl":"/avatars/461e6cc1c3441cde18192b080b0b8576.svg","isPro":false,"fullname":"Haoyuan Shi","user":"MrSunshy","type":"user","name":"MrSunshy"},"name":"Haoyuan Shi","status":"claimed_verified","statusLastChangedAt":"2026-09-02T16:45:04.895Z","hidden":false},{"_id":"6a9846f8fea818274321fceb","name":"Mingtao Chen","hidden":false},{"_id":"6a9846f8fea818274321fcec","name":"Shuo Jiang","hidden":false},{"_id":"6a9846f8fea818274321fced","name":"Ziyan Chen","hidden":false},{"_id":"6a9846f8fea818274321fcee","name":"Xuyi Sheng","hidden":false},{"_id":"6a9846f8fea818274321fcef","name":"Yiming Liu","hidden":false},{"_id":"6a9846f8fea818274321fcf0","name":"Ying Zhang","hidden":false},{"_id":"6a9846f8fea818274321fcf1","name":"Miao Wang","hidden":false},{"_id":"6a9846f8fea818274321fcf2","name":"Jianxiang Lu","hidden":false},{"_id":"6a9846f8fea818274321fcf3","name":"Fanyang Lu","hidden":false},{"_id":"6a9846f8fea818274321fcf4","name":"Songyuanyi Lu","hidden":false},{"_id":"6a9846f8fea818274321fcf5","name":"Xiele Wu","hidden":false},{"_id":"6a9846f8fea818274321fcf6","name":"Zhichao Hu","hidden":false},{"_id":"6a9846f8fea818274321fcf7","name":"Yuhong Liu","hidden":false},{"_id":"6a9846f8fea818274321fcf8","name":"Richeng Xuan","hidden":false}],"publishedAt":"2026-09-01T00:00:00.000Z","submittedOnDailyAt":"2026-09-02T00:00:00.000Z","title":"DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation","submittedOnDailyBy":{"_id":"652fb8bcc9dd2692a25ef2e3","avatarUrl":"/avatars/461e6cc1c3441cde18192b080b0b8576.svg","isPro":false,"fullname":"Haoyuan Shi","user":"MrSunshy","type":"user","name":"MrSunshy"},"summary":"Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.","upvotes":1,"discussionId":"6a9846f9fea818274321fcf9","projectPage":"https://dramachain-bench.github.io/","ai_summary":"DramaChain Bench evaluates the full short-drama production pipeline across stages, dimensions, and cascading defects using calibrated agents and expert annotations.","ai_keywords":["DramaChain Bench","multi-stage production chain","script adherence","cross-shot coherence","DramaChain Dimensions","DramaChain Agent","agentic judge","spatio-temporal defect localization","upstream defect cascade"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"652fb8bcc9dd2692a25ef2e3","avatarUrl":"/avatars/461e6cc1c3441cde18192b080b0b8576.svg","isPro":false,"fullname":"Haoyuan Shi","user":"MrSunshy","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6645f953c39288df638dbdd5","name":"Tencent-Hunyuan","fullname":"Tencent Hunyuan","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/62d22496c58f969c152bcefd/woKSjt2wXvBNKussyYPsa.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00646.md","query":{}}">
DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
Abstract
DramaChain Bench evaluates the full short-drama production pipeline across stages, dimensions, and cascading defects using calibrated agents and expert annotations.
Commercial short-drama production follows a multi-stage chain: script, storyboard, keyframe imagery, shot-level video, and the finished short drama. Most existing benchmarks evaluate solely the video-generation stage using pre-authored inputs instead of real upstream pipeline outputs. This leaves two critical questions unanswerable: whether each stage adheres to the original script intent (rather than only its immediate input prompt), and whether disparate shots remain coherent after assembly into multi-episode releases. We present DramaChain Bench, the first short-drama benchmark that evaluates every stage of the complete production chain. It is built upon three in-house systems sharing one dimension system, DramaChain Dimensions: five evaluation axes instantiated at every stage, resolving into 63 leaf dimensions. DramaChain Agent is calibrated against commercial short-drama platforms in both workflow and finished short-drama quality, enabling stage-wise fair comparison across models. DramaChain Labeling System has each of the 5,785 items scored independently by three professional annotators, with all defects spatio-temporally localised and selected from a predefined defect list. This process produces 17,488 valid scores and 255,925 traceable attribution records. The human annotations confirm that upstream defects cascade across the pipeline, demonstrating that final episode quality is not governed by video generation alone. DramaChain Agentic Judge then scores every leaf dimension automatically, gathering evidence over multiple agentic rounds before judging against a per-item checklist; it reproduces the model ranking at a mean PLCC of 0.918, enough to admit new models at no annotation cost.
Community
Most existing evaluations of short‑drama generation only assess the video‑generation stage in isolation. Their test samples are offline‑produced, rather than outputs from upstream modules of real‑world production pipelines. Consequently, it is difficult to observe the cross‑stage propagation effect of defects throughout the full production chain. To address this issue, we present DramaChain‑Bench. We build an end‑to‑end generation pipeline that emulates the industrial workflow of commercial short dramas, together with a professional human‑annotation system and an automated evaluation framework. This creates the industry’s first full‑chain benchmark for short‑drama generation covering storyboard design, key‑frame images, storyboard‑level videos, and final finished episodes.
Project page: https://dramachain-bench.github.io/
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.00646 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.00646 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.00646 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.