FlashRender retakes an input video along a target camera trajectory in seconds, using only 4 NFE.</p>\n","updatedAt":"2026-09-04T03:19:21.624Z","author":{"_id":"653929a66da48e0d21e65e17","avatarUrl":"/avatars/e34ae3411d689b4280ff34c1b680f283.svg","fullname":"Byeongjun Park","name":"byeongjun-park","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6293582320213318},"editors":["byeongjun-park"],"editorAvatarUrls":["/avatars/e34ae3411d689b4280ff34c1b680f283.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03563","authors":[{"_id":"6a9a24f88f7c3b7557239409","user":{"_id":"653929a66da48e0d21e65e17","avatarUrl":"/avatars/e34ae3411d689b4280ff34c1b680f283.svg","isPro":false,"fullname":"Byeongjun Park","user":"byeongjun-park","type":"user","name":"byeongjun-park"},"name":"Byeongjun Park","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.241Z","hidden":false},{"_id":"6a9a24f88f7c3b755723940a","name":"Byung-Hoon Kim","hidden":false},{"_id":"6a9a24f88f7c3b755723940b","name":"Hyungjin Chung","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow","submittedOnDailyBy":{"_id":"653929a66da48e0d21e65e17","avatarUrl":"/avatars/e34ae3411d689b4280ff34c1b680f283.svg","isPro":false,"fullname":"Byeongjun Park","user":"byeongjun-park","type":"user","name":"byeongjun-park"},"summary":"We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.","upvotes":3,"discussionId":"6a9a24f88f7c3b755723940c","projectPage":"https://byeongjun-park.github.io/FlashRender/","ai_summary":"FlashRender accelerates generative video rendering via representation alignment, a mean-flow objective, and on-policy distillation to achieve high-quality few-step camera-controlled synthesis.","ai_keywords":["FlashRender","generative rendering","sampling-step-dependent camera control","discretization error","Representation Transformation and Alignment (RETA)","visual geometry model","MeanFlow objective","on-policy flow map distillation","few-step sampling"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"64ab689073790912c7a8717a","name":"everex","fullname":"EverEx","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ab6841723beceb2f45c9da/cdIuwsqJw-2tlEhnco0kI.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"653929a66da48e0d21e65e17","avatarUrl":"/avatars/e34ae3411d689b4280ff34c1b680f283.svg","isPro":false,"fullname":"Byeongjun Park","user":"byeongjun-park","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"64ab689073790912c7a8717a","name":"everex","fullname":"EverEx","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64ab6841723beceb2f45c9da/cdIuwsqJw-2tlEhnco0kI.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03563.md","query":{}}">
FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow
Abstract
FlashRender accelerates generative video rendering via representation alignment, a mean-flow objective, and on-policy distillation to achieve high-quality few-step camera-controlled synthesis.
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignment (RETA), which aligns hidden source-video representations with target-video features from a frozen visual geometry model. This directly encodes the geometric transformation within the source-video stream, enabling sampling-step-consistent camera control. We then fine-tune the model with the MeanFlow objective on the lower-curvature denoising trajectory induced by RETA, allowing the model to more effectively address discretization error. Finally, we apply on-policy flow map distillation to correct self-rollout errors under fixed few-step sampling. Extensive experiments show that RETA, MeanFlow, and on-policy flow map distillation play complementary roles in few-step generative rendering. Together, they enable our approach to match multi-step baselines in video quality and geometric consistency at 25x lower sampling cost while achieving superior camera controllability, even under out-of-distribution target camera trajectories.
Community
FlashRender retakes an input video along a target camera trajectory in seconds, using only 4 NFE.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.03563 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.03563 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.03563 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.