Hugging Face Daily Papers · · 5 min read

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.</p>\n","updatedAt":"2026-07-10T03:30:41.617Z","author":{"_id":"6721dacfc5309c08451d21d5","avatarUrl":"/avatars/ac8be5ac8b8ee5b5533214e526b72dad.svg","fullname":"Huang Xuyao","name":"ElysiaTrue","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8479596376419067},"editors":["ElysiaTrue"],"editorAvatarUrls":["/avatars/ac8be5ac8b8ee5b5533214e526b72dad.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.03803","authors":[{"_id":"6a504e6175fd3d966bd45d43","name":"Xuyao Huang","hidden":false},{"_id":"6a504e6175fd3d966bd45d44","name":"Zelai Deng","hidden":false},{"_id":"6a504e6175fd3d966bd45d45","name":"Xu Wang","hidden":false},{"_id":"6a504e6175fd3d966bd45d46","name":"Xizhong Xiao","hidden":false},{"_id":"6a504e6175fd3d966bd45d47","name":"Zhijie Deng","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/6721dacfc5309c08451d21d5/4I-JH8itNF6IuVdUL4tVj.mp4","https://cdn-uploads.huggingface.co/production/uploads/6721dacfc5309c08451d21d5/ihD7yhWWOVjL7w4l0BKFy.mp4","https://cdn-uploads.huggingface.co/production/uploads/6721dacfc5309c08451d21d5/ifxnA0oW8lPq9PXnxJcYG.mp4"],"publishedAt":"2026-07-04T00:00:00.000Z","submittedOnDailyAt":"2026-07-10T00:00:00.000Z","title":"CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation","submittedOnDailyBy":{"_id":"6721dacfc5309c08451d21d5","avatarUrl":"/avatars/ac8be5ac8b8ee5b5533214e526b72dad.svg","isPro":false,"fullname":"Huang Xuyao","user":"ElysiaTrue","type":"user","name":"ElysiaTrue"},"summary":"The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.","upvotes":8,"discussionId":"6a504e6175fd3d966bd45d48","ai_summary":"CineMobile enables efficient image-to-video generation on mobile devices through distillation-guided pruning, diffusion distillation, and hybrid quantization techniques while maintaining visual quality and achieving significant speedup.","ai_keywords":["Diffusion Transformers","distillation-guided pruning","diffusion distillation","reinforcement learning","hybrid post-training quantization","cinematic motion effects","mobile device optimization","video generation","parameter efficiency"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","organization":{"_id":"673d5fe8d031224e947dc235","name":"SJTU-DENG-Lab","fullname":"DENG Lab @ SJTU","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bba541da140e461924dfed/_WPqM9jCqIIkS73aTeZP-.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64bba541da140e461924dfed","avatarUrl":"/avatars/367993765b0ca3734b2b100db33ed787.svg","isPro":true,"fullname":"zhijie deng","user":"zhijie3","type":"user"},{"_id":"627b04e09ef63e604f24d660","avatarUrl":"/avatars/6836932945dbe04c398bec23bcae6262.svg","isPro":false,"fullname":"dooho lee","user":"BlueYellowGreen","type":"user"},{"_id":"6644548a3a16452261cdb173","avatarUrl":"/avatars/4643db904204e3a60202a29e8c884139.svg","isPro":false,"fullname":"wangxu","user":"asunalove","type":"user"},{"_id":"6697e7e55ef2828a1ff371c3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6697e7e55ef2828a1ff371c3/U7-_BtDtSsrf02LIdUTN8.jpeg","isPro":false,"fullname":"Zetong Zhou","user":"Frywind","type":"user"},{"_id":"66d05337e62d6bbf50186c2f","avatarUrl":"/avatars/f5be15e754f0fbbb37d2cc5ea417f729.svg","isPro":false,"fullname":"Yijie Jin","user":"DrewJin0827","type":"user"},{"_id":"69b8babecbe5e9e908451999","avatarUrl":"/avatars/b7c3dfeadef59c1fe7e041e3ea4808b5.svg","isPro":false,"fullname":"Xia Xiao","user":"PolyU-Xia","type":"user"},{"_id":"6182f6630ca58240bd7d8139","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6182f6630ca58240bd7d8139/Ch1-KK89EDydBCSmFKsZj.png","isPro":false,"fullname":"John Pope","user":"johndpope","type":"user"},{"_id":"6311e75405cc08a1408f09c7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6311e75405cc08a1408f09c7/rzppl7_Ihm2Q5LBYlsGzR.jpeg","isPro":false,"fullname":"Xie Qinghe","user":"MikanAffine","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"673d5fe8d031224e947dc235","name":"SJTU-DENG-Lab","fullname":"DENG Lab @ SJTU","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/64bba541da140e461924dfed/_WPqM9jCqIIkS73aTeZP-.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.03803.md","query":{}}">
Papers
arxiv:2607.03803

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

Published on Jul 4
· Submitted by
Huang Xuyao
on Jul 10
Authors:
,

Abstract

CineMobile enables efficient image-to-video generation on mobile devices through distillation-guided pruning, diffusion distillation, and hybrid quantization techniques while maintaining visual quality and achieving significant speedup.

The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.

Community

Paper submitter about 22 hours ago

The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, slow motion, etc. While Diffusion Transformers (DiTs) exhibit strong performance in video generation, their large parameter sizes and multi-step iterative denoising processes lead to substantial computational overhead, making efficient generation on mobile devices challenging. We propose CineMobile to bridge the gap. In particular, CineMobile adopts a three-fold optimization strategy: (1) leveraging a distillation-guided pruning approach to derive a compact yet efficient model that retains the essential video generation capabilities required for cinematic effects; (2) optimizing the compressed model into a 4-step generator via a combination of diffusion distillation and reinforcement learning; (3) employing a hybrid post-training quantization strategy to compress the model footprint to under 1 GB. Experimental results show that compared to the teacher model with the Wan 2.1 architecture, CineMobile achieves a 40x speedup in generation while maintaining comparable visual quality. Specifically, CineMobile generates 49-frame 480p videos with a per-step denoising latency of 0.6s on an NVIDIA H200 GPU and 20s on the MediaTek Dimensity 8400 Ultimate 5G platform, with a peak memory usage of 1.8 GB, demonstrating its practical applicability for mobile-based image-to-video creation.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.03803
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.03803 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.03803 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.03803 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers