Hugging Face Daily Papers · · 5 min read

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

In this work, we present LiveAnimate, which brings 14B-parameter video Diffusion Transformers to real-time streaming inference (19.63 FPS on 2x H100) for interactive full-body animation. We focus on solving the quality degradation and memory explosion typical of long-form diffusion generation. We do this through two key designs: a Reference-Anchored training pipeline with block-wise distillation to hit a 3-step sampling budget, and PR-Sink Attention, a bounded KV-cache mechanism that keeps latency and memory constant over long streams by retrieving historical appearance context based on pose similarity.</p>\n","updatedAt":"2026-08-14T05:11:30.970Z","author":{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","fullname":"Jinpeng YU","name":"Jacob-Yu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8900464773178101},"editors":["Jacob-Yu"],"editorAvatarUrls":["/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.11745","authors":[{"_id":"6a7ea06a42823931a1f1772f","name":"Yuxuan Zhang","hidden":false},{"_id":"6a7ea06a42823931a1f17730","name":"Haozhong Xiong","hidden":false},{"_id":"6a7ea06a42823931a1f17731","name":"Yubo Huang","hidden":false},{"_id":"6a7ea06a42823931a1f17732","name":"Jiayi Song","hidden":false},{"_id":"6a7ea06a42823931a1f17733","user":{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","isPro":false,"fullname":"Jinpeng YU","user":"Jacob-Yu","type":"user","name":"Jacob-Yu"},"name":"Jinpeng Yu","status":"claimed_verified","statusLastChangedAt":"2026-08-14T08:45:04.765Z","hidden":false},{"_id":"6a7ea06a42823931a1f17734","name":"Haofan Wang","hidden":false},{"_id":"6a7ea06a42823931a1f17735","name":"Jiaming Liu","hidden":false},{"_id":"6a7ea06a42823931a1f17736","name":"Ruihua Huang","hidden":false},{"_id":"6a7ea06a42823931a1f17737","name":"Liwei Wang","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time","submittedOnDailyBy":{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","isPro":false,"fullname":"Jinpeng YU","user":"Jacob-Yu","type":"user","name":"Jacob-Yu"},"summary":"Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding responsive interaction. We present LiveAnimate, to our knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT). A two-stage training pipeline first adapts a pretrained bidirectional DiT into a block-causal autoregressive generator through Reference-Anchored Teacher-Forcing Adaptation, and then reduces the sampling budget to three steps through Block-wise Self-Forcing Distillation. To preserve appearance over extended streams, we introduce Pose-Retrieval Sink Attention (PR-Sink), a bounded KV-cache mechanism combining a Static Sink that permanently anchors the first generated block, a Dynamic Sink that holds a pose-retrieved historical block, and a three-slot Rolling Window. When a pose recurs, PR-Sink restores the relevant appearance context without retaining the entire sequence, so memory and per-block latency remain constant regardless of stream duration. Together with Ulysses sequence parallelism and operator fusion, these designs enable 19.63\\,FPS streaming inference on two NVIDIA H100 GPUs. On a three-minute benchmark, LiveAnimate maintains nearly constant perceptual quality and identity from the first 30 seconds to the final minute, while prior systems degrade substantially or require hours of offline computation for the same rollout. These results establish a new operating point in quality, latency, and duration for interactive full-body animation.","upvotes":9,"discussionId":"6a7ea06a42823931a1f17738","projectPage":"https://liveanimate.github.io/","githubRepo":"https://github.com/liveanimate/LiveAnimate","githubRepoAddedBy":"user","ai_summary":"LiveAnimate enables real-time, long-form pose-driven human animation via a 14B-parameter video diffusion transformer with specialized training, bounded attention caching, and sequence parallelism.","ai_keywords":["Diffusion Transformer (DiT)","Reference-Anchored Teacher-Forcing Adaptation","Block-wise Self-Forcing Distillation","Pose-Retrieval Sink Attention (PR-Sink)","Static Sink","Dynamic Sink","Rolling Window","Ulysses sequence parallelism","operator fusion"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":7,"organization":{"_id":"6a6841e7107886ba1a151b03","name":"QwenBusinessUnit","fullname":"Qwen Business Unit","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f79b323fe089b75e9e0c04/MlefZsdry-JuhKzAwxjQl.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","isPro":false,"fullname":"Jinpeng YU","user":"Jacob-Yu","type":"user"},{"_id":"66d53350ad293ffc4b178d10","avatarUrl":"/avatars/d1cb7c800d2af4081de62ef95dc2c44c.svg","isPro":false,"fullname":"jackylova","user":"jackylova","type":"user"},{"_id":"688cceac29c69ec01eb28644","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/wyL1U86pB5GIi4x3JflOa.png","isPro":false,"fullname":"Chuyue Li","user":"woody-woody","type":"user"},{"_id":"67051ad602b48c3fce373fe7","avatarUrl":"/avatars/ebf10382f7c34f41b77f30eb1eea784e.svg","isPro":false,"fullname":"sjy","user":"sjy92","type":"user"},{"_id":"636b3f9ce3ad78bc68b67541","avatarUrl":"/avatars/2b7e745953ae39e01222e99fb63b279e.svg","isPro":false,"fullname":"yuxuan","user":"zzyx","type":"user"},{"_id":"68f9a5cf4cbb5261e6a3cf78","avatarUrl":"/avatars/a0e9bcadd2b69e35ccca6f7d75c2b0b0.svg","isPro":false,"fullname":"Ethan Miller","user":"YaKaYC","type":"user"},{"_id":"69be48e4df85f5f792fc2a7e","avatarUrl":"/avatars/06505fd14a44973c6305ec3f23a67ee4.svg","isPro":false,"fullname":"None","user":"Gonzalo320","type":"user"},{"_id":"65c4eb7cd1dcbd30d86febec","avatarUrl":"/avatars/001c8f02e8ce794b2c21883628b2da72.svg","isPro":false,"fullname":"free-bit","user":"free-bit","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a6841e7107886ba1a151b03","name":"QwenBusinessUnit","fullname":"Qwen Business Unit","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f79b323fe089b75e9e0c04/MlefZsdry-JuhKzAwxjQl.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.11745.md","query":{}}">
Papers
arxiv:2608.11745

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

Published on Aug 13
· Submitted by
Jinpeng YU
on Aug 14
Authors:
,

Abstract

LiveAnimate enables real-time, long-form pose-driven human animation via a 14B-parameter video diffusion transformer with specialized training, bounded attention caching, and sequence parallelism.

Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding responsive interaction. We present LiveAnimate, to our knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT). A two-stage training pipeline first adapts a pretrained bidirectional DiT into a block-causal autoregressive generator through Reference-Anchored Teacher-Forcing Adaptation, and then reduces the sampling budget to three steps through Block-wise Self-Forcing Distillation. To preserve appearance over extended streams, we introduce Pose-Retrieval Sink Attention (PR-Sink), a bounded KV-cache mechanism combining a Static Sink that permanently anchors the first generated block, a Dynamic Sink that holds a pose-retrieved historical block, and a three-slot Rolling Window. When a pose recurs, PR-Sink restores the relevant appearance context without retaining the entire sequence, so memory and per-block latency remain constant regardless of stream duration. Together with Ulysses sequence parallelism and operator fusion, these designs enable 19.63\,FPS streaming inference on two NVIDIA H100 GPUs. On a three-minute benchmark, LiveAnimate maintains nearly constant perceptual quality and identity from the first 30 seconds to the final minute, while prior systems degrade substantially or require hours of offline computation for the same rollout. These results establish a new operating point in quality, latency, and duration for interactive full-body animation.

Community

Paper author Paper submitter about 7 hours ago

In this work, we present LiveAnimate, which brings 14B-parameter video Diffusion Transformers to real-time streaming inference (19.63 FPS on 2x H100) for interactive full-body animation. We focus on solving the quality degradation and memory explosion typical of long-form diffusion generation. We do this through two key designs: a Reference-Anchored training pipeline with block-wise distillation to hit a 3-step sampling budget, and PR-Sink Attention, a bounded KV-cache mechanism that keeps latency and memory constant over long streams by retrieving historical appearance context based on pose similarity.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.11745
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.11745 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.11745 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.11745 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers