Hugging Face Daily Papers · · 6 min read

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.</p>\n","updatedAt":"2026-09-10T02:30:30.966Z","author":{"_id":"65b43d30625ac670a71cbbf3","avatarUrl":"/avatars/1d6366ba0a7a829ed4d5cc483f4335ac.svg","fullname":"Vida Adeli","name":"vida-adl","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8679395914077759},"editors":["vida-adl"],"editorAvatarUrls":["/avatars/1d6366ba0a7a829ed4d5cc483f4335ac.svg"],"reactions":[],"isReport":false}},{"id":"6aa21a8461226b1f2dfc0c4f","author":{"_id":"65b43d30625ac670a71cbbf3","avatarUrl":"/avatars/1d6366ba0a7a829ed4d5cc483f4335ac.svg","fullname":"Vida Adeli","name":"vida-adl","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false},"createdAt":"2026-09-10T02:48:36.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"🎥 A quick look at Puppeteer!\nWe generate co-speech gestures that adapt not only to speech, but also to the character’s posture and surrounding objects, enabling physically grounded interactions across different scene configurations.\n\nMore details, results, and demos in the paper and project page.\n\nhttps://cdn-uploads.huggingface.co/production/uploads/65b43d30625ac670a71cbbf3/iGEEgOQVoOV4-2uf83vUb.mp4\n","html":"<p>🎥 A quick look at Puppeteer!<br>We generate co-speech gestures that adapt not only to speech, but also to the character’s posture and surrounding objects, enabling physically grounded interactions across different scene configurations.</p>\n<p>More details, results, and demos in the paper and project page.</p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/65b43d30625ac670a71cbbf3/iGEEgOQVoOV4-2uf83vUb.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-09-10T02:48:36.179Z","author":{"_id":"65b43d30625ac670a71cbbf3","avatarUrl":"/avatars/1d6366ba0a7a829ed4d5cc483f4335ac.svg","fullname":"Vida Adeli","name":"vida-adl","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.854752779006958},"editors":["vida-adl"],"editorAvatarUrls":["/avatars/1d6366ba0a7a829ed4d5cc483f4335ac.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.00369","authors":[{"_id":"6aa180e1d8c54e38c0a36a70","name":"Vida Adeli","hidden":false},{"_id":"6aa180e1d8c54e38c0a36a71","name":"Soroush Mehraban","hidden":false},{"_id":"6aa180e1d8c54e38c0a36a72","name":"Jacob Rommann","hidden":false},{"_id":"6aa180e1d8c54e38c0a36a73","name":"Harrison Sanborn","hidden":false},{"_id":"6aa180e1d8c54e38c0a36a74","name":"Cole Clifford","hidden":false},{"_id":"6aa180e1d8c54e38c0a36a75","name":"Babak Taati","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65b43d30625ac670a71cbbf3/LB8JQDuK4voTUa-UxuVmm.webp"],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-10T00:00:00.000Z","title":"Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation","submittedOnDailyBy":{"_id":"65b43d30625ac670a71cbbf3","avatarUrl":"/avatars/1d6366ba0a7a829ed4d5cc483f4335ac.svg","isPro":false,"fullname":"Vida Adeli","user":"vida-adl","type":"user","name":"vida-adl"},"summary":"Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.","upvotes":1,"discussionId":"6aa180e1d8c54e38c0a36a76","projectPage":"https://puppeteer.pickford.ai/","ai_summary":"Puppeteer is a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, physically grounded gestures.","ai_keywords":["co-speech gesture diffusion model","causal variational autoencoder","latent tokens","conditional diffusion","posture-aware","object-grounded","gesture primitives","temporal control","gesture in-betweening","SceneGes"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"67a5281da26e98ac622de2b4","name":"Pickford","fullname":"Pickford","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67a527a1c12b766246a29c63/1sJ-w4TlSPNvhrjxpmxlq.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65b43d30625ac670a71cbbf3","avatarUrl":"/avatars/1d6366ba0a7a829ed4d5cc483f4335ac.svg","isPro":false,"fullname":"Vida Adeli","user":"vida-adl","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67a5281da26e98ac622de2b4","name":"Pickford","fullname":"Pickford","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67a527a1c12b766246a29c63/1sJ-w4TlSPNvhrjxpmxlq.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.00369.md","query":{}}">
Papers
arxiv:2609.00369

Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

Published on Aug 31
· Submitted by
Vida Adeli
on Sep 10
Authors:
,

Abstract

Puppeteer is a diffusion-based co-speech gesture model that uses causal latent tokens and object geometry to generate temporally coherent, physically grounded gestures.

Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.

Community

Paper submitter about 6 hours ago

Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.

Paper submitter about 5 hours ago

🎥 A quick look at Puppeteer!
We generate co-speech gestures that adapt not only to speech, but also to the character’s posture and surrounding objects, enabling physically grounded interactions across different scene configurations.

More details, results, and demos in the paper and project page.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.00369
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.00369 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.00369 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.00369 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers