Hugging Face Daily Papers · · 3 min read

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

🚀 We’re excited to release W²-VLA! 🎉<br>Task-conditioned future wrist modeling for fine-grained robot manipulation.<br>&nbsp;<br>🌍 Global task context guides the prediction of task-relevant future wrist latents.<br>🧠 W²-CoT provides structured supervision to help shape the latent interface during training—without CoT decoding at inference.<br>&nbsp;<br>📈 98.5% average success rate on LIBERO
<br>🤖 60.71% Easy / 18.21% Hard on RoboTwin 2.0
<br>⚡ Real-time action generation at over 80 Hz</p>\n","updatedAt":"2026-08-07T03:05:45.998Z","author":{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","fullname":"PENGHAOSONG","name":"HarrisonPENG","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8058125972747803},"editors":["HarrisonPENG"],"editorAvatarUrls":["/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg"],"reactions":[{"reaction":"👍","users":["flameeee","HarrisonPENG"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05369","authors":[{"_id":"6a754b71e1228e04b32381d3","name":"Yuhao Pan","hidden":false},{"_id":"6a754b71e1228e04b32381d4","user":{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","isPro":false,"fullname":"PENGHAOSONG","user":"HarrisonPENG","type":"user","name":"HarrisonPENG"},"name":"Haosong Peng","status":"claimed_verified","statusLastChangedAt":"2026-08-07T16:45:28.179Z","hidden":false},{"_id":"6a754b71e1228e04b32381d5","user":{"_id":"68c14544a9a07d79e3e13166","avatarUrl":"/avatars/3c3f15bccb59d0a56c866b5c100cb35e.svg","isPro":false,"fullname":"Zhengshen Zhang","user":"flameeee","type":"user","name":"flameeee"},"name":"Zhengshen Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-07T08:45:04.545Z","hidden":false},{"_id":"6a754b71e1228e04b32381d6","name":"Zhengyang Yan","hidden":false},{"_id":"6a754b71e1228e04b32381d7","name":"Yalun Dai","hidden":false},{"_id":"6a754b71e1228e04b32381d8","name":"Fushuo Huo","hidden":false},{"_id":"6a754b71e1228e04b32381d9","name":"Chujie Wang","hidden":false},{"_id":"6a754b71e1228e04b32381da","name":"Tianyu Qi","hidden":false},{"_id":"6a754b71e1228e04b32381db","name":"Xiucheng Wang","hidden":false},{"_id":"6a754b71e1228e04b32381dc","name":"Nan Cheng","hidden":false},{"_id":"6a754b71e1228e04b32381dd","name":"Wenchao Xu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/69c8af92851b279c4da20fbb/GpVmryV-KVQfyzNKmRpUj.mp4"],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation","submittedOnDailyBy":{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","isPro":false,"fullname":"PENGHAOSONG","user":"HarrisonPENG","type":"user","name":"HarrisonPENG"},"summary":"Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.","upvotes":14,"discussionId":"6a754b71e1228e04b32381de","projectPage":"https://yyyyu120.github.io/W2-VLA/","githubRepo":"https://github.com/yyyyu120/W2-VLA","githubRepoAddedBy":"user","githubStars":15},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","isPro":false,"fullname":"PENGHAOSONG","user":"HarrisonPENG","type":"user"},{"_id":"64ad2f1f92772101d0394e43","avatarUrl":"/avatars/fa09408b65ed3d03a8aa8ba6965849f4.svg","isPro":false,"fullname":"Dai","user":"dialogueeeeee","type":"user"},{"_id":"64a18ee3d8aea615f31a7e73","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a18ee3d8aea615f31a7e73/BEkWtg6O49ETvVoJVH_pB.jpeg","isPro":false,"fullname":"Haosong Peng","user":"Livioni","type":"user"},{"_id":"63f1c1d72f7c0152e876c70d","avatarUrl":"/avatars/eed1ba2d29ceb7322d5ffdc387c6d11a.svg","isPro":false,"fullname":"Peirong Zheng","user":"zpr","type":"user"},{"_id":"668a6eb6d358e8fd17367126","avatarUrl":"/avatars/63808bd56d4cab9a32e9f67a58b1815b.svg","isPro":false,"fullname":"Z","user":"zg1018","type":"user"},{"_id":"6a1664189dc90b65ed40199b","avatarUrl":"/avatars/752c3f921ef05078bc5c17d3b4b2cd10.svg","isPro":false,"fullname":"ZHANG Meng","user":"truzw77","type":"user"},{"_id":"676139068cd4d1c2b607fa88","avatarUrl":"/avatars/11c254061b2c689d9f240c1261d8af2b.svg","isPro":false,"fullname":"mfy","user":"mmm001","type":"user"},{"_id":"667b8de7a68bf81afe668afe","avatarUrl":"/avatars/aeff10805ff858332e6f6a58735dbbd9.svg","isPro":false,"fullname":"leoli","user":"lifuguan","type":"user"},{"_id":"69d2c3173ce8c7cbe766830c","avatarUrl":"/avatars/e2dc0ed45579a1c2cdecee4edb4327c9.svg","isPro":false,"fullname":"Zien Wang","user":"Kensag","type":"user"},{"_id":"6860e12d23b92536086007c3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6860e12d23b92536086007c3/EdYK574oWqfTxBESgAQlg.jpeg","isPro":false,"fullname":"H","user":"trantor2nd","type":"user"},{"_id":"67dbcfa11033117925b40e34","avatarUrl":"/avatars/a4a9b1b9986b99613ef8b2dfd7bce0f8.svg","isPro":false,"fullname":"Liu Weiqing","user":"Tianhulove","type":"user"},{"_id":"68c14544a9a07d79e3e13166","avatarUrl":"/avatars/3c3f15bccb59d0a56c866b5c100cb35e.svg","isPro":false,"fullname":"Zhengshen Zhang","user":"flameeee","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05369.md","query":{}}">
Papers
arxiv:2608.05369

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Published on Aug 5
· Submitted by
PENGHAOSONG
on Aug 7

Abstract

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.

Community

Paper author Paper submitter about 14 hours ago

🚀 We’re excited to release W²-VLA! 🎉
Task-conditioned future wrist modeling for fine-grained robot manipulation.
 
🌍 Global task context guides the prediction of task-relevant future wrist latents.
🧠 W²-CoT provides structured supervision to help shape the latent interface during training—without CoT decoding at inference.
 
📈 98.5% average success rate on LIBERO

🤖 60.71% Easy / 18.21% Hard on RoboTwin 2.0

⚡ Real-time action generation at over 80 Hz

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.05369
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.05369 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.05369 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.05369 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers