🚀 We’re excited to release W²-VLA! 🎉<br>Task-conditioned future wrist modeling for fine-grained robot manipulation.<br> <br>🌍 Global task context guides the prediction of task-relevant future wrist latents.<br>🧠 W²-CoT provides structured supervision to help shape the latent interface during training—without CoT decoding at inference.<br> <br>📈 98.5% average success rate on LIBERO
<br>🤖 60.71% Easy / 18.21% Hard on RoboTwin 2.0
<br>⚡ Real-time action generation at over 80 Hz</p>\n","updatedAt":"2026-08-07T03:05:45.998Z","author":{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","fullname":"PENGHAOSONG","name":"HarrisonPENG","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8058125972747803},"editors":["HarrisonPENG"],"editorAvatarUrls":["/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg"],"reactions":[{"reaction":"👍","users":["flameeee","HarrisonPENG"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05369","authors":[{"_id":"6a754b71e1228e04b32381d3","name":"Yuhao Pan","hidden":false},{"_id":"6a754b71e1228e04b32381d4","user":{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","isPro":false,"fullname":"PENGHAOSONG","user":"HarrisonPENG","type":"user","name":"HarrisonPENG"},"name":"Haosong Peng","status":"claimed_verified","statusLastChangedAt":"2026-08-07T16:45:28.179Z","hidden":false},{"_id":"6a754b71e1228e04b32381d5","user":{"_id":"68c14544a9a07d79e3e13166","avatarUrl":"/avatars/3c3f15bccb59d0a56c866b5c100cb35e.svg","isPro":false,"fullname":"Zhengshen Zhang","user":"flameeee","type":"user","name":"flameeee"},"name":"Zhengshen Zhang","status":"claimed_verified","statusLastChangedAt":"2026-08-07T08:45:04.545Z","hidden":false},{"_id":"6a754b71e1228e04b32381d6","name":"Zhengyang Yan","hidden":false},{"_id":"6a754b71e1228e04b32381d7","name":"Yalun Dai","hidden":false},{"_id":"6a754b71e1228e04b32381d8","name":"Fushuo Huo","hidden":false},{"_id":"6a754b71e1228e04b32381d9","name":"Chujie Wang","hidden":false},{"_id":"6a754b71e1228e04b32381da","name":"Tianyu Qi","hidden":false},{"_id":"6a754b71e1228e04b32381db","name":"Xiucheng Wang","hidden":false},{"_id":"6a754b71e1228e04b32381dc","name":"Nan Cheng","hidden":false},{"_id":"6a754b71e1228e04b32381dd","name":"Wenchao Xu","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/69c8af92851b279c4da20fbb/GpVmryV-KVQfyzNKmRpUj.mp4"],"publishedAt":"2026-08-05T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation","submittedOnDailyBy":{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","isPro":false,"fullname":"PENGHAOSONG","user":"HarrisonPENG","type":"user","name":"HarrisonPENG"},"summary":"Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.","upvotes":14,"discussionId":"6a754b71e1228e04b32381de","projectPage":"https://yyyyu120.github.io/W2-VLA/","githubRepo":"https://github.com/yyyyu120/W2-VLA","githubRepoAddedBy":"user","githubStars":15},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"69c8af92851b279c4da20fbb","avatarUrl":"/avatars/6ae9352bf0b0f9391e3e0d388d4ad5d1.svg","isPro":false,"fullname":"PENGHAOSONG","user":"HarrisonPENG","type":"user"},{"_id":"64ad2f1f92772101d0394e43","avatarUrl":"/avatars/fa09408b65ed3d03a8aa8ba6965849f4.svg","isPro":false,"fullname":"Dai","user":"dialogueeeeee","type":"user"},{"_id":"64a18ee3d8aea615f31a7e73","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64a18ee3d8aea615f31a7e73/BEkWtg6O49ETvVoJVH_pB.jpeg","isPro":false,"fullname":"Haosong Peng","user":"Livioni","type":"user"},{"_id":"63f1c1d72f7c0152e876c70d","avatarUrl":"/avatars/eed1ba2d29ceb7322d5ffdc387c6d11a.svg","isPro":false,"fullname":"Peirong Zheng","user":"zpr","type":"user"},{"_id":"668a6eb6d358e8fd17367126","avatarUrl":"/avatars/63808bd56d4cab9a32e9f67a58b1815b.svg","isPro":false,"fullname":"Z","user":"zg1018","type":"user"},{"_id":"6a1664189dc90b65ed40199b","avatarUrl":"/avatars/752c3f921ef05078bc5c17d3b4b2cd10.svg","isPro":false,"fullname":"ZHANG Meng","user":"truzw77","type":"user"},{"_id":"676139068cd4d1c2b607fa88","avatarUrl":"/avatars/11c254061b2c689d9f240c1261d8af2b.svg","isPro":false,"fullname":"mfy","user":"mmm001","type":"user"},{"_id":"667b8de7a68bf81afe668afe","avatarUrl":"/avatars/aeff10805ff858332e6f6a58735dbbd9.svg","isPro":false,"fullname":"leoli","user":"lifuguan","type":"user"},{"_id":"69d2c3173ce8c7cbe766830c","avatarUrl":"/avatars/e2dc0ed45579a1c2cdecee4edb4327c9.svg","isPro":false,"fullname":"Zien Wang","user":"Kensag","type":"user"},{"_id":"6860e12d23b92536086007c3","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6860e12d23b92536086007c3/EdYK574oWqfTxBESgAQlg.jpeg","isPro":false,"fullname":"H","user":"trantor2nd","type":"user"},{"_id":"67dbcfa11033117925b40e34","avatarUrl":"/avatars/a4a9b1b9986b99613ef8b2dfd7bce0f8.svg","isPro":false,"fullname":"Liu Weiqing","user":"Tianhulove","type":"user"},{"_id":"68c14544a9a07d79e3e13166","avatarUrl":"/avatars/3c3f15bccb59d0a56c866b5c100cb35e.svg","isPro":false,"fullname":"Zhengshen Zhang","user":"flameeee","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05369.md","query":{}}">
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
Abstract
Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.
Community
🚀 We’re excited to release W²-VLA! 🎉
Task-conditioned future wrist modeling for fine-grained robot manipulation.
🌍 Global task context guides the prediction of task-relevant future wrist latents.
🧠 W²-CoT provides structured supervision to help shape the latent interface during training—without CoT decoding at inference.
📈 98.5% average success rate on LIBERO
🤖 60.71% Easy / 18.21% Hard on RoboTwin 2.0
⚡ Real-time action generation at over 80 Hz
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.05369 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.05369 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.05369 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.