Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/tcL426LGXbxo4WGaw0LRp.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/tcL426LGXbxo4WGaw0LRp.png\" alt=\"image\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/OIFv2chXvlXQQaQUrQ-Ps.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/OIFv2chXvlXQQaQUrQ-Ps.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-09-17T05:25:51.369Z","author":{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","fullname":"Shijie Lian","name":"LiamLian0727","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":15,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.846102774143219},"editors":["LiamLian0727"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.18487","authors":[{"_id":"6aab792f1d9cc4dec7962551","user":{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","isPro":false,"fullname":"Shijie Lian","user":"LiamLian0727","type":"user","name":"LiamLian0727"},"name":"Shijie Lian","status":"claimed_verified","statusLastChangedAt":"2026-09-17T08:45:04.533Z","hidden":false},{"_id":"6aab792f1d9cc4dec7962552","name":"Bin Yu","hidden":false},{"_id":"6aab792f1d9cc4dec7962553","name":"Zhaolong Shen","hidden":false},{"_id":"6aab792f1d9cc4dec7962554","name":"Xiaopeng Lin","hidden":false},{"_id":"6aab792f1d9cc4dec7962555","name":"Yichao Du","hidden":false},{"_id":"6aab792f1d9cc4dec7962556","name":"Zhirui Zhang","hidden":false},{"_id":"6aab792f1d9cc4dec7962557","name":"Laurence T. Yang","hidden":false},{"_id":"6aab792f1d9cc4dec7962558","name":"Kai Chen","hidden":false}],"publishedAt":"2026-09-16T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models","submittedOnDailyBy":{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","isPro":false,"fullname":"Shijie Lian","user":"LiamLian0727","type":"user","name":"LiamLian0727"},"summary":"Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.","upvotes":36,"discussionId":"6aab792f1d9cc4dec7962559","projectPage":"https://deepcybo-physai.github.io/ActionPiece/","githubRepo":"https://github.com/DeepCybo-PhysAI/ActionPiece","githubRepoAddedBy":"user","githubStars":13,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","isPro":false,"fullname":"Shijie Lian","user":"LiamLian0727","type":"user"},{"_id":"63e1d3451e5a4f34b7a728ef","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e1d3451e5a4f34b7a728ef/kHE5JHF9iZJOG0uBkvJr-.jpeg","isPro":false,"fullname":"Yichao Du","user":"yichaodu","type":"user"},{"_id":"6a686b4f7a167e16851568a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a686b4f7a167e16851568a9/1ykzaa_g8w8EFIJjYkdZI.jpeg","isPro":false,"fullname":"zhiyao","user":"STOPSUNS","type":"user"},{"_id":"67f8d3be1efce9e5cf4a3a76","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Sf3GAqwIr4eYdv6Mi-G72.png","isPro":false,"fullname":"zy","user":"Sofilyzia","type":"user"},{"_id":"66a1c2d6481acc844c2e7127","avatarUrl":"/avatars/aaf08491a27d0727d5a5cdec10ce06d0.svg","isPro":false,"fullname":"Ziyi Zhang","user":"Mbazze","type":"user"},{"_id":"662d166ba314b134a2e6dd89","avatarUrl":"/avatars/d1bacaa0caa5de609dd69a3328682859.svg","isPro":false,"fullname":"KaiHu","user":"KaiHuUTSC","type":"user"},{"_id":"655494c7914e42998e18060f","avatarUrl":"/avatars/145c765395ce1d2a71da74860a51ea6e.svg","isPro":false,"fullname":"DavidDeng","user":"ZiHDeng","type":"user"},{"_id":"6a03f485883427d8f45eeacc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/BWYIqDxGMmJ_Rl49OqKyP.jpeg","isPro":false,"fullname":"Zubin Zheng","user":"0SilverBullet","type":"user"},{"_id":"690b0d918c10011327247c2d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JtHAU2okZ-AlH20kawzqR.png","isPro":false,"fullname":"Zhaolong Shen","user":"majortom1330","type":"user"},{"_id":"6889e312c7792882e9128b29","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/kVizdw-Yfu0rDrf3W6CmW.png","isPro":false,"fullname":"Dingkun Liu","user":"DiMaria0817","type":"user"},{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","isPro":false,"fullname":"yubin","user":"VLyb","type":"user"},{"_id":"63e60ff62d704152abac8af8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e60ff62d704152abac8af8/5kX47xGSmw8sA57O3c_rs.jpeg","isPro":false,"fullname":"Qiuzhi Liu","user":"Dennis364","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.18487.md","query":{}}">
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
Abstract
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
Community
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.


Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.18487 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.18487 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.18487 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.