Hugging Face Daily Papers · · 6 min read

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/tcL426LGXbxo4WGaw0LRp.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/tcL426LGXbxo4WGaw0LRp.png\" alt=\"image\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/OIFv2chXvlXQQaQUrQ-Ps.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/65ec01fd770aa0e25d9374dc/OIFv2chXvlXQQaQUrQ-Ps.png\" alt=\"image\"></a></p>\n","updatedAt":"2026-09-17T05:25:51.369Z","author":{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","fullname":"Shijie Lian","name":"LiamLian0727","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":15,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.846102774143219},"editors":["LiamLian0727"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.18487","authors":[{"_id":"6aab792f1d9cc4dec7962551","user":{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","isPro":false,"fullname":"Shijie Lian","user":"LiamLian0727","type":"user","name":"LiamLian0727"},"name":"Shijie Lian","status":"claimed_verified","statusLastChangedAt":"2026-09-17T08:45:04.533Z","hidden":false},{"_id":"6aab792f1d9cc4dec7962552","name":"Bin Yu","hidden":false},{"_id":"6aab792f1d9cc4dec7962553","name":"Zhaolong Shen","hidden":false},{"_id":"6aab792f1d9cc4dec7962554","name":"Xiaopeng Lin","hidden":false},{"_id":"6aab792f1d9cc4dec7962555","name":"Yichao Du","hidden":false},{"_id":"6aab792f1d9cc4dec7962556","name":"Zhirui Zhang","hidden":false},{"_id":"6aab792f1d9cc4dec7962557","name":"Laurence T. Yang","hidden":false},{"_id":"6aab792f1d9cc4dec7962558","name":"Kai Chen","hidden":false}],"publishedAt":"2026-09-16T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models","submittedOnDailyBy":{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","isPro":false,"fullname":"Shijie Lian","user":"LiamLian0727","type":"user","name":"LiamLian0727"},"summary":"Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.","upvotes":36,"discussionId":"6aab792f1d9cc4dec7962559","projectPage":"https://deepcybo-physai.github.io/ActionPiece/","githubRepo":"https://github.com/DeepCybo-PhysAI/ActionPiece","githubRepoAddedBy":"user","githubStars":13,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65ec01fd770aa0e25d9374dc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/yvLWwBEdAdHb-8EdUHg3n.jpeg","isPro":false,"fullname":"Shijie Lian","user":"LiamLian0727","type":"user"},{"_id":"63e1d3451e5a4f34b7a728ef","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e1d3451e5a4f34b7a728ef/kHE5JHF9iZJOG0uBkvJr-.jpeg","isPro":false,"fullname":"Yichao Du","user":"yichaodu","type":"user"},{"_id":"6a686b4f7a167e16851568a9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a686b4f7a167e16851568a9/1ykzaa_g8w8EFIJjYkdZI.jpeg","isPro":false,"fullname":"zhiyao","user":"STOPSUNS","type":"user"},{"_id":"67f8d3be1efce9e5cf4a3a76","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Sf3GAqwIr4eYdv6Mi-G72.png","isPro":false,"fullname":"zy","user":"Sofilyzia","type":"user"},{"_id":"66a1c2d6481acc844c2e7127","avatarUrl":"/avatars/aaf08491a27d0727d5a5cdec10ce06d0.svg","isPro":false,"fullname":"Ziyi Zhang","user":"Mbazze","type":"user"},{"_id":"662d166ba314b134a2e6dd89","avatarUrl":"/avatars/d1bacaa0caa5de609dd69a3328682859.svg","isPro":false,"fullname":"KaiHu","user":"KaiHuUTSC","type":"user"},{"_id":"655494c7914e42998e18060f","avatarUrl":"/avatars/145c765395ce1d2a71da74860a51ea6e.svg","isPro":false,"fullname":"DavidDeng","user":"ZiHDeng","type":"user"},{"_id":"6a03f485883427d8f45eeacc","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/BWYIqDxGMmJ_Rl49OqKyP.jpeg","isPro":false,"fullname":"Zubin Zheng","user":"0SilverBullet","type":"user"},{"_id":"690b0d918c10011327247c2d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/JtHAU2okZ-AlH20kawzqR.png","isPro":false,"fullname":"Zhaolong Shen","user":"majortom1330","type":"user"},{"_id":"6889e312c7792882e9128b29","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/kVizdw-Yfu0rDrf3W6CmW.png","isPro":false,"fullname":"Dingkun Liu","user":"DiMaria0817","type":"user"},{"_id":"63d3b5f1640bb0f77173baea","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674819020331-noauth.jpeg","isPro":false,"fullname":"yubin","user":"VLyb","type":"user"},{"_id":"63e60ff62d704152abac8af8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63e60ff62d704152abac8af8/5kX47xGSmw8sA57O3c_rs.jpeg","isPro":false,"fullname":"Qiuzhi Liu","user":"Dennis364","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6948d884070dda0c2ae35a78","name":"DeepCybo","fullname":"DeepCybo","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ec01fd770aa0e25d9374dc/QOsz6P_7AxyqGrjsRHTGk.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.18487.md","query":{}}">
Papers
arxiv:2609.18487

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

Published on Sep 16
· Submitted by
Shijie Lian
on Sep 17
Authors:

Abstract

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.

Community

Paper author Paper submitter about 10 hours ago

Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.

image

image

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.18487
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2609.18487 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.18487 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2609.18487 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers