Plug-and-play 2D motion interface for real-world Motion Language Models</p>\n<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/67f5d9f4f7a398c2a27166d3/5BX8q7bT-kZ31RJuKGWcD.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>","updatedAt":"2026-08-18T08:59:00.135Z","author":{"_id":"67f5d9f4f7a398c2a27166d3","avatarUrl":"/avatars/4d339fe7b9b760c2f936823f6de9eab2.svg","fullname":"KANAME YOKOYAMA","name":"KanameYOkoYAMA","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.6087909936904907},"editors":["KanameYOkoYAMA"],"editorAvatarUrls":["/avatars/4d339fe7b9b760c2f936823f6de9eab2.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.15984","authors":[{"_id":"6a83bcb8675db694db8cd45d","user":{"_id":"67f5d9f4f7a398c2a27166d3","avatarUrl":"/avatars/4d339fe7b9b760c2f936823f6de9eab2.svg","isPro":false,"fullname":"KANAME YOKOYAMA","user":"KanameYOkoYAMA","type":"user","name":"KanameYOkoYAMA"},"name":"Kaname Yokoyama","status":"claimed_verified","statusLastChangedAt":"2026-08-18T08:45:05.100Z","hidden":false},{"_id":"6a83bcb8675db694db8cd45e","name":"Norimichi Ukita","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models","submittedOnDailyBy":{"_id":"67f5d9f4f7a398c2a27166d3","avatarUrl":"/avatars/4d339fe7b9b760c2f936823f6de9eab2.svg","isPro":false,"fullname":"KANAME YOKOYAMA","user":"KanameYOkoYAMA","type":"user","name":"KanameYOkoYAMA"},"summary":"Motion Language Models (MoLMs) typically understand human motions by tokenizing 3D motion and processing the resulting tokens using a language model. However, obtaining accurate 3D motions from monocular videos is challenging, limiting their real-world applicability. To address this issue, we introduce a plug-and-play 2D Motion Interface that enables 3D-pretrained MoLMs to accept 2D motion inputs without modifying or fine-tuning the original models.\n Experiments on public datasets show that our method achieves performance comparable to 3D motion inputs across multiple MoLMs and outperforms training MoLMs from scratch on 2D motions. We further construct a monocular real-world video motion evaluation dataset and introduce a real-video adapter, demonstrating the usefulness of 2D motions over 3D motions under the evaluated monocular pose-estimation setting. These results suggest that 2D motion provides a practical interface for deploying MoLMs in real-world motion understanding settings. Code is available at https://github.com/irajisamurai/2D-Motion-Interface.","upvotes":1,"discussionId":"6a83bcb8675db694db8cd45f","githubRepo":"https://github.com/irajisamurai/2D-Motion-Interface","githubRepoAddedBy":"user","ai_summary":"A plug-and-play 2D motion interface allows pretrained motion language models to process 2D inputs without retraining, improving real-world applicability.","ai_keywords":["Motion Language Models","2D Motion Interface","3D motion tokenization","monocular video","real-video adapter"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"67f5d9f4f7a398c2a27166d3","avatarUrl":"/avatars/4d339fe7b9b760c2f936823f6de9eab2.svg","isPro":false,"fullname":"KANAME YOKOYAMA","user":"KanameYOkoYAMA","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models
Abstract
A plug-and-play 2D motion interface allows pretrained motion language models to process 2D inputs without retraining, improving real-world applicability.
Motion Language Models (MoLMs) typically understand human motions by tokenizing 3D motion and processing the resulting tokens using a language model. However, obtaining accurate 3D motions from monocular videos is challenging, limiting their real-world applicability. To address this issue, we introduce a plug-and-play 2D Motion Interface that enables 3D-pretrained MoLMs to accept 2D motion inputs without modifying or fine-tuning the original models.
Experiments on public datasets show that our method achieves performance comparable to 3D motion inputs across multiple MoLMs and outperforms training MoLMs from scratch on 2D motions. We further construct a monocular real-world video motion evaluation dataset and introduce a real-video adapter, demonstrating the usefulness of 2D motions over 3D motions under the evaluated monocular pose-estimation setting. These results suggest that 2D motion provides a practical interface for deploying MoLMs in real-world motion understanding settings. Code is available at https://github.com/irajisamurai/2D-Motion-Interface.
Community
Plug-and-play 2D motion interface for real-world Motion Language Models
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.15984 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.