Hugging Face Daily Papers · · 4 min read

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Website: <a href=\"https://robodojo-benchmark.com/\" rel=\"nofollow\">https://robodojo-benchmark.com/</a><br>arXiv: <a href=\"https://arxiv.org/abs/2607.04434\" rel=\"nofollow\">https://arxiv.org/abs/2607.04434</a><br>Leaderboard: <a href=\"https://robodojo-benchmark.com/LeaderBoard\" rel=\"nofollow\">https://robodojo-benchmark.com/LeaderBoard</a><br>Benchmark code: <a href=\"https://github.com/RoboDojo-Benchmark/RoboDojo\" rel=\"nofollow\">https://github.com/RoboDojo-Benchmark/RoboDojo</a><br>XPolicyLab code: <a href=\"https://github.com/XPolicyLab/XPolicyLab\" rel=\"nofollow\">https://github.com/XPolicyLab/XPolicyLab</a><br>Community: <a href=\"https://robodojo-benchmark.com/community\" rel=\"nofollow\">https://robodojo-benchmark.com/community</a></p>\n","updatedAt":"2026-07-09T02:48:48.892Z","author":{"_id":"65b37a9b06d8b55123ef8921","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg","fullname":"Tianxing Chen","name":"TianxingChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":17,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.5292746424674988},"editors":["TianxingChen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.04434","authors":[{"_id":"6a4dcd4e25849b193a834c06","name":"Tianxing Chen","hidden":false},{"_id":"6a4dcd4e25849b193a834c07","name":"Yue Chen","hidden":false},{"_id":"6a4dcd4e25849b193a834c08","name":"Zixuan Li","hidden":false},{"_id":"6a4dcd4e25849b193a834c09","name":"Junyuan Tang","hidden":false},{"_id":"6a4dcd4e25849b193a834c0a","name":"Kailun Su","hidden":false},{"_id":"6a4dcd4e25849b193a834c0b","name":"Haoran Lu","hidden":false},{"_id":"6a4dcd4e25849b193a834c0c","name":"Weijie Wan","hidden":false},{"_id":"6a4dcd4e25849b193a834c0d","name":"Baijun Chen","hidden":false},{"_id":"6a4dcd4e25849b193a834c0e","name":"Songling Liu","hidden":false},{"_id":"6a4dcd4e25849b193a834c0f","name":"Haowen Yan","hidden":false},{"_id":"6a4dcd4e25849b193a834c10","name":"Honghao Su","hidden":false},{"_id":"6a4dcd4e25849b193a834c11","name":"Zhiyang Dou","hidden":false},{"_id":"6a4dcd4e25849b193a834c12","name":"Kaixuan Wang","hidden":false},{"_id":"6a4dcd4e25849b193a834c13","name":"Dandan Zhang","hidden":false},{"_id":"6a4dcd4e25849b193a834c14","name":"Yunze Liu","hidden":false},{"_id":"6a4dcd4e25849b193a834c15","name":"Yan Qin","hidden":false},{"_id":"6a4dcd4e25849b193a834c16","name":"Qiwei Liang","hidden":false},{"_id":"6a4dcd4e25849b193a834c17","name":"Qiwei Wu","hidden":false},{"_id":"6a4dcd4e25849b193a834c18","name":"Zijian Lin","hidden":false},{"_id":"6a4dcd4e25849b193a834c19","name":"Wenwei Lin","hidden":false},{"_id":"6a4dcd4e25849b193a834c1a","name":"Yuran Wang","hidden":false},{"_id":"6a4dcd4e25849b193a834c1b","name":"Minghua He","hidden":false},{"_id":"6a4dcd4e25849b193a834c1c","name":"Tianshu Wu","hidden":false},{"_id":"6a4dcd4e25849b193a834c1d","name":"Ruihai Wu","hidden":false},{"_id":"6a4dcd4e25849b193a834c1e","name":"Jingquan Zhou","hidden":false},{"_id":"6a4dcd4e25849b193a834c1f","name":"Kai-Chong Lei","hidden":false},{"_id":"6a4dcd4e25849b193a834c20","name":"Haibao Yu","hidden":false},{"_id":"6a4dcd4e25849b193a834c21","name":"Yuanfeng Ji","hidden":false},{"_id":"6a4dcd4e25849b193a834c22","name":"Weiyang Jin","hidden":false},{"_id":"6a4dcd4e25849b193a834c23","name":"Guanyu Lin","hidden":false},{"_id":"6a4dcd4e25849b193a834c24","name":"Xiaofan Li","hidden":false},{"_id":"6a4dcd4e25849b193a834c25","name":"Qi Xiong","hidden":false},{"_id":"6a4dcd4e25849b193a834c26","name":"Renjing Xu","hidden":false},{"_id":"6a4dcd4e25849b193a834c27","name":"Zhongyu Li","hidden":false},{"_id":"6a4dcd4e25849b193a834c28","name":"Wenhao Chai","hidden":false},{"_id":"6a4dcd4e25849b193a834c29","name":"Enze Xie","hidden":false},{"_id":"6a4dcd4e25849b193a834c2a","name":"Ziwei Wang","hidden":false},{"_id":"6a4dcd4e25849b193a834c2b","name":"Yao Mu","hidden":false},{"_id":"6a4dcd4e25849b193a834c2c","name":"Hao Dong","hidden":false},{"_id":"6a4dcd4e25849b193a834c2d","name":"Wojciech Matusik","hidden":false},{"_id":"6a4dcd4e25849b193a834c2e","name":"Mingyu Ding","hidden":false},{"_id":"6a4dcd4e25849b193a834c2f","name":"Wenbo Ding","hidden":false},{"_id":"6a4dcd4e25849b193a834c30","name":"Ping Luo","hidden":false},{"_id":"6a4dcd4e25849b193a834c31","name":"Masayoshi Tomizuka","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65b37a9b06d8b55123ef8921/M-5QhZgUyukN0I4dEpyL8.mp4"],"publishedAt":"2026-07-07T00:00:00.000Z","submittedOnDailyAt":"2026-07-09T00:00:00.000Z","title":"RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies","submittedOnDailyBy":{"_id":"65b37a9b06d8b55123ef8921","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg","isPro":false,"fullname":"Tianxing Chen","user":"TianxingChen","type":"user","name":"TianxingChen"},"summary":"Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.","upvotes":4,"discussionId":"6a4dcd4f25849b193a834c32","projectPage":"https://robodojo-benchmark.com/","githubRepo":"https://github.com/RoboDojo-Benchmark/RoboDojo","githubRepoAddedBy":"user","ai_summary":"RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions.","ai_keywords":[""],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":92,"organization":{"_id":"6a4a92e27391729bb22ddf43","name":"RoboDojo-Benchmark","fullname":"RoboDojo-Benchmark","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67cc5e0232aeea9209d35033/evaqVzuHh_3XFQ5HRSxo5.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65b37a9b06d8b55123ef8921","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg","isPro":false,"fullname":"Tianxing Chen","user":"TianxingChen","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"642e97dbc1b0f8e4e76c2b30","avatarUrl":"/avatars/60adf4470baf12d5687d53a6c3299bcd.svg","isPro":false,"fullname":"james curry","user":"ainbo","type":"user"},{"_id":"645d7f107c7258d904e82749","avatarUrl":"/avatars/a4e9d47b281f18616c522c1a8b8ee7e5.svg","isPro":false,"fullname":"HuichiZhou","user":"Zhouhc","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a4a92e27391729bb22ddf43","name":"RoboDojo-Benchmark","fullname":"RoboDojo-Benchmark","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67cc5e0232aeea9209d35033/evaqVzuHh_3XFQ5HRSxo5.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.04434.md","query":{}}">
Papers
arxiv:2607.04434

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Published on Jul 7
· Submitted by
Tianxing Chen
on Jul 9
Authors:
,

Abstract

RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions.

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.04434
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.04434 in a model README.md to link it from this page.

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.04434 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers