Website: <a href=\"https://robodojo-benchmark.com/\" rel=\"nofollow\">https://robodojo-benchmark.com/</a><br>arXiv: <a href=\"https://arxiv.org/abs/2607.04434\" rel=\"nofollow\">https://arxiv.org/abs/2607.04434</a><br>Leaderboard: <a href=\"https://robodojo-benchmark.com/LeaderBoard\" rel=\"nofollow\">https://robodojo-benchmark.com/LeaderBoard</a><br>Benchmark code: <a href=\"https://github.com/RoboDojo-Benchmark/RoboDojo\" rel=\"nofollow\">https://github.com/RoboDojo-Benchmark/RoboDojo</a><br>XPolicyLab code: <a href=\"https://github.com/XPolicyLab/XPolicyLab\" rel=\"nofollow\">https://github.com/XPolicyLab/XPolicyLab</a><br>Community: <a href=\"https://robodojo-benchmark.com/community\" rel=\"nofollow\">https://robodojo-benchmark.com/community</a></p>\n","updatedAt":"2026-07-09T02:48:48.892Z","author":{"_id":"65b37a9b06d8b55123ef8921","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg","fullname":"Tianxing Chen","name":"TianxingChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":17,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.5292746424674988},"editors":["TianxingChen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.04434","authors":[{"_id":"6a4dcd4e25849b193a834c06","name":"Tianxing Chen","hidden":false},{"_id":"6a4dcd4e25849b193a834c07","name":"Yue Chen","hidden":false},{"_id":"6a4dcd4e25849b193a834c08","name":"Zixuan Li","hidden":false},{"_id":"6a4dcd4e25849b193a834c09","name":"Junyuan Tang","hidden":false},{"_id":"6a4dcd4e25849b193a834c0a","name":"Kailun Su","hidden":false},{"_id":"6a4dcd4e25849b193a834c0b","name":"Haoran Lu","hidden":false},{"_id":"6a4dcd4e25849b193a834c0c","name":"Weijie Wan","hidden":false},{"_id":"6a4dcd4e25849b193a834c0d","name":"Baijun Chen","hidden":false},{"_id":"6a4dcd4e25849b193a834c0e","name":"Songling Liu","hidden":false},{"_id":"6a4dcd4e25849b193a834c0f","name":"Haowen Yan","hidden":false},{"_id":"6a4dcd4e25849b193a834c10","name":"Honghao Su","hidden":false},{"_id":"6a4dcd4e25849b193a834c11","name":"Zhiyang Dou","hidden":false},{"_id":"6a4dcd4e25849b193a834c12","name":"Kaixuan Wang","hidden":false},{"_id":"6a4dcd4e25849b193a834c13","name":"Dandan Zhang","hidden":false},{"_id":"6a4dcd4e25849b193a834c14","name":"Yunze Liu","hidden":false},{"_id":"6a4dcd4e25849b193a834c15","name":"Yan Qin","hidden":false},{"_id":"6a4dcd4e25849b193a834c16","name":"Qiwei Liang","hidden":false},{"_id":"6a4dcd4e25849b193a834c17","name":"Qiwei Wu","hidden":false},{"_id":"6a4dcd4e25849b193a834c18","name":"Zijian Lin","hidden":false},{"_id":"6a4dcd4e25849b193a834c19","name":"Wenwei Lin","hidden":false},{"_id":"6a4dcd4e25849b193a834c1a","name":"Yuran Wang","hidden":false},{"_id":"6a4dcd4e25849b193a834c1b","name":"Minghua He","hidden":false},{"_id":"6a4dcd4e25849b193a834c1c","name":"Tianshu Wu","hidden":false},{"_id":"6a4dcd4e25849b193a834c1d","name":"Ruihai Wu","hidden":false},{"_id":"6a4dcd4e25849b193a834c1e","name":"Jingquan Zhou","hidden":false},{"_id":"6a4dcd4e25849b193a834c1f","name":"Kai-Chong Lei","hidden":false},{"_id":"6a4dcd4e25849b193a834c20","name":"Haibao Yu","hidden":false},{"_id":"6a4dcd4e25849b193a834c21","name":"Yuanfeng Ji","hidden":false},{"_id":"6a4dcd4e25849b193a834c22","name":"Weiyang Jin","hidden":false},{"_id":"6a4dcd4e25849b193a834c23","name":"Guanyu Lin","hidden":false},{"_id":"6a4dcd4e25849b193a834c24","name":"Xiaofan Li","hidden":false},{"_id":"6a4dcd4e25849b193a834c25","name":"Qi Xiong","hidden":false},{"_id":"6a4dcd4e25849b193a834c26","name":"Renjing Xu","hidden":false},{"_id":"6a4dcd4e25849b193a834c27","name":"Zhongyu Li","hidden":false},{"_id":"6a4dcd4e25849b193a834c28","name":"Wenhao Chai","hidden":false},{"_id":"6a4dcd4e25849b193a834c29","name":"Enze Xie","hidden":false},{"_id":"6a4dcd4e25849b193a834c2a","name":"Ziwei Wang","hidden":false},{"_id":"6a4dcd4e25849b193a834c2b","name":"Yao Mu","hidden":false},{"_id":"6a4dcd4e25849b193a834c2c","name":"Hao Dong","hidden":false},{"_id":"6a4dcd4e25849b193a834c2d","name":"Wojciech Matusik","hidden":false},{"_id":"6a4dcd4e25849b193a834c2e","name":"Mingyu Ding","hidden":false},{"_id":"6a4dcd4e25849b193a834c2f","name":"Wenbo Ding","hidden":false},{"_id":"6a4dcd4e25849b193a834c30","name":"Ping Luo","hidden":false},{"_id":"6a4dcd4e25849b193a834c31","name":"Masayoshi Tomizuka","hidden":false}],"mediaUrls":["https://cdn-uploads.huggingface.co/production/uploads/65b37a9b06d8b55123ef8921/M-5QhZgUyukN0I4dEpyL8.mp4"],"publishedAt":"2026-07-07T00:00:00.000Z","submittedOnDailyAt":"2026-07-09T00:00:00.000Z","title":"RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies","submittedOnDailyBy":{"_id":"65b37a9b06d8b55123ef8921","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg","isPro":false,"fullname":"Tianxing Chen","user":"TianxingChen","type":"user","name":"TianxingChen"},"summary":"Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.","upvotes":4,"discussionId":"6a4dcd4f25849b193a834c32","projectPage":"https://robodojo-benchmark.com/","githubRepo":"https://github.com/RoboDojo-Benchmark/RoboDojo","githubRepoAddedBy":"user","ai_summary":"RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions.","ai_keywords":[""],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":92,"organization":{"_id":"6a4a92e27391729bb22ddf43","name":"RoboDojo-Benchmark","fullname":"RoboDojo-Benchmark","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67cc5e0232aeea9209d35033/evaqVzuHh_3XFQ5HRSxo5.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65b37a9b06d8b55123ef8921","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65b37a9b06d8b55123ef8921/CT5tLwezjXct1eTszA8sO.jpeg","isPro":false,"fullname":"Tianxing Chen","user":"TianxingChen","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"642e97dbc1b0f8e4e76c2b30","avatarUrl":"/avatars/60adf4470baf12d5687d53a6c3299bcd.svg","isPro":false,"fullname":"james curry","user":"ainbo","type":"user"},{"_id":"645d7f107c7258d904e82749","avatarUrl":"/avatars/a4e9d47b281f18616c522c1a8b8ee7e5.svg","isPro":false,"fullname":"HuichiZhou","user":"Zhouhc","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a4a92e27391729bb22ddf43","name":"RoboDojo-Benchmark","fullname":"RoboDojo-Benchmark","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67cc5e0232aeea9209d35033/evaqVzuHh_3XFQ5HRSxo5.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.04434.md","query":{}}">
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Abstract
RoboDojo presents a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across diverse tasks and evaluation dimensions.
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage, and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce. We introduce RoboDojo, a unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies. RoboDojo includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions. RoboDojo supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. Together with XPolicyLab, policies can be integrated once and evaluated across simulation and real-world settings with minimal adaptation. We integrate 30 policies into XPolicyLab and evaluate them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance. The website is available at http://robodojo-benchmark.com/.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.04434 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.04434 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.