Hugging Face Daily Papers · · 6 min read

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it difficult to identify the root causes of agent failures. To address these limitations, we introduce UniClawBench, the first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings. UniClawBench is built around five foundational model capabilities: Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, and Cross-Platform Coordination. Based on these capabilities, we design 400 bilingual real-world tasks. Unlike previous benchmarks that rely on static, pre-recorded answers, our benchmark evaluates agents in live Docker containers using fine-grained, step-by-step completion checkpoints. Furthermore, we design a closed-loop evaluation strategy comprising an executor agent, a hidden supervisor agent, and a user agent to simulate realistic multi-turn human feedback without leaking grading criteria. To disentangle base model capabilities from framework-level design choices, we evaluate state-of-the-art models under multiple agent frameworks. Through comprehensive comparisons across both models and frameworks, we show how base model capabilities and agent framework designs jointly shape performance in real-world environments. To facilitate future research, we make our benchmark and code publicly available at <a href=\"https://github.com/HKU-MMLab/UniClawBench\" rel=\"nofollow\">https://github.com/HKU-MMLab/UniClawBench</a></p>\n","updatedAt":"2026-07-10T04:13:58.201Z","author":{"_id":"62d812e143df7719860d05d1","avatarUrl":"/avatars/412f7ec5c9f54990f4b562652d3e2c59.svg","fullname":"zhekai chen","name":"Azily","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8888607621192932},"editors":["Azily"],"editorAvatarUrls":["/avatars/412f7ec5c9f54990f4b562652d3e2c59.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.08768","authors":[{"_id":"6a506fe675fd3d966bd45e15","name":"Zhekai Chen","hidden":false},{"_id":"6a506fe675fd3d966bd45e16","name":"Chengqi Duan","hidden":false},{"_id":"6a506fe675fd3d966bd45e17","name":"Kaiyue Sun","hidden":false},{"_id":"6a506fe675fd3d966bd45e18","name":"Bohao Li","hidden":false},{"_id":"6a506fe675fd3d966bd45e19","name":"Yuqing Wang","hidden":false},{"_id":"6a506fe675fd3d966bd45e1a","name":"Manyuan Zhang","hidden":false},{"_id":"6a506fe675fd3d966bd45e1b","name":"Xihui Liu","hidden":false}],"publishedAt":"2026-07-09T00:00:00.000Z","submittedOnDailyAt":"2026-07-10T00:00:00.000Z","title":"UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks","submittedOnDailyBy":{"_id":"62d812e143df7719860d05d1","avatarUrl":"/avatars/412f7ec5c9f54990f4b562652d3e2c59.svg","isPro":false,"fullname":"zhekai chen","user":"Azily","type":"user","name":"Azily"},"summary":"The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it difficult to identify the root causes of agent failures. To address these limitations, we introduce UniClawBench, the first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings. UniClawBench is built around five foundational model capabilities: Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, and Cross-Platform Coordination. Based on these capabilities, we design 400 bilingual real-world tasks. Unlike previous benchmarks that rely on static, pre-recorded answers, our benchmark evaluates agents in live Docker containers using fine-grained, step-by-step completion checkpoints. Furthermore, we design a closed-loop evaluation strategy comprising an executor agent, a hidden supervisor agent, and a user agent to simulate realistic multi-turn human feedback without leaking grading criteria. To disentangle base model capabilities from framework-level design choices, we evaluate state-of-the-art models under multiple agent frameworks. Through comprehensive comparisons across both models and frameworks, we show how base model capabilities and agent framework designs jointly shape performance in real-world environments. To facilitate future research, we make our benchmark and code publicly available at https://github.com/HKU-MMLab/UniClawBench.","upvotes":22,"discussionId":"6a506fe775fd3d966bd45e1c","projectPage":"https://uniclawbench.github.io/","githubRepo":"https://github.com/HKU-MMLab/UniClawBench","githubRepoAddedBy":"user","ai_summary":"UniClawBench introduces a capability-driven benchmark for evaluating proactive agents in real-world environments using live Docker container evaluation and closed-loop assessment with multiple agent roles.","ai_keywords":["large language models","multimodal large language models","proactive agents","real-world environments","capability-driven benchmark","skill usage","exploration","long-context reasoning","multimodal understanding","cross-platform coordination","Docker containers","closed-loop evaluation","executor agent","supervisor agent","user agent"],"ai_summary_model":"Qwen/Qwen2.5-Coder-32B-Instruct","githubStars":20,"organization":{"_id":"67ea9ecfc234715db8dbf339","name":"hkuhk","fullname":"The University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ea9e8d2d95c10a0da11b0c/FNnR4M7YqKRuG43N5771B.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62d812e143df7719860d05d1","avatarUrl":"/avatars/412f7ec5c9f54990f4b562652d3e2c59.svg","isPro":false,"fullname":"zhekai chen","user":"Azily","type":"user"},{"_id":"63ea23b9dedfeebe54d02bdf","avatarUrl":"/avatars/4d9f9a546aa8c63e277161ea700075c4.svg","isPro":false,"fullname":"Yuqing Wang","user":"Epiphqny","type":"user"},{"_id":"667a7f1d3b78e49a81ab02c2","avatarUrl":"/avatars/b60cb5b9070c7573a7f241407d705ecc.svg","isPro":false,"fullname":"Zhihang Liu","user":"lntzm","type":"user"},{"_id":"6492a0d8d4ae24c933ace44d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6492a0d8d4ae24c933ace44d/FXYIucGnWkMDu4gHWy1qw.jpeg","isPro":false,"fullname":"Longxiang Tang","user":"lloong","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64a2b496e2e19de17db7de65","avatarUrl":"/avatars/241448ca487833d6cc5d57bb1fdb6ee5.svg","isPro":false,"fullname":"Duan Chengqi","user":"gogoduan","type":"user"},{"_id":"64d5c6acdd57652c1a472f2d","avatarUrl":"/avatars/358ea808645b5bf72dd82b07cacf7a78.svg","isPro":false,"fullname":"Xiong Xuyuan","user":"xjxyys","type":"user"},{"_id":"66ce751a8ec9fda2cf5a9e85","avatarUrl":"/avatars/c17093ca81dad007b3e50bae503955a7.svg","isPro":false,"fullname":"Haocheng Xi","user":"xihc-ucb","type":"user"},{"_id":"6485b08e687d9e0c759121b0","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6485b08e687d9e0c759121b0/P_9F0izrQgUfEd-VEbhg8.jpeg","isPro":false,"fullname":"sijin","user":"CH3COOK","type":"user"},{"_id":"60d045c4778bafd0fbcfa3f5","avatarUrl":"/avatars/0cc0c2739c1934430ea09df7e9668c80.svg","isPro":false,"fullname":"Yi Chen","user":"ChenYi99","type":"user"},{"_id":"637cba13b8e573d75be96ea6","avatarUrl":"/avatars/5eca230e63d66947b2a05c1ff964a96c.svg","isPro":false,"fullname":"Nina","user":"NinaKarine","type":"user"},{"_id":"6310b7e70a43f97f6c56191e","avatarUrl":"/avatars/4a24c76e34d12c3d6230a4a081115f72.svg","isPro":false,"fullname":"Bohao Li","user":"BreakLee","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67ea9ecfc234715db8dbf339","name":"hkuhk","fullname":"The University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/67ea9e8d2d95c10a0da11b0c/FNnR4M7YqKRuG43N5771B.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.08768.md","query":{}}">
Papers
arxiv:2607.08768

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

Published on Jul 9
· Submitted by
zhekai chen
on Jul 10
Authors:
,

Abstract

UniClawBench introduces a capability-driven benchmark for evaluating proactive agents in real-world environments using live Docker container evaluation and closed-loop assessment with multiple agent roles.

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it difficult to identify the root causes of agent failures. To address these limitations, we introduce UniClawBench, the first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings. UniClawBench is built around five foundational model capabilities: Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, and Cross-Platform Coordination. Based on these capabilities, we design 400 bilingual real-world tasks. Unlike previous benchmarks that rely on static, pre-recorded answers, our benchmark evaluates agents in live Docker containers using fine-grained, step-by-step completion checkpoints. Furthermore, we design a closed-loop evaluation strategy comprising an executor agent, a hidden supervisor agent, and a user agent to simulate realistic multi-turn human feedback without leaking grading criteria. To disentangle base model capabilities from framework-level design choices, we evaluate state-of-the-art models under multiple agent frameworks. Through comprehensive comparisons across both models and frameworks, we show how base model capabilities and agent framework designs jointly shape performance in real-world environments. To facilitate future research, we make our benchmark and code publicly available at https://github.com/HKU-MMLab/UniClawBench.

Community

Paper submitter about 21 hours ago

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it difficult to identify the root causes of agent failures. To address these limitations, we introduce UniClawBench, the first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings. UniClawBench is built around five foundational model capabilities: Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, and Cross-Platform Coordination. Based on these capabilities, we design 400 bilingual real-world tasks. Unlike previous benchmarks that rely on static, pre-recorded answers, our benchmark evaluates agents in live Docker containers using fine-grained, step-by-step completion checkpoints. Furthermore, we design a closed-loop evaluation strategy comprising an executor agent, a hidden supervisor agent, and a user agent to simulate realistic multi-turn human feedback without leaking grading criteria. To disentangle base model capabilities from framework-level design choices, we evaluate state-of-the-art models under multiple agent frameworks. Through comprehensive comparisons across both models and frameworks, we show how base model capabilities and agent framework designs jointly shape performance in real-world environments. To facilitate future research, we make our benchmark and code publicly available at https://github.com/HKU-MMLab/UniClawBench

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.08768
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2607.08768 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.08768 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2607.08768 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers