Hugging Face Daily Papers · · 6 min read

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.</p>\n","updatedAt":"2026-08-31T05:30:14.850Z","author":{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","fullname":"chendy25","name":"DongyangChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8366963863372803},"editors":["DongyangChen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.25417","authors":[{"_id":"6a8fc4613bd48bb654ea689a","name":"Shudong Liu","hidden":false},{"_id":"6a8fc4613bd48bb654ea689b","user":{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","isPro":false,"fullname":"chendy25","user":"DongyangChen","type":"user","name":"DongyangChen"},"name":"Dongyang Chen","status":"claimed_verified","statusLastChangedAt":"2026-08-28T13:21:44.021Z","hidden":false},{"_id":"6a8fc4613bd48bb654ea689c","name":"Enci Zhang","hidden":false},{"_id":"6a8fc4613bd48bb654ea689d","name":"Jinwei Liang","hidden":false},{"_id":"6a8fc4613bd48bb654ea689e","name":"Zheng Ma","hidden":false},{"_id":"6a8fc4613bd48bb654ea689f","name":"Lewei Lu","hidden":false}],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-31T00:00:00.000Z","title":"Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents","submittedOnDailyBy":{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","isPro":false,"fullname":"chendy25","user":"DongyangChen","type":"user","name":"DongyangChen"},"summary":"Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.","upvotes":12,"discussionId":"6a8fc4613bd48bb654ea68a0","projectPage":"https://huggingface.co/EASEL-Bench","githubRepo":"https://github.com/OOOHS/EASEL","githubRepoAddedBy":"user","ai_summary":"EASEL benchmarks fine-grained visual tool use through reference-guided reconstruction and semantic tasks, revealing that current multimodal agents struggle with closed-loop precision and trajectory stability.","ai_keywords":["dexterous visual tool use","reference-guided visual reconstruction","closed-loop parameterized visual action","multimodal agents","trajectory supervision","EASEL-Data","EASEL-9B"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6a8c13bf21ce9c13bb18af82","name":"EASEL-Bench","fullname":"EASEL-Bench","avatar":"https://www.gravatar.com/avatar/9ba401f6bca264382c5eb05a1b8ea467?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","isPro":false,"fullname":"chendy25","user":"DongyangChen","type":"user"},{"_id":"6913fc439e7cd70fcf343d2a","avatarUrl":"/avatars/13d9e345f8f22ed748b6e41e9c2cca53.svg","isPro":false,"fullname":"Shudong Liu","user":"sh0ooo","type":"user"},{"_id":"6a85c1ed85b5753bc01b6b82","avatarUrl":"/avatars/2648377c44b5fa7ed4a1510d213a5343.svg","isPro":false,"fullname":"Rongwei Wang","user":"Vigor020509","type":"user"},{"_id":"645755a7182c64e98984a59d","avatarUrl":"/avatars/7dd613e31f8a948a81b1b5cfa20ea620.svg","isPro":false,"fullname":"liuhyer","user":"yerr2","type":"user"},{"_id":"66a117a3b824c037736ce39a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/cTxQus1FBVbpX8_48iqGv.jpeg","isPro":false,"fullname":"苏德昭","user":"ssdddzzzz","type":"user"},{"_id":"68f9d57f60cfbcef7539e31b","avatarUrl":"/avatars/fce994d1f45d1fc0ea66f94f316bf461.svg","isPro":false,"fullname":"Ken Tanaka","user":"BruceWayneJP","type":"user"},{"_id":"6a41aae5c65231a89caf7c0f","avatarUrl":"/avatars/64635d50f3512161bdbc2739f5148a6b.svg","isPro":false,"fullname":"Shubo Li","user":"lishubo","type":"user"},{"_id":"66ef7fdaad0dc3f0eb03d413","avatarUrl":"/avatars/63dd709f86a8d5b80915a797b651501d.svg","isPro":false,"fullname":"Chaoyang Wang","user":"ChaoyangWang","type":"user"},{"_id":"65ee88ab6b0697f087dca7a9","avatarUrl":"/avatars/4adca10037ed5ec161c0d50929ed7cbf.svg","isPro":false,"fullname":"jimmy liang","user":"jimmy77","type":"user"},{"_id":"68ac1c1a6e734f57c81b44cb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FfhSc_mF3KxSm_ehNrLqq.png","isPro":false,"fullname":"pangcong","user":"ppcc1135","type":"user"},{"_id":"65717368be66cd9b65a8201c","avatarUrl":"/avatars/fe945828eec9ded4cfa3b89d48a64d90.svg","isPro":false,"fullname":"Wu Zehuan","user":"wzhgba","type":"user"},{"_id":"6603b5eb85170ce50847e066","avatarUrl":"/avatars/1b57ec39df511e536d747741dd970696.svg","isPro":false,"fullname":"Zhouhanyu Shen","user":"Shenzhou02","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a8c13bf21ce9c13bb18af82","name":"EASEL-Bench","fullname":"EASEL-Bench","avatar":"https://www.gravatar.com/avatar/9ba401f6bca264382c5eb05a1b8ea467?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.25417.md","query":{}}">
Papers
arxiv:2608.25417

Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

Published on Aug 26
· Submitted by
chendy25
on Aug 31
Authors:
,

Abstract

EASEL benchmarks fine-grained visual tool use through reference-guided reconstruction and semantic tasks, revealing that current multimodal agents struggle with closed-loop precision and trajectory stability.

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.

Community

Paper author Paper submitter about 3 hours ago

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.25417
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.25417 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.25417 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.25417 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers