Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.</p>\n","updatedAt":"2026-08-31T05:30:14.850Z","author":{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","fullname":"chendy25","name":"DongyangChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8366963863372803},"editors":["DongyangChen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.25417","authors":[{"_id":"6a8fc4613bd48bb654ea689a","name":"Shudong Liu","hidden":false},{"_id":"6a8fc4613bd48bb654ea689b","user":{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","isPro":false,"fullname":"chendy25","user":"DongyangChen","type":"user","name":"DongyangChen"},"name":"Dongyang Chen","status":"claimed_verified","statusLastChangedAt":"2026-08-28T13:21:44.021Z","hidden":false},{"_id":"6a8fc4613bd48bb654ea689c","name":"Enci Zhang","hidden":false},{"_id":"6a8fc4613bd48bb654ea689d","name":"Jinwei Liang","hidden":false},{"_id":"6a8fc4613bd48bb654ea689e","name":"Zheng Ma","hidden":false},{"_id":"6a8fc4613bd48bb654ea689f","name":"Lewei Lu","hidden":false}],"publishedAt":"2026-08-26T00:00:00.000Z","submittedOnDailyAt":"2026-08-31T00:00:00.000Z","title":"Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents","submittedOnDailyBy":{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","isPro":false,"fullname":"chendy25","user":"DongyangChen","type":"user","name":"DongyangChen"},"summary":"Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.","upvotes":12,"discussionId":"6a8fc4613bd48bb654ea68a0","projectPage":"https://huggingface.co/EASEL-Bench","githubRepo":"https://github.com/OOOHS/EASEL","githubRepoAddedBy":"user","ai_summary":"EASEL benchmarks fine-grained visual tool use through reference-guided reconstruction and semantic tasks, revealing that current multimodal agents struggle with closed-loop precision and trajectory stability.","ai_keywords":["dexterous visual tool use","reference-guided visual reconstruction","closed-loop parameterized visual action","multimodal agents","trajectory supervision","EASEL-Data","EASEL-9B"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":0,"organization":{"_id":"6a8c13bf21ce9c13bb18af82","name":"EASEL-Bench","fullname":"EASEL-Bench","avatar":"https://www.gravatar.com/avatar/9ba401f6bca264382c5eb05a1b8ea467?d=retro&size=100"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"68bdc19f50490484f15e3a97","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/mtRipayDWN1hEL1Yx5xWg.png","isPro":false,"fullname":"chendy25","user":"DongyangChen","type":"user"},{"_id":"6913fc439e7cd70fcf343d2a","avatarUrl":"/avatars/13d9e345f8f22ed748b6e41e9c2cca53.svg","isPro":false,"fullname":"Shudong Liu","user":"sh0ooo","type":"user"},{"_id":"6a85c1ed85b5753bc01b6b82","avatarUrl":"/avatars/2648377c44b5fa7ed4a1510d213a5343.svg","isPro":false,"fullname":"Rongwei Wang","user":"Vigor020509","type":"user"},{"_id":"645755a7182c64e98984a59d","avatarUrl":"/avatars/7dd613e31f8a948a81b1b5cfa20ea620.svg","isPro":false,"fullname":"liuhyer","user":"yerr2","type":"user"},{"_id":"66a117a3b824c037736ce39a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/cTxQus1FBVbpX8_48iqGv.jpeg","isPro":false,"fullname":"苏德昭","user":"ssdddzzzz","type":"user"},{"_id":"68f9d57f60cfbcef7539e31b","avatarUrl":"/avatars/fce994d1f45d1fc0ea66f94f316bf461.svg","isPro":false,"fullname":"Ken Tanaka","user":"BruceWayneJP","type":"user"},{"_id":"6a41aae5c65231a89caf7c0f","avatarUrl":"/avatars/64635d50f3512161bdbc2739f5148a6b.svg","isPro":false,"fullname":"Shubo Li","user":"lishubo","type":"user"},{"_id":"66ef7fdaad0dc3f0eb03d413","avatarUrl":"/avatars/63dd709f86a8d5b80915a797b651501d.svg","isPro":false,"fullname":"Chaoyang Wang","user":"ChaoyangWang","type":"user"},{"_id":"65ee88ab6b0697f087dca7a9","avatarUrl":"/avatars/4adca10037ed5ec161c0d50929ed7cbf.svg","isPro":false,"fullname":"jimmy liang","user":"jimmy77","type":"user"},{"_id":"68ac1c1a6e734f57c81b44cb","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/FfhSc_mF3KxSm_ehNrLqq.png","isPro":false,"fullname":"pangcong","user":"ppcc1135","type":"user"},{"_id":"65717368be66cd9b65a8201c","avatarUrl":"/avatars/fe945828eec9ded4cfa3b89d48a64d90.svg","isPro":false,"fullname":"Wu Zehuan","user":"wzhgba","type":"user"},{"_id":"6603b5eb85170ce50847e066","avatarUrl":"/avatars/1b57ec39df511e536d747741dd970696.svg","isPro":false,"fullname":"Zhouhanyu Shen","user":"Shenzhou02","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a8c13bf21ce9c13bb18af82","name":"EASEL-Bench","fullname":"EASEL-Bench","avatar":"https://www.gravatar.com/avatar/9ba401f6bca264382c5eb05a1b8ea467?d=retro&size=100"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.25417.md","query":{}}">
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
Abstract
EASEL benchmarks fine-grained visual tool use through reference-guided reconstruction and semantic tasks, revealing that current multimodal agents struggle with closed-loop precision and trajectory stability.
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.
Community
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.25417 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.25417 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.25417 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.