Hugging Face Daily Papers · · 5 min read

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., \"subject composition\"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (⊕), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (⊕) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/kClIRhwUtCL4ww5OmhHQc.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/kClIRhwUtCL4ww5OmhHQc.png\" alt=\"Clipboard_Screenshot_1787039539\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/J-6YdI7bXKrVHQo2CaMGT.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/J-6YdI7bXKrVHQo2CaMGT.png\" alt=\"Clipboard_Screenshot_1787039572\"></a></p>\n","updatedAt":"2026-08-18T07:53:20.272Z","author":{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","fullname":"chaofanma","name":"chaofanma","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8343113660812378},"editors":["chaofanma"],"editorAvatarUrls":["/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16765","authors":[{"_id":"6a83f656675db694db8cd628","name":"Haoran Wang","hidden":false},{"_id":"6a83f656675db694db8cd629","name":"Chaofan Ma","hidden":false},{"_id":"6a83f656675db694db8cd62a","name":"Ran Yi","hidden":false},{"_id":"6a83f656675db694db8cd62b","name":"Lizhuang Ma","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation","submittedOnDailyBy":{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","isPro":false,"fullname":"chaofanma","user":"chaofanma","type":"user","name":"chaofanma"},"summary":"Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., \"subject composition\"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench","upvotes":7,"discussionId":"6a83f657675db694db8cd62c","projectPage":"https://amuseum-whr.github.io/TraceBench/","ai_summary":"This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations.","ai_keywords":["multi-reference image generation","compositional formula","Anchor","Disentangle","Apply","Compose","TRACE-Bench","operator-aligned evaluation","diagnostic tree analysis","disentanglement","attribute binding"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","isPro":false,"fullname":"chaofanma","user":"chaofanma","type":"user"},{"_id":"64ca30b1e3cc4a476d40138f","avatarUrl":"/avatars/bf7f5347279744260f25f3c553e0de97.svg","isPro":false,"fullname":"王浩然","user":"AMuseum","type":"user"},{"_id":"66c4b72abaf18141ad340727","avatarUrl":"/avatars/74ad246ac09037424ec6d39b3a88f71a.svg","isPro":false,"fullname":"wenhao","user":"bluixe","type":"user"},{"_id":"6376eeb5fe88a92b3a8cde4d","avatarUrl":"/avatars/706065aae14113c33de673a72bd009d2.svg","isPro":false,"fullname":"Wei Jiang","user":"Ailon-Island","type":"user"},{"_id":"677c155cc8551c58d2673522","avatarUrl":"/avatars/025d10d80d222535798ba88e4c161434.svg","isPro":false,"fullname":"Haoyu Zhen","user":"haoyuzhen","type":"user"},{"_id":"6437c7dae282b4a48eaf065e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6437c7dae282b4a48eaf065e/AxodKQXyrviTFQRyjnL01.jpeg","isPro":true,"fullname":"Haoyu Zhen","user":"anyeZHY","type":"user"},{"_id":"6847a76a223d5a02bf0e8ae7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/lFUjbQkJ1_wBSQVnen-gA.png","isPro":false,"fullname":"Zhaoyu Zeng","user":"zengzhaoyu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16765.md","query":{}}">
Papers
arxiv:2608.16765

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

Published on Aug 17
· Submitted by
chaofanma
on Aug 18
Authors:
,

Abstract

This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations.

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench

Community

Paper submitter about 1 hour ago

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (⊕), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (⊕) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity.

Clipboard_Screenshot_1787039539

Clipboard_Screenshot_1787039572

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16765
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.16765 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.16765 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.16765 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers