Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., \"subject composition\"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (⊕), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (⊕) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity.</p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/kClIRhwUtCL4ww5OmhHQc.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/kClIRhwUtCL4ww5OmhHQc.png\" alt=\"Clipboard_Screenshot_1787039539\"></a></p>\n<p><a href=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/J-6YdI7bXKrVHQo2CaMGT.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/63048965eb6d777a838cb7a8/J-6YdI7bXKrVHQo2CaMGT.png\" alt=\"Clipboard_Screenshot_1787039572\"></a></p>\n","updatedAt":"2026-08-18T07:53:20.272Z","author":{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","fullname":"chaofanma","name":"chaofanma","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8343113660812378},"editors":["chaofanma"],"editorAvatarUrls":["/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.16765","authors":[{"_id":"6a83f656675db694db8cd628","name":"Haoran Wang","hidden":false},{"_id":"6a83f656675db694db8cd629","name":"Chaofan Ma","hidden":false},{"_id":"6a83f656675db694db8cd62a","name":"Ran Yi","hidden":false},{"_id":"6a83f656675db694db8cd62b","name":"Lizhuang Ma","hidden":false}],"publishedAt":"2026-08-17T00:00:00.000Z","submittedOnDailyAt":"2026-08-18T00:00:00.000Z","title":"TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation","submittedOnDailyBy":{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","isPro":false,"fullname":"chaofanma","user":"chaofanma","type":"user","name":"chaofanma"},"summary":"Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., \"subject composition\"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench","upvotes":7,"discussionId":"6a83f657675db694db8cd62c","projectPage":"https://amuseum-whr.github.io/TraceBench/","ai_summary":"This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations.","ai_keywords":["multi-reference image generation","compositional formula","Anchor","Disentangle","Apply","Compose","TRACE-Bench","operator-aligned evaluation","diagnostic tree analysis","disentanglement","attribute binding"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63048965eb6d777a838cb7a8","avatarUrl":"/avatars/b987fb7f630443bf94a03daf8dcbffe9.svg","isPro":false,"fullname":"chaofanma","user":"chaofanma","type":"user"},{"_id":"64ca30b1e3cc4a476d40138f","avatarUrl":"/avatars/bf7f5347279744260f25f3c553e0de97.svg","isPro":false,"fullname":"王浩然","user":"AMuseum","type":"user"},{"_id":"66c4b72abaf18141ad340727","avatarUrl":"/avatars/74ad246ac09037424ec6d39b3a88f71a.svg","isPro":false,"fullname":"wenhao","user":"bluixe","type":"user"},{"_id":"6376eeb5fe88a92b3a8cde4d","avatarUrl":"/avatars/706065aae14113c33de673a72bd009d2.svg","isPro":false,"fullname":"Wei Jiang","user":"Ailon-Island","type":"user"},{"_id":"677c155cc8551c58d2673522","avatarUrl":"/avatars/025d10d80d222535798ba88e4c161434.svg","isPro":false,"fullname":"Haoyu Zhen","user":"haoyuzhen","type":"user"},{"_id":"6437c7dae282b4a48eaf065e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6437c7dae282b4a48eaf065e/AxodKQXyrviTFQRyjnL01.jpeg","isPro":true,"fullname":"Haoyu Zhen","user":"anyeZHY","type":"user"},{"_id":"6847a76a223d5a02bf0e8ae7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/lFUjbQkJ1_wBSQVnen-gA.png","isPro":false,"fullname":"Zhaoyu Zeng","user":"zengzhaoyu","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"63e5ef7bf2e9a8f22c515654","name":"SJTU","fullname":"Shanghai Jiao Tong University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1676013394657-63e5ee22b6a40bf941da0928.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.16765.md","query":{}}">
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
Abstract
This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations.
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (oplus), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (oplus) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity. Project page: https://amuseum-whr.github.io/TraceBench
Community
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting and lead to fragmented coverage, uncontrolled complexity, and little diagnostic value. Recognizing that diverse multi-reference tasks share a common set of atomic operations, we adopt a capability-oriented perspective and formalize four operators: Anchor (f), Disentangle (g), Apply (⊕), and Compose (C). Any multi-reference prompt can then be represented as a compositional formula over these operators, whose structural complexity is quantified by the number of operator slots. Building on this formulation, we construct TRACE-Bench, comprising approximately 1,600 evaluation cases across slot counts 1--8, built from 631 formula templates and around 4,000 reference images spanning diverse artistic styles and real-world subjects. The formula structure directly drives an operator-aligned evaluation protocol for per-capability scoring and a diagnostic tree analysis for recursive failure localization. Evaluating 9 leading models reveals insights invisible to holistic scoring: the primary bottleneck lies in disentanglement (g) and attribute binding (⊕) rather than scene-level composition (C), with even the best model scoring only 0.74 on attribute fidelity.


Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.16765 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.16765 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.16765 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.