Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire figure, structural fidelity is inherently conjunctive: correctness on one axis cannot compensate for failure on another. Current open-source models fail to satisfy this criterion. Supervised fine-tuning (SFT) learns plausible layouts but cannot reliably ensure structural correctness, while scalar reward-based post-training obscures which structural dimension has failed. To address this, we introduce SciForma, a framework for the structure-faithful generation of scientific methodology diagrams. Specifically, SciForma decomposes diagram quality into three structural axes: Component, Arrow, and Text, guided by a structural inventory. Built on this foundation, we curate SciFormaData-700K for structured training and SciFormaBench-2K for logic-verified evaluation. To close the gap left by SFT, we develop Multi-Dimensional Conjunctive Preference Optimization (M-DPO), which enforces simultaneous correctness across all axes and adaptively routes gradients to the most deficient dimension in post-training. The same structural inventory also enables iterative editing at inference time to correct residual errors. This combination allows SciForma-9B to exceed all open-source baselines and GPT-Image-1.5 on both SciFormaBench-2K and AIBench, bringing open scientific diagram generation close to proprietary-level structural fidelity. Our code and data will be available at: <a href=\"https://github.com/microsoft/SciForma\" rel=\"nofollow\">https://github.com/microsoft/SciForma</a>.</p>\n","updatedAt":"2026-07-22T02:56:23.187Z","author":{"_id":"66e391a5021730e4ead995eb","avatarUrl":"/avatars/43ea3085b77ad9b08d21a8642156e574.svg","fullname":"Luo Yuxuan","name":"LoYuXrqw","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":2,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8406214714050293},"editors":["LoYuXrqw"],"editorAvatarUrls":["/avatars/43ea3085b77ad9b08d21a8642156e574.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.18091","authors":[{"_id":"6a6030b27e7f152167e470bf","user":{"_id":"66e391a5021730e4ead995eb","avatarUrl":"/avatars/43ea3085b77ad9b08d21a8642156e574.svg","isPro":false,"fullname":"Luo Yuxuan","user":"LoYuXrqw","type":"user","name":"LoYuXrqw"},"name":"Yuxuan Luo","status":"claimed_verified","statusLastChangedAt":"2026-07-22T07:39:29.321Z","hidden":false},{"_id":"6a6030b27e7f152167e470c0","name":"Peng Zhang","hidden":false},{"_id":"6a6030b27e7f152167e470c1","name":"Xinjie Zhang","hidden":false},{"_id":"6a6030b27e7f152167e470c2","name":"Xun Guo","hidden":false},{"_id":"6a6030b27e7f152167e470c3","name":"Zhouhui Lian","hidden":false},{"_id":"6a6030b27e7f152167e470c4","name":"Yan Lu","hidden":false}],"publishedAt":"2026-07-20T00:00:00.000Z","submittedOnDailyAt":"2026-07-22T00:00:00.000Z","title":"SciForma: Structure-Faithful Generation of Scientific Diagrams","submittedOnDailyBy":{"_id":"66e391a5021730e4ead995eb","avatarUrl":"/avatars/43ea3085b77ad9b08d21a8642156e574.svg","isPro":false,"fullname":"Luo Yuxuan","user":"LoYuXrqw","type":"user","name":"LoYuXrqw"},"summary":"Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire figure, structural fidelity is inherently conjunctive: correctness on one axis cannot compensate for failure on another. Current open-source models fail to satisfy this criterion. Supervised fine-tuning (SFT) learns plausible layouts but cannot reliably ensure structural correctness, while scalar reward-based post-training obscures which structural dimension has failed. To address this, we introduce SciForma, a framework for the structure faithful generation of scientific methodology diagrams. Specifically, SciForma decomposes diagram quality into three structural axes: Component, Arrow, and Text, guided by a structural inventory. Built on this foundation, we curate SciFormaData-700K for structured training and SciFormaBench-2K for logic-verified evaluation. To close the gap left by SFT, we develop Multi-Dimensional Conjunctive Preference Optimization (M-DPO), which enforces simultaneous correctness across all axes and adaptively routes gradients to the most deficient dimension in post-training. The same structural inventory also enables iterative editing at inference time to correct residual errors. This combination allows SciForma-9B to exceed all open-source baselines and GPT-Image-1.5 on both SciFormaBench-2K and AIBench, bringing open scientific diagram generation close to proprietary-level structural fidelity. Our code and data will be available at: https://github.com/microsoft/SciForma.","upvotes":11,"discussionId":"6a6030b37e7f152167e470c5"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66e391a5021730e4ead995eb","avatarUrl":"/avatars/43ea3085b77ad9b08d21a8642156e574.svg","isPro":false,"fullname":"Luo Yuxuan","user":"LoYuXrqw","type":"user"},{"_id":"64338d1c4521083b9d2d21da","avatarUrl":"/avatars/54b809021d794f1c4b762fbc5d0c7c90.svg","isPro":false,"fullname":"Xinjie","user":"Xinjie-Q","type":"user"},{"_id":"661230e7941ed67394a873e7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/661230e7941ed67394a873e7/MoQ-FHlTu1In-UpG6E9j2.jpeg","isPro":false,"fullname":"Li Yu","user":"phxember","type":"user"},{"_id":"67a2bd1ad1ff276e6c84971b","avatarUrl":"/avatars/6366e5bd558a31c38cb722445d0d457c.svg","isPro":false,"fullname":"wang chuang","user":"learnandturn","type":"user"},{"_id":"66828cc2fcfdfa8c5a509499","avatarUrl":"/avatars/c3421d9c51ca0e91226e3468971197e1.svg","isPro":false,"fullname":"xTryer","user":"xTryer","type":"user"},{"_id":"6666a432b78b8b6b34a816e9","avatarUrl":"/avatars/ee6d98e57658ae82ba7e6fc2592e258b.svg","isPro":false,"fullname":"yuwei yang","user":"Ruler138","type":"user"},{"_id":"649aa367c6cf3cc95bc1b7f6","avatarUrl":"/avatars/4bf5446c261eab08fc06caebf4c5779a.svg","isPro":false,"fullname":"Yifei Shen","user":"yshenaw","type":"user"},{"_id":"65a2a384bfaec7e7cae41d27","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/65a2a384bfaec7e7cae41d27/pXmEXk_7q9-Y-ceLrR9ef.png","isPro":false,"fullname":"Tianci Bi","user":"tiancibi","type":"user"},{"_id":"677921d46f370093aa0d26e4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/_lV-a6412BAICEBjrU1zM.png","isPro":false,"fullname":"LIU Zhening","user":"ZincL","type":"user"},{"_id":"654344009698c3b984cadb99","avatarUrl":"/avatars/8bb3f1a91de4f58a00e219cdcc17f61d.svg","isPro":false,"fullname":"PengZhang","user":"zzzpeng","type":"user"},{"_id":"64b8331277ae61bcc8ee96ac","avatarUrl":"/avatars/93cc624338538d86548b017227cd1841.svg","isPro":false,"fullname":"Nathan Guilhot","user":"nguilhot","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2607/2607.18091.md","query":{}}">
SciForma: Structure-Faithful Generation of Scientific Diagrams
Abstract
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire figure, structural fidelity is inherently conjunctive: correctness on one axis cannot compensate for failure on another. Current open-source models fail to satisfy this criterion. Supervised fine-tuning (SFT) learns plausible layouts but cannot reliably ensure structural correctness, while scalar reward-based post-training obscures which structural dimension has failed. To address this, we introduce SciForma, a framework for the structure faithful generation of scientific methodology diagrams. Specifically, SciForma decomposes diagram quality into three structural axes: Component, Arrow, and Text, guided by a structural inventory. Built on this foundation, we curate SciFormaData-700K for structured training and SciFormaBench-2K for logic-verified evaluation. To close the gap left by SFT, we develop Multi-Dimensional Conjunctive Preference Optimization (M-DPO), which enforces simultaneous correctness across all axes and adaptively routes gradients to the most deficient dimension in post-training. The same structural inventory also enables iterative editing at inference time to correct residual errors. This combination allows SciForma-9B to exceed all open-source baselines and GPT-Image-1.5 on both SciFormaBench-2K and AIBench, bringing open scientific diagram generation close to proprietary-level structural fidelity. Our code and data will be available at: https://github.com/microsoft/SciForma.
Community
Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire figure, structural fidelity is inherently conjunctive: correctness on one axis cannot compensate for failure on another. Current open-source models fail to satisfy this criterion. Supervised fine-tuning (SFT) learns plausible layouts but cannot reliably ensure structural correctness, while scalar reward-based post-training obscures which structural dimension has failed. To address this, we introduce SciForma, a framework for the structure-faithful generation of scientific methodology diagrams. Specifically, SciForma decomposes diagram quality into three structural axes: Component, Arrow, and Text, guided by a structural inventory. Built on this foundation, we curate SciFormaData-700K for structured training and SciFormaBench-2K for logic-verified evaluation. To close the gap left by SFT, we develop Multi-Dimensional Conjunctive Preference Optimization (M-DPO), which enforces simultaneous correctness across all axes and adaptively routes gradients to the most deficient dimension in post-training. The same structural inventory also enables iterative editing at inference time to correct residual errors. This combination allows SciForma-9B to exceed all open-source baselines and GPT-Image-1.5 on both SciFormaBench-2K and AIBench, bringing open scientific diagram generation close to proprietary-level structural fidelity. Our code and data will be available at: https://github.com/microsoft/SciForma.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.18091 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.18091 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.18091 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.