Hugging Face Daily Papers · · 3 min read

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

Boogu-Image-0.1 is a strongly competitive Apache-2.0 open-source unified image generation and editing model family, including Base, Turbo, Edit, and other variants that provide stable, practical capabilities for high-quality text-to-image generation, fast generation, image editing, and Chinese-English text rendering, with performance that matches top closed-source models in many scenarios.</p>\n","updatedAt":"2026-07-16T04:31:20.959Z","author":{"_id":"63a2ca27dca41424cb8387a3","avatarUrl":"/avatars/ae9bec9c2326184e94faa8d6fd98f2da.svg","fullname":"X","name":"Chufeng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8987017273902893},"editors":["Chufeng"],"editorAvatarUrls":["/avatars/ae9bec9c2326184e94faa8d6fd98f2da.svg"],"reactions":[{"reaction":"🔥","users":["EurekaY","CostaliyA","cynricfu","ChenyangLei"],"count":4}],"isReport":false}},{"id":"6a587398a48c80795ece21d0","author":{"_id":"63423d3e9948f573f378df4b","avatarUrl":"/avatars/d1e28f4d675bf0eeaf98c3686ad1b45d.svg","fullname":"Perhacept","name":"Perhacept","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-07-16T06:00:56.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"impressive!","html":"<p>impressive!</p>\n","updatedAt":"2026-07-16T06:00:56.730Z","author":{"_id":"63423d3e9948f573f378df4b","avatarUrl":"/avatars/d1e28f4d675bf0eeaf98c3686ad1b45d.svg","fullname":"Perhacept","name":"Perhacept","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7108615636825562},"editors":["Perhacept"],"editorAvatarUrls":["/avatars/d1e28f4d675bf0eeaf98c3686ad1b45d.svg"],"reactions":[{"reaction":"🚀","users":["EurekaY","ChenyangLei"],"count":2}],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.13125","authors":[{"_id":"6a585e6f69aa0f8878bddbe2","name":"Guoxuan Chen","hidden":false},{"_id":"6a585e6f69aa0f8878bddbe3","name":"Chufeng Xiao","hidden":false},{"_id":"6a585e6f69aa0f8878bddbe4","name":"Haoran Yang","hidden":false},{"_id":"6a585e6f69aa0f8878bddbe5","name":"Siyue Xie","hidden":false},{"_id":"6a585e6f69aa0f8878bddbe6","name":"Binxiao Huang","hidden":false},{"_id":"6a585e6f69aa0f8878bddbe7","name":"Ming Zhang","hidden":false},{"_id":"6a585e6f69aa0f8878bddbe8","name":"Cheuk Him Chau","hidden":false},{"_id":"6a585e6f69aa0f8878bddbe9","name":"Xinyu Fu","hidden":false},{"_id":"6a585e6f69aa0f8878bddbea","name":"Yingzhao Lian","hidden":false},{"_id":"6a585e6f69aa0f8878bddbeb","name":"Tom S. Y. Li","hidden":false},{"_id":"6a585e6f69aa0f8878bddbec","name":"Jintao Lin","hidden":false},{"_id":"6a585e6f69aa0f8878bddbed","name":"Bowen Dong","hidden":false},{"_id":"6a585e6f69aa0f8878bddbee","name":"Zian Qian","hidden":false},{"_id":"6a585e6f69aa0f8878bddbef","name":"Yuhao Liu","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf0","name":"Yuxuan Hu","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf1","name":"Weikang Shi","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf2","name":"Bin Zou","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf3","name":"Bowen Zheng","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf4","name":"Haoxuan Che","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf5","name":"Chang Chen","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf6","name":"Yuyang He","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf7","name":"Heyang Sun","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf8","name":"Tianyu Huang","hidden":false},{"_id":"6a585e6f69aa0f8878bddbf9","name":"Chong Hou Choi","hidden":false},{"_id":"6a585e6f69aa0f8878bddbfa","name":"Cheng Gong","hidden":false},{"_id":"6a585e6f69aa0f8878bddbfb","name":"Han Shi","hidden":false},{"_id":"6a585e6f69aa0f8878bddbfc","name":"Haoli Bai","hidden":false},{"_id":"6a585e6f69aa0f8878bddbfd","name":"Xihui Liu","hidden":false},{"_id":"6a585e6f69aa0f8878bddbfe","name":"Hongsheng Li","hidden":false},{"_id":"6a585e6f69aa0f8878bddbff","name":"Qifeng Chen","hidden":false},{"_id":"6a585e6f69aa0f8878bddc00","name":"Chao Huang","hidden":false},{"_id":"6a585e6f69aa0f8878bddc01","name":"Rui Liu","hidden":false},{"_id":"6a585e6f69aa0f8878bddc02","name":"Chenyang Lei","hidden":false}],"publishedAt":"2026-07-14T00:00:00.000Z","submittedOnDailyAt":"2026-07-16T00:00:00.000Z","title":"Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation","submittedOnDailyBy":{"_id":"63a2ca27dca41424cb8387a3","avatarUrl":"/avatars/ae9bec9c2326184e94faa8d6fd98f2da.svg","isPro":false,"fullname":"X","user":"Chufeng","type":"user","name":"Chufeng"},"summary":"We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \\$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.","upvotes":107,"discussionId":"6a585e7069aa0f8878bddc03","projectPage":"https://boogu.org/","githubRepo":"https://github.com/boogu-project/Boogu-Image","githubRepoAddedBy":"user","githubStars":755,"organization":{"_id":"6a2f735add4271ad7948febc","name":"Boogu","fullname":"Boogu","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a2f728163c271161df8f4a7/gRDLtMmEnc3HIVldi5HjH.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6687f9a71309e08b1f84bdc6","avatarUrl":"/avatars/f947ec9fe620ae4cffa83b371acdd571.svg","isPro":false,"fullname":"MeiYi","user":"natalie5","type":"user"},{"_id":"633c4c33d5935998f751081c","avatarUrl":"/avatars/cb093ebc750e274ba21828838732a430.svg","isPro":true,"fullname":"Lei","user":"ChenyangLei","type":"user"},{"_id":"656326571338610184c16448","avatarUrl":"/avatars/e5fb4fdf86706363ad931962b34925b3.svg","isPro":false,"fullname":"cera","user":"cccera","type":"user"},{"_id":"63a2ca27dca41424cb8387a3","avatarUrl":"/avatars/ae9bec9c2326184e94faa8d6fd98f2da.svg","isPro":false,"fullname":"X","user":"Chufeng","type":"user"},{"_id":"656844522d73834278a17db9","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/656844522d73834278a17db9/TA-moNnbtUqY52cwPcjUO.jpeg","isPro":false,"fullname":"Elian","user":"same899","type":"user"},{"_id":"66d59dc9b005ad82ca6fc61d","avatarUrl":"/avatars/0ba424690afd1144a89665c5bacdfde7.svg","isPro":false,"fullname":"Runyi YU","user":"IngridYU","type":"user"},{"_id":"6451fa0dc5fc2fed144a97ef","avatarUrl":"/avatars/d0a9a39f4cf4ec392600c58393d169fb.svg","isPro":true,"fullname":"Xinyu Fu","user":"cynricfu","type":"user"},{"_id":"69cb24029f710900fbf3a706","avatarUrl":"/avatars/d4976f3983ddc322412fbd191de30343.svg","isPro":false,"fullname":"lian","user":"yzleo","type":"user"},{"_id":"653e646589f7466f2cb851cf","avatarUrl":"/avatars/be15092a86f70e0b0671e7d02c3f8105.svg","isPro":false,"fullname":"Yunkang Tao","user":"LuffyGear5","type":"user"},{"_id":"64b751baa00eab5bcde04edb","avatarUrl":"/avatars/7c422f736ffcc219c88d57062e52ad81.svg","isPro":false,"fullname":"Yiyang Li","user":"EricLee6936","type":"user"},{"_id":"6450d2ce302fecf037af302e","avatarUrl":"/avatars/e493ebd06a7f70d459bb847db518b3db.svg","isPro":false,"fullname":"He","user":"OctRev","type":"user"},{"_id":"63423d3e9948f573f378df4b","avatarUrl":"/avatars/d1e28f4d675bf0eeaf98c3686ad1b45d.svg","isPro":false,"fullname":"Perhacept","user":"Perhacept","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":2,"organization":{"_id":"6a2f735add4271ad7948febc","name":"Boogu","fullname":"Boogu","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a2f728163c271161df8f4a7/gRDLtMmEnc3HIVldi5HjH.png"},"query":{}}">
Papers
arxiv:2607.13125

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

Published on Jul 14
· Submitted by
X
on Jul 16
#2 Paper of the day
Authors:
,

Abstract

We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.

Community

Paper submitter about 16 hours ago

Boogu-Image-0.1 is a strongly competitive Apache-2.0 open-source unified image generation and editing model family, including Base, Turbo, Edit, and other variants that provide stable, practical capabilities for high-quality text-to-image generation, fast generation, image editing, and Chinese-English text rendering, with performance that matches top closed-source models in many scenarios.

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Models citing this paper

Browse 10 models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2607.13125 in a dataset README.md to link it from this page.

Spaces citing this paper

Browse 17 spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers