LLaDA-Image is a competitive 6B-parameter open-source unified image generation and editing model family. It includes LLaDA-Image, a 50-step Base model for high-quality text-to-image generation and instruction-guided editing, and LLaDA-Image-Turbo, a 4-step distilled model for fast generation and editing. Both variants support practical text-to-image generation, VQ-conditioned generation, reference-image editing, and Chinese--English text rendering.</p>\n","updatedAt":"2026-09-04T01:47:21.373Z","author":{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","fullname":"Haoxing chen","name":"HaoxingChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8637649416923523},"editors":["HaoxingChen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg"],"reactions":[],"isReport":false}},{"id":"6a9a71b48a8188b33ab65144","author":{"_id":"6a3ef2989c8c0cbd8c05a214","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/y666uRzeluX9pn3cyrq27.png","fullname":"TJ June","name":"tjjune123","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-04T07:22:28.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is insane \nSome just dropped a podcast discussion about this paper on https://tensorbrife.site/\nAnd there are also several podcast on other papers\nTo get direct link for this paper discussion access Listen to \"A Powerful, Fully Open-Source Image Generation and Editing Model Family\" on TensorBrief https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b","html":"<p>This is insane<br>Some just dropped a podcast discussion about this paper on <a href=\"https://tensorbrife.site/\" rel=\"nofollow\">https://tensorbrife.site/</a><br>And there are also several podcast on other papers<br>To get direct link for this paper discussion access Listen to \"A Powerful, Fully Open-Source Image Generation and Editing Model Family\" on TensorBrief <a href=\"https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b\" rel=\"nofollow\">https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b</a></p>\n","updatedAt":"2026-09-04T07:22:28.262Z","author":{"_id":"6a3ef2989c8c0cbd8c05a214","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/y666uRzeluX9pn3cyrq27.png","fullname":"TJ June","name":"tjjune123","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8213039040565491},"editors":["tjjune123"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/y666uRzeluX9pn3cyrq27.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03796","authors":[{"_id":"6a9a20ae8f7c3b75572393dc","name":"Chuyan Chen","hidden":false},{"_id":"6a9a20ae8f7c3b75572393dd","user":{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","isPro":false,"fullname":"Haoxing chen","user":"HaoxingChen","type":"user","name":"HaoxingChen"},"name":"Haoxing Chen","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.215Z","hidden":false},{"_id":"6a9a20ae8f7c3b75572393de","name":"Kun Chen","hidden":false},{"_id":"6a9a20ae8f7c3b75572393df","name":"Zhenglin Cheng","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e0","name":"Long Cui","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e1","name":"Ruishan Fang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e2","user":{"_id":"60d2a2984956988b63753371","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d2a2984956988b63753371/apXIcWbi7jnLVH37CdMTV.jpeg","isPro":false,"fullname":"Zhangxuan Gu","user":"zhangxgu","type":"user","name":"zhangxgu"},"name":"Zhangxuan Gu","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.226Z","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e3","user":{"_id":"6321dc2615b7beab57c3afa1","avatarUrl":"/avatars/68b440d813c6d6252aa8afe25d2aba47.svg","isPro":false,"fullname":"Zhicheng Huang","user":"David-huang","type":"user","name":"David-huang"},"name":"Zhicheng Huang","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.222Z","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e4","name":"Zhenzhong Lan","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e5","name":"Yuanting Lei","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e6","name":"Haoquan Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e7","name":"Jianguo Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e8","name":"Rongchuan Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e9","name":"Sidu Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ea","name":"Tao Lin","hidden":false},{"_id":"6a9a20ae8f7c3b75572393eb","name":"Deyuan Liu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ec","name":"Jiacheng Liu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ed","name":"Lin Liu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ee","name":"Yuxuan Lou","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ef","name":"Zhisheng Lu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f0","name":"Yuxin Ma","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f1","name":"Shuheng Shen","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f2","name":"Peng Sun","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f3","name":"Chaoyang Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f4","name":"Hongjun Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f5","name":"Xiaomei Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f6","name":"Yongxin Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f7","name":"Chengzhang Wu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f8","name":"Hongru Wu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f9","name":"Jun Xie","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes","submittedOnDailyBy":{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","isPro":false,"fullname":"Haoxing chen","user":"HaoxingChen","type":"user","name":"HaoxingChen"},"summary":"We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.","upvotes":82,"discussionId":"6a9a20af8f7c3b75572393fa","projectPage":"https://github.com/inclusionAI/LLaDA-Image","ai_summary":"LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art open-source results.","ai_keywords":["Diffusion Transformer","DiT","vision-language understanding","LLaDA2.0-Mini","diffusion language model","image-only pre-training","RMSNorm","Muon optimizer","parameter-free","LLaDA-Image-Turbo","distillation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"67aea5c8f086ab0f70ed97c9","name":"inclusionAI","fullname":"inclusionAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/fyKuazRifqiaIO34xrhhm.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","isPro":false,"fullname":"Haoxing chen","user":"HaoxingChen","type":"user"},{"_id":"664e0c3042c1e249238facc7","avatarUrl":"/avatars/9efb7b1d89996b9d6af0c8f288c9c900.svg","isPro":false,"fullname":"Stormy (SII)","user":"StormyX","type":"user"},{"_id":"65028e8389707f182386588c","avatarUrl":"/avatars/86a748a3264e6e0f4ee5eaf8f7032ecb.svg","isPro":true,"fullname":"oneko","user":"kenshinn","type":"user"},{"_id":"6a6a81ef4c287dbb8e7ea111","avatarUrl":"/avatars/b360809862f7e66d9662afd87c2e488a.svg","isPro":false,"fullname":"Patricia Smith","user":"patricia-smith","type":"user"},{"_id":"6a6a8282c125cc860a923126","avatarUrl":"/avatars/54c8a9aa6343e151f0877178d6ce4fe2.svg","isPro":false,"fullname":"Michael Davis","user":"michael-davis","type":"user"},{"_id":"6a6a91d04076c9a06636ccec","avatarUrl":"/avatars/ba9022fbb629ae3e6b71a4f7479f28b8.svg","isPro":false,"fullname":"Thomas Perez","user":"thomas-perez","type":"user"},{"_id":"6a6a92d228e0925e730258ad","avatarUrl":"/avatars/e1857df7419f398b90ea1cb3e28cdb55.svg","isPro":false,"fullname":"Thomas Sanchez","user":"thomas-sanchez","type":"user"},{"_id":"6a6a9417a569279e390e1343","avatarUrl":"/avatars/4866974d3fab4ba081b094cd0819de75.svg","isPro":false,"fullname":"Mark Davis","user":"silverArcL","type":"user"},{"_id":"6a6a9b15ba280f1955230f14","avatarUrl":"/avatars/a5b5a109e7bbccac09e88a81872662aa.svg","isPro":false,"fullname":"Edward Williams","user":"ZenithMap","type":"user"},{"_id":"6a6aa0a8b172d8c070b6b478","avatarUrl":"/avatars/a3e77445bdf99a9d00175ac942c99e81.svg","isPro":false,"fullname":"Edward Wilson","user":"Indigo-Flow","type":"user"},{"_id":"6a6a9d781fcc333c083e0a2f","avatarUrl":"/avatars/31b1c26caa8af0dd423cc1ff6ed9af43.svg","isPro":false,"fullname":"Brian Jones","user":"brian-jones","type":"user"},{"_id":"6a6aa1c6557526be22f38a09","avatarUrl":"/avatars/8677177172aaef08cbc5b35680dce73a.svg","isPro":false,"fullname":"Steven Hernandez","user":"orbitglade","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"67aea5c8f086ab0f70ed97c9","name":"inclusionAI","fullname":"inclusionAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/fyKuazRifqiaIO34xrhhm.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03796.md","query":{}}">
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
Abstract
LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art open-source results.
We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.
Community
LLaDA-Image is a competitive 6B-parameter open-source unified image generation and editing model family. It includes LLaDA-Image, a 50-step Base model for high-quality text-to-image generation and instruction-guided editing, and LLaDA-Image-Turbo, a 4-step distilled model for fast generation and editing. Both variants support practical text-to-image generation, VQ-conditioned generation, reference-image editing, and Chinese--English text rendering.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.03796 in a dataset README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.