Hugging Face Daily Papers · · 4 min read

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

LLaDA-Image is a competitive 6B-parameter open-source unified image generation and editing model family. It includes LLaDA-Image, a 50-step Base model for high-quality text-to-image generation and instruction-guided editing, and LLaDA-Image-Turbo, a 4-step distilled model for fast generation and editing. Both variants support practical text-to-image generation, VQ-conditioned generation, reference-image editing, and Chinese--English text rendering.</p>\n","updatedAt":"2026-09-04T01:47:21.373Z","author":{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","fullname":"Haoxing chen","name":"HaoxingChen","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":5,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8637649416923523},"editors":["HaoxingChen"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg"],"reactions":[],"isReport":false}},{"id":"6a9a71b48a8188b33ab65144","author":{"_id":"6a3ef2989c8c0cbd8c05a214","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/y666uRzeluX9pn3cyrq27.png","fullname":"TJ June","name":"tjjune123","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-09-04T07:22:28.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is insane \nSome just dropped a podcast discussion about this paper on https://tensorbrife.site/\nAnd there are also several podcast on other papers\nTo get direct link for this paper discussion access Listen to \"A Powerful, Fully Open-Source Image Generation and Editing Model Family\" on TensorBrief https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b","html":"<p>This is insane<br>Some just dropped a podcast discussion about this paper on <a href=\"https://tensorbrife.site/\" rel=\"nofollow\">https://tensorbrife.site/</a><br>And there are also several podcast on other papers<br>To get direct link for this paper discussion access Listen to \"A Powerful, Fully Open-Source Image Generation and Editing Model Family\" on TensorBrief <a href=\"https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b\" rel=\"nofollow\">https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b</a></p>\n","updatedAt":"2026-09-04T07:22:28.262Z","author":{"_id":"6a3ef2989c8c0cbd8c05a214","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/y666uRzeluX9pn3cyrq27.png","fullname":"TJ June","name":"tjjune123","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8213039040565491},"editors":["tjjune123"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/y666uRzeluX9pn3cyrq27.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.03796","authors":[{"_id":"6a9a20ae8f7c3b75572393dc","name":"Chuyan Chen","hidden":false},{"_id":"6a9a20ae8f7c3b75572393dd","user":{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","isPro":false,"fullname":"Haoxing chen","user":"HaoxingChen","type":"user","name":"HaoxingChen"},"name":"Haoxing Chen","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.215Z","hidden":false},{"_id":"6a9a20ae8f7c3b75572393de","name":"Kun Chen","hidden":false},{"_id":"6a9a20ae8f7c3b75572393df","name":"Zhenglin Cheng","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e0","name":"Long Cui","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e1","name":"Ruishan Fang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e2","user":{"_id":"60d2a2984956988b63753371","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/60d2a2984956988b63753371/apXIcWbi7jnLVH37CdMTV.jpeg","isPro":false,"fullname":"Zhangxuan Gu","user":"zhangxgu","type":"user","name":"zhangxgu"},"name":"Zhangxuan Gu","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.226Z","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e3","user":{"_id":"6321dc2615b7beab57c3afa1","avatarUrl":"/avatars/68b440d813c6d6252aa8afe25d2aba47.svg","isPro":false,"fullname":"Zhicheng Huang","user":"David-huang","type":"user","name":"David-huang"},"name":"Zhicheng Huang","status":"claimed_verified","statusLastChangedAt":"2026-09-04T08:45:04.222Z","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e4","name":"Zhenzhong Lan","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e5","name":"Yuanting Lei","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e6","name":"Haoquan Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e7","name":"Jianguo Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e8","name":"Rongchuan Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393e9","name":"Sidu Li","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ea","name":"Tao Lin","hidden":false},{"_id":"6a9a20ae8f7c3b75572393eb","name":"Deyuan Liu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ec","name":"Jiacheng Liu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ed","name":"Lin Liu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ee","name":"Yuxuan Lou","hidden":false},{"_id":"6a9a20ae8f7c3b75572393ef","name":"Zhisheng Lu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f0","name":"Yuxin Ma","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f1","name":"Shuheng Shen","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f2","name":"Peng Sun","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f3","name":"Chaoyang Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f4","name":"Hongjun Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f5","name":"Xiaomei Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f6","name":"Yongxin Wang","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f7","name":"Chengzhang Wu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f8","name":"Hongru Wu","hidden":false},{"_id":"6a9a20ae8f7c3b75572393f9","name":"Jun Xie","hidden":false}],"publishedAt":"2026-09-03T00:00:00.000Z","submittedOnDailyAt":"2026-09-04T00:00:00.000Z","title":"LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes","submittedOnDailyBy":{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","isPro":false,"fullname":"Haoxing chen","user":"HaoxingChen","type":"user","name":"HaoxingChen"},"summary":"We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.","upvotes":82,"discussionId":"6a9a20af8f7c3b75572393fa","projectPage":"https://github.com/inclusionAI/LLaDA-Image","ai_summary":"LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art open-source results.","ai_keywords":["Diffusion Transformer","DiT","vision-language understanding","LLaDA2.0-Mini","diffusion language model","image-only pre-training","RMSNorm","Muon optimizer","parameter-free","LLaDA-Image-Turbo","distillation"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"67aea5c8f086ab0f70ed97c9","name":"inclusionAI","fullname":"inclusionAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/fyKuazRifqiaIO34xrhhm.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63898c562a897944ea5f07a2","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63898c562a897944ea5f07a2/oG1ia44ICpAk1tReIaX0G.jpeg","isPro":false,"fullname":"Haoxing chen","user":"HaoxingChen","type":"user"},{"_id":"664e0c3042c1e249238facc7","avatarUrl":"/avatars/9efb7b1d89996b9d6af0c8f288c9c900.svg","isPro":false,"fullname":"Stormy (SII)","user":"StormyX","type":"user"},{"_id":"65028e8389707f182386588c","avatarUrl":"/avatars/86a748a3264e6e0f4ee5eaf8f7032ecb.svg","isPro":true,"fullname":"oneko","user":"kenshinn","type":"user"},{"_id":"6a6a81ef4c287dbb8e7ea111","avatarUrl":"/avatars/b360809862f7e66d9662afd87c2e488a.svg","isPro":false,"fullname":"Patricia Smith","user":"patricia-smith","type":"user"},{"_id":"6a6a8282c125cc860a923126","avatarUrl":"/avatars/54c8a9aa6343e151f0877178d6ce4fe2.svg","isPro":false,"fullname":"Michael Davis","user":"michael-davis","type":"user"},{"_id":"6a6a91d04076c9a06636ccec","avatarUrl":"/avatars/ba9022fbb629ae3e6b71a4f7479f28b8.svg","isPro":false,"fullname":"Thomas Perez","user":"thomas-perez","type":"user"},{"_id":"6a6a92d228e0925e730258ad","avatarUrl":"/avatars/e1857df7419f398b90ea1cb3e28cdb55.svg","isPro":false,"fullname":"Thomas Sanchez","user":"thomas-sanchez","type":"user"},{"_id":"6a6a9417a569279e390e1343","avatarUrl":"/avatars/4866974d3fab4ba081b094cd0819de75.svg","isPro":false,"fullname":"Mark Davis","user":"silverArcL","type":"user"},{"_id":"6a6a9b15ba280f1955230f14","avatarUrl":"/avatars/a5b5a109e7bbccac09e88a81872662aa.svg","isPro":false,"fullname":"Edward Williams","user":"ZenithMap","type":"user"},{"_id":"6a6aa0a8b172d8c070b6b478","avatarUrl":"/avatars/a3e77445bdf99a9d00175ac942c99e81.svg","isPro":false,"fullname":"Edward Wilson","user":"Indigo-Flow","type":"user"},{"_id":"6a6a9d781fcc333c083e0a2f","avatarUrl":"/avatars/31b1c26caa8af0dd423cc1ff6ed9af43.svg","isPro":false,"fullname":"Brian Jones","user":"brian-jones","type":"user"},{"_id":"6a6aa1c6557526be22f38a09","avatarUrl":"/avatars/8677177172aaef08cbc5b35680dce73a.svg","isPro":false,"fullname":"Steven Hernandez","user":"orbitglade","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"organization":{"_id":"67aea5c8f086ab0f70ed97c9","name":"inclusionAI","fullname":"inclusionAI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/662e1f9da266499277937d33/fyKuazRifqiaIO34xrhhm.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.03796.md","query":{}}">
Papers
arxiv:2609.03796

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

Published on Sep 3
· Submitted by
Haoxing chen
on Sep 4
#1 Paper of the day

Abstract

LLaDA-Image unifies a 6B diffusion transformer with a frozen vision-language module, using image-only pre-training and a Muon optimizer to generate photorealistic images with precise editing, and is distilled into a fast 2-4 step variant that achieves state-of-the-art open-source results.

We introduce LLaDA-Image, a unified framework that pairs a 6B Diffusion Transformer (DiT) trained from scratch with a frozen vision-language understanding module built on the LLaDA2.0-Mini diffusion language model backbone. Instead of relying heavily on paired image-text data from the beginning, we first build a strong visual generative prior through image-only pre-training and mid-training. The generation pipeline comprises 220M samples, 98 of which are real images. For efficient and scalable optimization, we use parameter-free RMSNorm throughout the DiT together with the Muon optimizer. The resulting unified model produces highly photorealistic images while accurately following fine-grained editing instructions. We further distill LLaDA-Image into LLaDA-Image-Turbo, enabling fast inference in 2-4 sampling steps. On Qwen-Image-Bench, LLaDA-Image achieves overall scores of 53.53 and 53.38 on the English and Chinese tracks, respectively, setting a new state-of-the-art among open-source models on both tracks. To support further research on capable and efficient generative models, we release our model weights, training code, and detailed recipes.

Community

Paper author Paper submitter about 7 hours ago

LLaDA-Image is a competitive 6B-parameter open-source unified image generation and editing model family. It includes LLaDA-Image, a 50-step Base model for high-quality text-to-image generation and instruction-guided editing, and LLaDA-Image-Turbo, a 4-step distilled model for fast generation and editing. Both variants support practical text-to-image generation, VQ-conditioned generation, reference-image editing, and Chinese--English text rendering.

This is insane
Some just dropped a podcast discussion about this paper on https://tensorbrife.site/
And there are also several podcast on other papers
To get direct link for this paper discussion access Listen to "A Powerful, Fully Open-Source Image Generation and Editing Model Family" on TensorBrief https://www.tensorbrife.site/podcast/c01b7c2d-9e9b-436c-a6f9-aaeca35c797b

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.03796
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2609.03796 in a dataset README.md to link it from this page.

Spaces citing this paper

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers