Hugging Face Daily Papers · · 3 min read

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

paper: <a href=\"https://arxiv.org/abs/2608.03457\" rel=\"nofollow\">https://arxiv.org/abs/2608.03457</a></p>\n","updatedAt":"2026-08-05T06:47:12.083Z","author":{"_id":"624f909eac5dd186b01ac3f5","avatarUrl":"/avatars/0aafdb1cbb492fda52a0303031cc6c14.svg","fullname":"Zebin You","name":"yyyou","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":4,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.49811187386512756},"editors":["yyyou"],"editorAvatarUrls":["/avatars/0aafdb1cbb492fda52a0303031cc6c14.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.03457","authors":[{"_id":"6a72cc621a375f948521c52a","name":"Fengqi Zhu","hidden":false},{"_id":"6a72cc621a375f948521c52b","name":"Shaoxuan Xu","hidden":false},{"_id":"6a72cc621a375f948521c52c","name":"Jingyang Ou","hidden":false},{"_id":"6a72cc621a375f948521c52d","name":"Zebin You","hidden":false},{"_id":"6a72cc621a375f948521c52e","name":"Yipeng Xing","hidden":false},{"_id":"6a72cc621a375f948521c52f","name":"Huabin Liu","hidden":false},{"_id":"6a72cc621a375f948521c530","name":"Xiaolu Zhang","hidden":false},{"_id":"6a72cc621a375f948521c531","name":"Jun Zhou","hidden":false},{"_id":"6a72cc621a375f948521c532","name":"Zhenzhong Lan","hidden":false},{"_id":"6a72cc621a375f948521c533","name":"Yankai Lin","hidden":false},{"_id":"6a72cc621a375f948521c534","name":"Wayne Xin Zhao","hidden":false},{"_id":"6a72cc621a375f948521c535","name":"Jianguo Li","hidden":false},{"_id":"6a72cc621a375f948521c536","name":"Chongxuan Li","hidden":false},{"_id":"6a72cc621a375f948521c537","name":"Ji-Rong Wen","hidden":false}],"publishedAt":"2026-08-04T00:00:00.000Z","submittedOnDailyAt":"2026-08-05T00:00:00.000Z","title":"LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models","submittedOnDailyBy":{"_id":"624f909eac5dd186b01ac3f5","avatarUrl":"/avatars/0aafdb1cbb492fda52a0303031cc6c14.svg","isPro":false,"fullname":"Zebin You","user":"yyyou","type":"user","name":"yyyou"},"summary":"Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.","upvotes":12,"discussionId":"6a72cc621a375f948521c538","organization":{"_id":"67b58a7e7aaae267db2fa046","name":"GSAI-ML","fullname":"GSAI-ML","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/624f909eac5dd186b01ac3f5/6tfvx3XT5Sx6YDGl7KUAU.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"65faa2f9d5a985179868dcf7","avatarUrl":"/avatars/16f9f236ce98fcd57f99a56a1b1a5915.svg","isPro":false,"fullname":"Monohydroxides","user":"Monohydroxides","type":"user"},{"_id":"656c3e7c9c8778992f9cf83e","avatarUrl":"/avatars/4a60d75663877a238cb6abcafccdc3cc.svg","isPro":false,"fullname":"Ziheng Peng","user":"SKwra","type":"user"},{"_id":"6a073249dbb75fbd0f1dc28a","avatarUrl":"/avatars/859cbf1a275c06911e9a60d8d376eb4c.svg","isPro":false,"fullname":"Rongzhen Wang","user":"rongzhenwang","type":"user"},{"_id":"682e8e6d007cd8c2f2cd0afd","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/az0B5USOPlLR3vq872JeN.png","isPro":false,"fullname":"Chenyu Zheng","user":"ChenyuZheng","type":"user"},{"_id":"66e01987c52ff2985f2dcf76","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/uDEaluRKI7N2oYepbz1s3.jpeg","isPro":false,"fullname":"ShaoXuan Xu","user":"WriteMath","type":"user"},{"_id":"67be0888267276f5a2f9ce71","avatarUrl":"/avatars/c27dc0003351d2e59d7361064d572e55.svg","isPro":false,"fullname":"Jia-Nan Li","user":"JinaLeejnl","type":"user"},{"_id":"624f909eac5dd186b01ac3f5","avatarUrl":"/avatars/0aafdb1cbb492fda52a0303031cc6c14.svg","isPro":false,"fullname":"Zebin You","user":"yyyou","type":"user"},{"_id":"66bb04f6447411b9c0125570","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/66bb04f6447411b9c0125570/AS3V1Gk7Ed408yS4Ekuap.jpeg","isPro":false,"fullname":"Yu-Yang Qian","user":"d3LLM-model","type":"user"},{"_id":"652246392d3ba46ccd0a6042","avatarUrl":"/avatars/c7353f27a0767847c15720c1adcf8f20.svg","isPro":false,"fullname":"Lingjie Chen","user":"OnAnOrange","type":"user"},{"_id":"6554681f849e7971cb69eb8d","avatarUrl":"/avatars/dc8c7289c887e9adca6082c36823f94c.svg","isPro":false,"fullname":"Yuxin Li","user":"AndyLYX","type":"user"},{"_id":"66b1dce83158b1fadffb0d9d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/RbWKIxCjawVaSidwZLZA1.png","isPro":false,"fullname":"Kaixu Chen","user":"kaixuChen","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"67b58a7e7aaae267db2fa046","name":"GSAI-ML","fullname":"GSAI-ML","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/624f909eac5dd186b01ac3f5/6tfvx3XT5Sx6YDGl7KUAU.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.03457.md","query":{}}">
Papers
arxiv:2608.03457

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

Published on Aug 4
· Submitted by
Zebin You
on Aug 5
Authors:
,

Abstract

Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dLLMs remains poorly understood. We systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models. Specifically, for optimization, the optimal nominal batch size grows faster, while the optimal learning rate decays more rapidly with compute. For model--data allocation, IsoFLOP analysis reveals a slight data-side tilt: the optimal token budget grows faster than activated model-side computation. For MoE architecture, larger scales increasingly favor larger expert pools at fixed activated capacity, while moderate expert granularity remains consistently effective and the preferred fraction of activated capacity assigned to shared experts remains stable across scales. Guided by these findings, we train LLaDA MoE v2, a 30B-A3B dLLM, from scratch on 23.5T tokens. With approximately 65\% as many pretraining tokens as Qwen3, LLaDA MoE v2 approaches Qwen3 on several knowledge, reasoning, and coding benchmarks. After supervised fine-tuning alone, it outperforms SDAR Chat on seven of eight reasoning and coding benchmarks and remains close to Qwen3 on several tasks. These results establish practical scaling laws and design principles for MoE dLLMs.

Community

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.03457
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.03457 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.03457 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.03457 in a Space README.md to link it from this page.

Collections including this paper

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers