We propose 4 different ways to lower peak memory consumption and improve throughput in large-scale long-context MoE distributed training.</p>\n<ul>\n<li>PipelinedLLEP: Achieve balanced expert parallelism with chunk-wise comm-compute overlap</li>\n<li>Ring-DTP: break-up large dense layer (vocab) across data-tensor-parallel dimensions, 85%+ lower-memory without losing throughput.</li>\n<li>Selective checkpoint offload (SCO) keeps the one long-lived tensor of each checkpoint boundary in CPU memory</li>\n<li>OffloadStreamAdamW turns the serial CPU Adam update of optimizer offload into a bucket pipeline</li>\n</ul>\n<p>All four change only the order and granularity of computation and data movement, so the loss and gradients stay exact. In matched component tests, they cut the MoE dispatch peak by up to 59.3%<br>without losing throughput, the vocabulary projection peak by 86.6%, and the offloaded optimizer<br>step by 2.05× faster. Composed on MoE models from 120B to 667B parameters, they train at<br>1M context length, 8–32× the reach of a tuned FSDP2 baseline, and up to 10.4× its throughput<br><a href=\"https://cdn-uploads.huggingface.co/production/uploads/6393f04df7e70dd0166c004e/lChfmFO9KnwiwWEnVh8bG.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/6393f04df7e70dd0166c004e/lChfmFO9KnwiwWEnVh8bG.png\" alt=\"throughput_and_flops\"></a></p>\n","updatedAt":"2026-09-17T07:03:10.223Z","author":{"_id":"6393f04df7e70dd0166c004e","avatarUrl":"/avatars/acf4e9e0204a7ff7445aecc4102700cd.svg","fullname":"Phi Nguyen","name":"nxphi47","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":11,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8544101119041443},"editors":["nxphi47"],"editorAvatarUrls":["/avatars/acf4e9e0204a7ff7445aecc4102700cd.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.14306","authors":[{"_id":"6aab8f6d1d9cc4dec79625d5","name":"Shrey Pandit","hidden":false},{"_id":"6aab8f6d1d9cc4dec79625d6","name":"Xuan-Phi Nguyen","hidden":false},{"_id":"6aab8f6d1d9cc4dec79625d7","name":"Yiran Zhao","hidden":false},{"_id":"6aab8f6d1d9cc4dec79625d8","name":"Shafiq Joty","hidden":false}],"publishedAt":"2026-09-13T00:00:00.000Z","submittedOnDailyAt":"2026-09-17T00:00:00.000Z","title":"Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training","submittedOnDailyBy":{"_id":"6393f04df7e70dd0166c004e","avatarUrl":"/avatars/acf4e9e0204a7ff7445aecc4102700cd.svg","isPro":false,"fullname":"Phi Nguyen","user":"nxphi47","type":"user","name":"nxphi47"},"summary":"Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint. Four are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count. Which one runs out first changes with the model, the context length, and the device count, so lowering the largest only exposes the next. We bound all four with schedules whose GPU working set is fixed at launch: PipelinedLLEP extends least-loaded expert parallelism with a cap on the tokens each source contributes to a dispatch chunk, Ring-DTP circulates activations or weight shards around a ring at the vocabulary projection and folds each block of logits into an online log-sum-exp, Selective checkpoint offload (SCO) keeps the one long-lived tensor of each checkpoint boundary in CPU memory, and OffloadStreamAdamW turns the serial CPU Adam update of optimizer offload into a bucket pipeline. All four change only the order and granularity of computation and data movement, so the loss and gradients stay exact. In matched component tests, they cut the MoE dispatch peak by up to 59.3% without losing throughput, the vocabulary projection peak by 86.6%, and the offloaded optimizer step by 2.05times faster. Composed on MoE models from 120B to 667B parameters, they train at 1M context length, 8--32times the reach of a tuned FSDP2 baseline, and up to 10.4times its throughput.","upvotes":7,"discussionId":"6aab8f6d1d9cc4dec79625d9","organization":{"_id":"5f6d64475e78cc6b0ed31e4c","name":"Salesforce","fullname":"Salesforce AI Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1602756670970-noauth.jpeg"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6393f04df7e70dd0166c004e","avatarUrl":"/avatars/acf4e9e0204a7ff7445aecc4102700cd.svg","isPro":false,"fullname":"Phi Nguyen","user":"nxphi47","type":"user"},{"_id":"69f93f2be748ce28c92b80e8","avatarUrl":"/avatars/ccedfb62f35b5d08e75294ab0c49b07c.svg","isPro":false,"fullname":"U2-Bench","user":"u2bench-anon","type":"user"},{"_id":"67a79d33a4e7c29abb4bbf3c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/qc3DLYZiGQgD9SxUVzavd.jpeg","isPro":false,"fullname":"Marius Dinca","user":"Puddings22","type":"user"},{"_id":"684d57f26e04c265777ead3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/cuOj-bQqukSZreXgUJlfm.png","isPro":false,"fullname":"Joakim Lee","user":"Reinforcement4All","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"6a2da6c8ca070ee12c6e396c","avatarUrl":"/avatars/0355287dcabaa67dbc7f0b10b87451f9.svg","isPro":false,"fullname":"Joe Mama","user":"JoeMama123123123","type":"user"},{"_id":"6270324ebecab9e2dcf245de","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6270324ebecab9e2dcf245de/cMbtWSasyNlYc9hvsEEzt.jpeg","isPro":false,"fullname":"Kye Gomez","user":"kye","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"5f6d64475e78cc6b0ed31e4c","name":"Salesforce","fullname":"Salesforce AI Research","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1602756670970-noauth.jpeg"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.14306.md","query":{}}">
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
Abstract
Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint. Four are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count. Which one runs out first changes with the model, the context length, and the device count, so lowering the largest only exposes the next. We bound all four with schedules whose GPU working set is fixed at launch: PipelinedLLEP extends least-loaded expert parallelism with a cap on the tokens each source contributes to a dispatch chunk, Ring-DTP circulates activations or weight shards around a ring at the vocabulary projection and folds each block of logits into an online log-sum-exp, Selective checkpoint offload (SCO) keeps the one long-lived tensor of each checkpoint boundary in CPU memory, and OffloadStreamAdamW turns the serial CPU Adam update of optimizer offload into a bucket pipeline. All four change only the order and granularity of computation and data movement, so the loss and gradients stay exact. In matched component tests, they cut the MoE dispatch peak by up to 59.3% without losing throughput, the vocabulary projection peak by 86.6%, and the offloaded optimizer step by 2.05times faster. Composed on MoE models from 120B to 667B parameters, they train at 1M context length, 8--32times the reach of a tuned FSDP2 baseline, and up to 10.4times its throughput.
Community
We propose 4 different ways to lower peak memory consumption and improve throughput in large-scale long-context MoE distributed training.
- PipelinedLLEP: Achieve balanced expert parallelism with chunk-wise comm-compute overlap
- Ring-DTP: break-up large dense layer (vocab) across data-tensor-parallel dimensions, 85%+ lower-memory without losing throughput.
- Selective checkpoint offload (SCO) keeps the one long-lived tensor of each checkpoint boundary in CPU memory
- OffloadStreamAdamW turns the serial CPU Adam update of optimizer offload into a bucket pipeline
All four change only the order and granularity of computation and data movement, so the loss and gradients stay exact. In matched component tests, they cut the MoE dispatch peak by up to 59.3%
without losing throughput, the vocabulary projection peak by 86.6%, and the offloaded optimizer
step by 2.05× faster. Composed on MoE models from 120B to 667B parameters, they train at
1M context length, 8–32× the reach of a tuned FSDP2 baseline, and up to 10.4× its throughput

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.14306 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.14306 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.14306 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.