Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.</p>\n","updatedAt":"2026-07-23T02:23:11.330Z","author":{"_id":"62e38a3ef3b208e2aecf2c84","avatarUrl":"/avatars/eee9d7d53f3f9b7a18a21d289abd3c64.svg","fullname":"Dongfang Li","name":"crazyofapple","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":1,"identifiedLanguage":{"language":"en","probability":0.18271753191947937},"editors":["crazyofapple"],"editorAvatarUrls":["/avatars/eee9d7d53f3f9b7a18a21d289abd3c64.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2607.20145","authors":[{"_id":"6a617adec3792c34f5a04055","name":"Dongfang Li","hidden":false},{"_id":"6a617adec3792c34f5a04056","name":"Xiaodong Luo","hidden":false},{"_id":"6a617adec3792c34f5a04057","name":"Ruoyu Sun","hidden":false},{"_id":"6a617adec3792c34f5a04058","name":"Xuhui Chen","hidden":false},{"_id":"6a617adec3792c34f5a04059","name":"Linyuan Qiu","hidden":false},{"_id":"6a617adec3792c34f5a0405a","name":"Jian Meng","hidden":false},{"_id":"6a617adec3792c34f5a0405b","name":"Zhengxuan Lu","hidden":false},{"_id":"6a617adec3792c34f5a0405c","name":"Yiting Wang","hidden":false},{"_id":"6a617adec3792c34f5a0405d","name":"Yucheng Xie","hidden":false},{"_id":"6a617adec3792c34f5a0405e","name":"Tao Guo","hidden":false},{"_id":"6a617adec3792c34f5a0405f","name":"Tianxiang Fang","hidden":false},{"_id":"6a617adec3792c34f5a04060","name":"Jing Li","hidden":false},{"_id":"6a617adec3792c34f5a04061","name":"Sihang Chen","hidden":false},{"_id":"6a617adec3792c34f5a04062","name":"Shihao Hong","hidden":false},{"_id":"6a617adec3792c34f5a04063","name":"Chang Liu","hidden":false},{"_id":"6a617adec3792c34f5a04064","name":"Weihua Dai","hidden":false},{"_id":"6a617adec3792c34f5a04065","name":"Zirong Zeng","hidden":false},{"_id":"6a617adec3792c34f5a04066","name":"Ziwei Zhu","hidden":false},{"_id":"6a617adec3792c34f5a04067","name":"Zhuohan Wang","hidden":false},{"_id":"6a617adec3792c34f5a04068","name":"Zhengjun Yue","hidden":false},{"_id":"6a617adec3792c34f5a04069","name":"Igor Vasilyev","hidden":false},{"_id":"6a617adec3792c34f5a0406a","name":"Min Liu","hidden":false},{"_id":"6a617adec3792c34f5a0406b","name":"Weijian Sun","hidden":false},{"_id":"6a617adec3792c34f5a0406c","name":"Xin Chen","hidden":false},{"_id":"6a617adec3792c34f5a0406d","name":"Yingmeng Gao","hidden":false},{"_id":"6a617adec3792c34f5a0406e","name":"Jinhua Zhou","hidden":false},{"_id":"6a617adec3792c34f5a0406f","name":"Taolue Chen","hidden":false},{"_id":"6a617adec3792c34f5a04070","name":"Chenwei Wu","hidden":false},{"_id":"6a617adec3792c34f5a04071","name":"Dong Zhang","hidden":false},{"_id":"6a617adec3792c34f5a04072","name":"Wenlong Jin","hidden":false},{"_id":"6a617adec3792c34f5a04073","name":"Jinmin Xiang","hidden":false},{"_id":"6a617adec3792c34f5a04074","name":"Barkova Maria","hidden":false},{"_id":"6a617adec3792c34f5a04075","name":"Ushakov Anton","hidden":false},{"_id":"6a617adec3792c34f5a04076","name":"Xianfei Jin","hidden":false},{"_id":"6a617adec3792c34f5a04077","name":"Tian Ding","hidden":false},{"_id":"6a617adec3792c34f5a04078","name":"Zhihang Lin","hidden":false},{"_id":"6a617adec3792c34f5a04079","name":"Qian Chen","hidden":false},{"_id":"6a617adec3792c34f5a0407a","name":"Linxin Yang","hidden":false},{"_id":"6a617adec3792c34f5a0407b","name":"Mingzhe Yang","hidden":false},{"_id":"6a617adec3792c34f5a0407c","name":"Bingwei Zhang","hidden":false},{"_id":"6a617adec3792c34f5a0407d","name":"Hongzhang Yang","hidden":false},{"_id":"6a617adec3792c34f5a0407e","name":"Fangxue Zhang","hidden":false},{"_id":"6a617adec3792c34f5a0407f","name":"Shijun Qin","hidden":false},{"_id":"6a617adec3792c34f5a04080","name":"Jie Yu","hidden":false},{"_id":"6a617adec3792c34f5a04081","name":"Cuihua Hu","hidden":false},{"_id":"6a617adec3792c34f5a04082","name":"Tolstykh Vasiliy","hidden":false},{"_id":"6a617adec3792c34f5a04083","name":"Nosov Ivan","hidden":false},{"_id":"6a617adec3792c34f5a04084","name":"Abdullin Amir","hidden":false},{"_id":"6a617adec3792c34f5a04085","name":"Zhichen Zhou","hidden":false},{"_id":"6a617adec3792c34f5a04086","name":"Xin Zhang","hidden":false},{"_id":"6a617adec3792c34f5a04087","name":"Zhixiong Ning","hidden":false},{"_id":"6a617adec3792c34f5a04088","name":"Xutong Zhao","hidden":false},{"_id":"6a617adec3792c34f5a04089","name":"Junjie Huang","hidden":false},{"_id":"6a617adec3792c34f5a0408a","name":"Jiajun Liu","hidden":false},{"_id":"6a617adec3792c34f5a0408b","name":"Weiyan Kong","hidden":false},{"_id":"6a617adec3792c34f5a0408c","name":"Zheng Zhang","hidden":false},{"_id":"6a617adec3792c34f5a0408d","name":"Wenhan Luo","hidden":false},{"_id":"6a617adec3792c34f5a0408e","name":"Lin Hu","hidden":false},{"_id":"6a617adec3792c34f5a0408f","name":"Yangbo Guo","hidden":false},{"_id":"6a617adec3792c34f5a04090","name":"Li Zeng","hidden":false},{"_id":"6a617adec3792c34f5a04091","name":"Shihao Zeng","hidden":false},{"_id":"6a617adec3792c34f5a04092","name":"Baotian Hu","hidden":false},{"_id":"6a617adec3792c34f5a04093","name":"Min Zhang","hidden":false},{"_id":"6a617adec3792c34f5a04094","name":"Haizhou Li","hidden":false},{"_id":"6a617adec3792c34f5a04095","name":"Zhiquan Luo","hidden":false}],"publishedAt":"2026-07-22T00:00:00.000Z","submittedOnDailyAt":"2026-07-23T00:00:00.000Z","title":"SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD","submittedOnDailyBy":{"_id":"62e38a3ef3b208e2aecf2c84","avatarUrl":"/avatars/eee9d7d53f3f9b7a18a21d289abd3c64.svg","isPro":false,"fullname":"Dongfang Li","user":"crazyofapple","type":"user","name":"crazyofapple"},"summary":"Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.","upvotes":28,"discussionId":"6a617adfc3792c34f5a04096","githubRepo":"https://github.com/SLAI-AITP/SLAI-T-Rex","githubRepoAddedBy":"user","githubStars":12},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"62e38a3ef3b208e2aecf2c84","avatarUrl":"/avatars/eee9d7d53f3f9b7a18a21d289abd3c64.svg","isPro":false,"fullname":"Dongfang Li","user":"crazyofapple","type":"user"},{"_id":"63b6dbc8ccebeadccc888456","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1673396893898-63b6dbc8ccebeadccc888456.jpeg","isPro":false,"fullname":"Xin Zhang","user":"izhx","type":"user"},{"_id":"6a4c9bba52fc7b4faa5f9cf6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6a4c9bba52fc7b4faa5f9cf6/6LyZLtM9GsYCL_54S_fiL.png","isPro":false,"fullname":"SLAI-AITP","user":"SLAI-AITP","type":"user"},{"_id":"692d9fcdecb95bcaedb8db50","avatarUrl":"/avatars/a6b2cd6bcce530e07588dadd2c69fb03.svg","isPro":false,"fullname":"Timwell2604","user":"Llover2604","type":"user"},{"_id":"67c5b0a69c3163ab4854c960","avatarUrl":"/avatars/787e7a5aef3c3582aecded1f544240f1.svg","isPro":false,"fullname":"Chen","user":"We11man","type":"user"},{"_id":"67a20dbbfda6bc91694013c6","avatarUrl":"/avatars/41fcf39054aa925979175082d45c534e.svg","isPro":false,"fullname":"Zhengxuan Lu","user":"Toleco","type":"user"},{"_id":"691c51a2ebd2ae5b32008ac7","avatarUrl":"/avatars/61a99082e787243c94fed90796ea334a.svg","isPro":false,"fullname":"Tongshu Bian","user":"BeatlesBIAN","type":"user"},{"_id":"671ceb1dc7ffaca1a3ea3f7e","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/0OTGLfqRPL9TN14nez_DW.png","isPro":false,"fullname":"no1cu pa","user":"No1cu","type":"user"},{"_id":"658c0b0574e79b9a8e9de89a","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/658c0b0574e79b9a8e9de89a/8HeMEOT5cLEzauXGlrqF_.jpeg","isPro":false,"fullname":"Xinping Zhao","user":"Yuki131","type":"user"},{"_id":"6a617e1d3157217e5cdc6ce3","avatarUrl":"/avatars/299843793ecb184e121416dc146c9ac3.svg","isPro":false,"fullname":"Raymond Tse","user":"raymond0612","type":"user"},{"_id":"69b8f36c289868e8a606f45d","avatarUrl":"/avatars/8c461ff0de56f2ef3e8d6506572c31fc.svg","isPro":false,"fullname":"Shibo Su","user":"Su4o","type":"user"},{"_id":"6a617cc180280a26ac9cf943","avatarUrl":"/avatars/cbafc6988e0a20768c9214e0f968bed2.svg","isPro":false,"fullname":"qiulinyuan","user":"qiulinyuan","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":1,"query":{}}">
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Abstract
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
Community
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2607.20145 in a model README.md to link it from this page.
Cite arxiv.org/abs/2607.20145 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2607.20145 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.