Miles is open-sourced at <a href=\"https://github.com/radixark/miles\" rel=\"nofollow\">https://github.com/radixark/miles</a>, with the project website at <a href=\"https://miles.radixark.com/\" rel=\"nofollow\">https://miles.radixark.com/</a></p>\n","updatedAt":"2026-09-09T03:52:36.344Z","author":{"_id":"66038f8e90c62bb38f06eb9f","avatarUrl":"/avatars/e67137b47f45f580bae7ea6472671359.svg","fullname":"Zhichen Zeng","name":"CharyZeng","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.845340371131897},"editors":["CharyZeng"],"editorAvatarUrls":["/avatars/e67137b47f45f580bae7ea6472671359.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2609.08368","authors":[{"_id":"6aa0d54bd0174964227bed36","name":"RadixArk","hidden":false},{"_id":"6aa0d54bd0174964227bed38","name":"Tom Chen","hidden":false},{"_id":"6aa0d54bd0174964227bed39","name":"Mao Cheng","hidden":false},{"_id":"6aa0d54bd0174964227bed3a","name":"Shi Dong","hidden":false},{"_id":"6aa0d54bd0174964227bed3b","name":"Kangrui Du","hidden":false},{"_id":"6aa0d54bd0174964227bed3c","name":"Yanbin Jiang","hidden":false},{"_id":"6aa0d54bd0174964227bed3d","name":"Jiajun Li","hidden":false},{"_id":"6aa0d54bd0174964227bed3e","name":"Yiming Li","hidden":false},{"_id":"6aa0d54bd0174964227bed3f","user":{"_id":"6aa0d94c7f8bc24c79f0d4ec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6aa0d94c7f8bc24c79f0d4ec/PPZnlhQrdTlfLNOKTuywA.jpeg","isPro":false,"fullname":"Tao Lin","user":"nblintao","type":"user","name":"nblintao"},"name":"Tao Lin","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.659Z","hidden":false},{"_id":"6aa0d54bd0174964227bed40","name":"Yusheng Su","hidden":false},{"_id":"6aa0d54bd0174964227bed41","name":"Andy Ye","hidden":false},{"_id":"6aa0d54bd0174964227bed42","name":"Yueming Yuan","hidden":false},{"_id":"6aa0d54bd0174964227bed43","user":{"_id":"66038f8e90c62bb38f06eb9f","avatarUrl":"/avatars/e67137b47f45f580bae7ea6472671359.svg","isPro":false,"fullname":"Zhichen Zeng","user":"CharyZeng","type":"user","name":"CharyZeng"},"name":"Zhichen Zeng","status":"claimed_verified","statusLastChangedAt":"2026-09-09T08:45:04.651Z","hidden":false}],"publishedAt":"2026-09-08T00:00:00.000Z","submittedOnDailyAt":"2026-09-09T00:00:00.000Z","title":"Miles v0.1: Production-Level Post-Training","submittedOnDailyBy":{"_id":"66038f8e90c62bb38f06eb9f","avatarUrl":"/avatars/e67137b47f45f580bae7ea6472671359.svg","isPro":false,"fullname":"Zhichen Zeng","user":"CharyZeng","type":"user","name":"CharyZeng"},"summary":"We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.","upvotes":41,"discussionId":"6aa0d54bd0174964227bed44","projectPage":"https://miles.radixark.com/","githubRepo":"https://github.com/radixark/miles","githubRepoAddedBy":"user","ai_summary":"Miles is an open-source, production-ready system for large-scale reinforcement learning and post-training that supports diverse backends, weight synchronization, LoRA, distillation, and diffusion models.","ai_keywords":["reinforcement-learning","rollout engines","SGLang","Megatron-LM","PyTorch FSDP","weight-synchronization","LoRA RL","on-policy distillation","supervised fine-tuning","true-on-policy rollout-training alignment","diffusion models","agentic RL"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2681,"organization":{"_id":"692f7057b350eac7e5075674","name":"RadixArk","fullname":"RadixArk","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/692f6fa22fe03c133986059c/uVcDantaDXJOwc_T93yei.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"66038f8e90c62bb38f06eb9f","avatarUrl":"/avatars/e67137b47f45f580bae7ea6472671359.svg","isPro":false,"fullname":"Zhichen Zeng","user":"CharyZeng","type":"user"},{"_id":"60f1abe7544c2adfd699860c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674929746905-60f1abe7544c2adfd699860c.jpeg","isPro":false,"fullname":"AK","user":"akhaliq","type":"user"},{"_id":"644bf6ef778ecbfb977e8e84","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644bf6ef778ecbfb977e8e84/x2zFOkSIhDHf0LEZ1wzbU.jpeg","isPro":false,"fullname":"Adarsh AS","user":"adarshxs","type":"user"},{"_id":"6aa0d94c7f8bc24c79f0d4ec","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6aa0d94c7f8bc24c79f0d4ec/PPZnlhQrdTlfLNOKTuywA.jpeg","isPro":false,"fullname":"Tao Lin","user":"nblintao","type":"user"},{"_id":"6838d22a62d46c89d3aef8fd","avatarUrl":"/avatars/c00cd39954f0714d3cf2e759dd39c8cc.svg","isPro":false,"fullname":"Shi","user":"sdong15","type":"user"},{"_id":"6a46b51fa28951ed1e12ea86","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/JVXuBeU32ISbkQeG9v_gS.jpeg","isPro":false,"fullname":"Haoguang Cai","user":"unseenmars","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"64b805561d913123e6ed3335","avatarUrl":"/avatars/9a394aee4398cbcac2013288f1f85ade.svg","isPro":false,"fullname":"Yanbin Jiang","user":"jybsuper","type":"user"},{"_id":"666e81b1067382b3e917892f","avatarUrl":"/avatars/b0a6e7219da61d5754dd6038e492b88a.svg","isPro":false,"fullname":"Haolin Fu","user":"hlFu","type":"user"},{"_id":"6836b9390ad5a1a7f3595b53","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/99d1Bps3eC0Rx681zm9Lv.png","isPro":false,"fullname":"yang yuhao","user":"yhyang201","type":"user"},{"_id":"65f758eebc6b396606fe5b67","avatarUrl":"/avatars/e16ca327bdc9832f39644bee7d821d25.svg","isPro":false,"fullname":"Lingyan Hao","user":"LyH88","type":"user"},{"_id":"6a6d0ef3dd326b5ce47f4f4d","avatarUrl":"/avatars/f66b8f8c26a6c6782656892fedfada99.svg","isPro":false,"fullname":"A S","user":"a-shao","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"692f7057b350eac7e5075674","name":"RadixArk","fullname":"RadixArk","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/692f6fa22fe03c133986059c/uVcDantaDXJOwc_T93yei.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2609/2609.08368.md","query":{}}">
Miles v0.1: Production-Level Post-Training
Abstract
Miles is an open-source, production-ready system for large-scale reinforcement learning and post-training that supports diverse backends, weight synchronization, LoRA, distillation, and diffusion models.
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2609.08368 in a model README.md to link it from this page.
Cite arxiv.org/abs/2609.08368 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2609.08368 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.