to be better</p>\n","updatedAt":"2026-09-01T04:31:50.452Z","author":{"_id":"64b496c9bcfd8542d63802ff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b496c9bcfd8542d63802ff/r4cSujcWckHz0eRKZ4NWH.jpeg","fullname":"J.L","name":"Jo1uck","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":6,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9550177454948425},"editors":["Jo1uck"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/64b496c9bcfd8542d63802ff/r4cSujcWckHz0eRKZ4NWH.jpeg"],"reactions":[],"isReport":false}},{"id":"6a965b7d304c39bfc7dbdef0","author":{"_id":"60c82ac0dfbd57a384a01127","avatarUrl":"/avatars/5422e1be5f9eb06ac396fcd2430641c5.svg","fullname":"Ji Seunghyun","name":"sorryhyun","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":12,"isUserFollowing":false},"createdAt":"2026-09-01T04:58:37.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"I think this paper has in common with idea in https://arxiv.org/abs/2510.01938","html":"<p>I think this paper has in common with idea in <a href=\"https://arxiv.org/abs/2510.01938\" rel=\"nofollow\">https://arxiv.org/abs/2510.01938</a></p>\n","updatedAt":"2026-09-01T04:58:37.249Z","author":{"_id":"60c82ac0dfbd57a384a01127","avatarUrl":"/avatars/5422e1be5f9eb06ac396fcd2430641c5.svg","fullname":"Ji Seunghyun","name":"sorryhyun","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":12,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9812636971473694},"editors":["sorryhyun"],"editorAvatarUrls":["/avatars/5422e1be5f9eb06ac396fcd2430641c5.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.31036","authors":[{"_id":"6a96533dcd6ebc484732edcc","name":"Jiale Kang","hidden":false},{"_id":"6a96533dcd6ebc484732edcd","name":"Ziyin Yue","hidden":false},{"_id":"6a96533dcd6ebc484732edce","name":"Zheng Zhan","hidden":false},{"_id":"6a96533dcd6ebc484732edcf","name":"Yangyi Huang","hidden":false},{"_id":"6a96533dcd6ebc484732edd0","name":"Weiyang Liu","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-01T00:00:00.000Z","title":"Normalized Low-Rank Adaptation","submittedOnDailyBy":{"_id":"64b496c9bcfd8542d63802ff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b496c9bcfd8542d63802ff/r4cSujcWckHz0eRKZ4NWH.jpeg","isPro":false,"fullname":"J.L","user":"Jo1uck","type":"user","name":"Jo1uck"},"summary":"While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. Building on this observation, we introduce Normalized Low-Rank Adaptation (NoRA), a simple yet effective method that normalizes the down-projection matrices during training. We further show that the same normalization can be applied only at initialization, improving standard LoRA without requiring repeated normalization throughout training. Across pretraining, supervised finetuning, and reinforcement learning, NoRA consistently accelerates convergence, improves performance and training stability, and mitigates catastrophic forgetting. These benefits require neither additional trainable parameters nor inference-time computation, making NoRA a simple and broadly applicable enhancement to LoRA.","upvotes":23,"discussionId":"6a96533dcd6ebc484732edd1","githubRepo":"https://github.com/Joluck/NoRA","githubRepoAddedBy":"user","ai_summary":"Normalized Low-Rank Adaptation stabilizes LoRA training by normalizing down-projection matrices, accelerating convergence and improving performance without extra parameters or inference cost.","ai_keywords":["low-rank adaptation","LoRA","Normalized Low-Rank Adaptation","NoRA","down-projection","catastrophic forgetting","reinforcement learning"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64b496c9bcfd8542d63802ff","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64b496c9bcfd8542d63802ff/r4cSujcWckHz0eRKZ4NWH.jpeg","isPro":false,"fullname":"J.L","user":"Jo1uck","type":"user"},{"_id":"64bfa52094c0e3be4a33863e","avatarUrl":"/avatars/bc0737bf4481433b669433facf6ce6f4.svg","isPro":false,"fullname":"Shenzhennan","user":"5456es","type":"user"},{"_id":"655646baf8a2d3c020546ec8","avatarUrl":"/avatars/4ca8de82745bb5a4fda511569bb6bd94.svg","isPro":false,"fullname":"Niclas P","user":"NPBP26","type":"user"},{"_id":"64c4accbb496b4e176786171","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/64c4accbb496b4e176786171/eyh6ulc4fufpGmf_DO2K3.jpeg","isPro":false,"fullname":"Molly Sophia","user":"mollysama","type":"user"},{"_id":"64fe4fd79132c7f62a56efb7","avatarUrl":"/avatars/222487e18bc7e8300310f0a759fa4e27.svg","isPro":false,"fullname":"zheng zhan","user":"zz8585","type":"user"},{"_id":"66de61d7174e9c6971dbb253","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/sM0xfS7HAkf_6GmkEGjDk.png","isPro":false,"fullname":"Alic Li","user":"Alic-Li","type":"user"},{"_id":"642d6902179328845319fc77","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642d6902179328845319fc77/WoIplMsTyKEenwxchDiQL.jpeg","isPro":false,"fullname":"cgisky","user":"cgisky","type":"user"},{"_id":"642c2398eb6e214d4f8962c6","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/642c2398eb6e214d4f8962c6/26cc7i_jtx7X9ugSgbEyd.png","isPro":false,"fullname":"Beortust","user":"Beortust","type":"user"},{"_id":"65372bae0d973d3fee4131c7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/OeBjr1hlUINmjUJ8iu42I.png","isPro":false,"fullname":"BadCat","user":"Foresta","type":"user"},{"_id":"67a63df4463947b63eacc9c8","avatarUrl":"/avatars/7b5097444185ee46dbf464a9aca95f81.svg","isPro":false,"fullname":"ricky","user":"rickywangricky","type":"user"},{"_id":"6821898e0097a361d0b79abc","avatarUrl":"/avatars/3f43f9276394fd4ede9801ea12aeabf7.svg","isPro":false,"fullname":"manjuan","user":"Ehoon","type":"user"},{"_id":"671e61c1a95cf4695495654d","avatarUrl":"/avatars/c83f205a2d8b2a8a64dafb856a78529a.svg","isPro":false,"fullname":"lin","user":"leo42580","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"query":{}}">
Normalized Low-Rank Adaptation
Published on Aug 31
· Submitted by J.L on Sep 1 Abstract
Normalized Low-Rank Adaptation stabilizes LoRA training by normalizing down-projection matrices, accelerating convergence and improving performance without extra parameters or inference cost.
While low-rank adaptation (LoRA) is widely used for parameter-efficient model adaptation, how to regularize its training dynamics for stable and effective optimization remains underexplored. Because LoRA initializes the up-projection to zero, its early optimization dynamics are largely governed by the down-projection. Building on this observation, we introduce Normalized Low-Rank Adaptation (NoRA), a simple yet effective method that normalizes the down-projection matrices during training. We further show that the same normalization can be applied only at initialization, improving standard LoRA without requiring repeated normalization throughout training. Across pretraining, supervised finetuning, and reinforcement learning, NoRA consistently accelerates convergence, improves performance and training stability, and mitigates catastrophic forgetting. These benefits require neither additional trainable parameters nor inference-time computation, making NoRA a simple and broadly applicable enhancement to LoRA.
Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.31036 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.31036 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.31036 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.