NoRA: Normalized LoRA Method Speeds Convergence and Cuts Catastrophic Forgetting at Zero Extra Cost
Summary
A new arXiv paper (2608.31036) introduces Normalized Low-Rank Adaptation (NoRA), which stabilizes LoRA training by normalizing down-projection matrices. Authors report faster convergence, better training stability, and reduced catastrophic forgetting across pretraining, SFT, and RL — with no extra trainable parameters and no inference-time overhead.
Originally reported by huggingface.co
Read the original article →Original headline: NoRA: Normalized LoRA Method Speeds Convergence and Cuts Catastrophic Forgetting at Zero Extra Cost