arXiv cs.LG
7/24/2026

The Active Ingredient in Muon's Grokking
Short summary
This paper isolates the mechanism behind the Muon optimizer's faster grokking on modular arithmetic, showing that orthogonalization (Newton-Schulz iteration) is the active ingredient while spectral-norm constraints alone provide no speedup over AdamW. Orthogonalizing optimizers reach generalization at ~3x lower spectral norm and settle into lower-norm solutions. Reducing Newton-Schulz iterations from five to one accelerates threshold crossing but makes the grokked solution fragile, while five iterations remain robust across learning rates. The authors release full training and analysis code.
- •Orthogonalization via Newton-Schulz iteration is the key mechanism; spectral-norm constraints alone match AdamW speed
- •Orthogonalizing optimizers achieve generalization at ~3x lower spectral norm than alternatives
- •Reducing iterations from 5 to 1 speeds up grokking but causes transient collapse at higher learning rates
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
