How Far is Adam from Natural Gradient Descent?
Mirrored from arXiv — Machine Learning for archival readability. Support the source by reading on the original site.
Computer Science > Machine Learning
Title:How Far is Adam from Natural Gradient Descent?
Abstract:Adam is the standard optimizer in deep learning, yet its geometric relationship to natural gradient descent (NGD) contains unresolved questions. We study Adam's full update rule, including momentum, as a diagonal empirical Fisher approximation subject to diagonal truncation, empirical label substitution, and temporal lag. Using the scale-invariant $\gamma(\Delta\theta)$ metric, we measure Adam's geometric deviation from true NGD across four loss landscapes: well-conditioned linear regression, ill-conditioned linear regression, logistic regression, and a non-convex small neural network. Adam's geometric trajectory is context-dependent. Deviation remains low in well-conditioned settings but rises significantly under ill-conditioning, reaching misalignments of $\approx 10^3$ in the neural network. Higher geometric drift correlates with slower initial optimization but does not degrade final objective minimization; Adam consistently reaches low loss. Furthermore, the improved empirical Fisher (iEF) tracks more stable paths than the standard empirical Fisher (EF), which frequently oscillates or diverges. Our results suggest Adam's practical optimization power may stem from a balance of structural approximation errors and momentum smoothing rather than close tracking of the natural gradient path.
| Comments: | 9 pages, 4 figures, 2 tables |
| Subjects: | Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE) |
| MSC classes: | 68T07 |
| ACM classes: | I.2.6 |
| Cite as: | arXiv:2610.00004 [cs.LG] |
| (or arXiv:2610.00004v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00004
arXiv-issued DOI via DataCite
|
|
| Related DOI: | https://doi.org/10.5281/zenodo.20466393
DOI(s) linking to related resources
|
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — Machine Learning
-
Reverse Item Response Theory for Sparsity-Robust Ranking in Fragmented Cancer Drug-Response Matrices
Oct 3
-
FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law
Oct 3
-
Integrating Fairness and Explainability in a Multiple Instance Reinforcement Learning System
Oct 3
-
Fast Polynomial Transcendentals for LLMs
Oct 3
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.