Strong error analysis for the stochastic momentum optimizer
Stochastic gradient descent ({SGD}) optimization schemes are the methods of choice for the optimization of deep neural networks ({DNNs}) in artificial intelligence ({AI}) systems. Often not the standard {SGD} method is used but instead suitable accelerated, adaptive, and/or normalized variants of standard {SGD} such as Adam, {AdamW}, and {MUON} are employed to train large scale {AI} systems in practically relevant settings. The acceleration (higher order convergence speed) in all these popular optimizers relies on the momentum {SGD} optimizer. In this work we provide a rigorous error analysis for the momentum {SGD} optimizer. In particular, we establish convergence rates for the momentum optimizer in terms of the size of the learning rate (step size), the size of the mini-batch, and the size of the one-point convexity constant.
- Published in:
arXiv - Type:
Article - Authors:
- Year:
2026 - Source:
http://arxiv.org/abs/2608.04245
Citation information
: Strong error analysis for the stochastic momentum optimizer, arXiv, 2026, {arXiv}:2608.04245, August, {arXiv}, http://arxiv.org/abs/2608.04245, Gallon.Jentzen.2026a,
@Article{Gallon.Jentzen.2026a,
author={Gallon, Davide; Jentzen, Arnulf},
title={Strong error analysis for the stochastic momentum optimizer},
journal={arXiv},
number={{arXiv}:2608.04245},
month={August},
publisher={{arXiv}},
url={http://arxiv.org/abs/2608.04245},
year={2026},
abstract={Stochastic gradient descent ({SGD}) optimization schemes are the methods of choice for the optimization of deep neural networks ({DNNs}) in artificial intelligence ({AI}) systems. Often not the standard {SGD} method is used but instead suitable accelerated, adaptive, and/or normalized variants of standard {SGD} such as Adam, {AdamW}, and {MUON} are employed to train large scale {AI} systems in...}}