Strong error analysis for the stochastic momentum optimizer

Stochastic gradient descent ({SGD}) optimization schemes are the methods of choice for the optimization of deep neural networks ({DNNs}) in artificial intelligence ({AI}) systems. Often not the standard {SGD} method is used but instead suitable accelerated, adaptive, and/or normalized variants of standard {SGD} such as Adam, {AdamW}, and {MUON} are employed to train large scale {AI} systems in practically relevant settings. The acceleration (higher order convergence speed) in all these popular optimizers relies on the momentum {SGD} optimizer. In this work we provide a rigorous error analysis for the momentum {SGD} optimizer. In particular, we establish convergence rates for the momentum optimizer in terms of the size of the learning rate (step size), the size of the mini-batch, and the size of the one-point convexity constant.

Citation information

Gallon, Davide; Jentzen, Arnulf: Strong error analysis for the stochastic momentum optimizer, arXiv, 2026, {arXiv}:2608.04245, August, {arXiv}, http://arxiv.org/abs/2608.04245, Gallon.Jentzen.2026a,

Associated Lamarr Researchers

Photo. Portrait of Arnulf Jentzen.

Prof. Dr. Arnulf Jentzen

Lamarr Fellow to the profile