Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks

Shokhrukh Ibragimov, Arnulf Jentzen

Abstract

Gradient based optimization methods are nowadays the methods of choice for training deep neural networks (DNNs) in artificial intelligence (AI) systems. In practically relevant DNN training problems, one does usually not apply the standard gradient descent (GD) optimization method but instead one employs suitable sophisticated GD optimization methods, which incorporate adaptivity and/or acceleration techniques, such as the famous Adam optimizer. It is a key contribution of this work to provide a general unified convergence analysis for GD optimization methods in the training of DNNs with analytic activations such as the softplus and the popular Gaussian error linear unit (GeLU) activation. Our general unified convergence result applies to a large class of gradient based optimization methods such as the standard GD, the momentum, the Nesterov accelerated gradient (NAG), the RMSprop, the Adam, the Adamax, the Nadam, the Nadamax, the Adan, the AdaBelief, the AMSGrad, and the Yogi optimizers. Our analysis employs the theory of Kurdyka-Łojasiewicz (KL) inequalities to establish convergence to critical points in the training of DNNs. To the best of our knowledge, the generality of our convergence analysis is also just in the special situation of the Adam optimizer a new contribution to the literature on the analysis of AI optimization algorithms.

Disclosure

“90685587, Mathematics Münster: Dynamics-Geometry-Structure funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation). Most of the specific formulations in the proofs of this work have been created using [39]. Use of large language models Google Gemini has significantly supported us in creating the literature review in Subsec- tion 1.3. The specific formulations in Subsection 1.3 are due to the authors and all formula- tions and references in Subsection 1.3 and the entire a”

PDF page 47
Classification
Literature search
Multiplier
2
Verified

Structural counts

Pages 51 pdf
Theorems 0 source
Lemmas 0 source
Propositions 0 source
Corollaries 0 source
Definitions 2 source
Displayed equations 269 source
Bibliography entries 72 source
Appendix pages 0 estimated

Count notes

  • Source counts use the expanded primary TeX file Universal_Optimizer_2026_07_05_arXiv.tex.
  • Appendix pages include the first PDF page with an explicit Appendix heading through the final page.