Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Understanding the Role of Momentum in Stochastic Gradient Methods

Oct 30, 2019

Igor Gitman, Hunter Lang, Pengchuan Zhang, Lin Xiao

Figure 1 for Understanding the Role of Momentum in Stochastic Gradient Methods

Figure 2 for Understanding the Role of Momentum in Stochastic Gradient Methods

Figure 3 for Understanding the Role of Momentum in Stochastic Gradient Methods

Figure 4 for Understanding the Role of Momentum in Stochastic Gradient Methods

Share this with someone who'll enjoy it:

Abstract:The use of momentum in stochastic gradient methods has become a widespread practice in machine learning. Different variants of momentum, including heavy-ball momentum, Nesterov's accelerated gradient (NAG), and quasi-hyperbolic momentum (QHM), have demonstrated success on various tasks. Despite these empirical successes, there is a lack of clear understanding of how the momentum parameters affect convergence and various performance measures of different algorithms. In this paper, we use the general formulation of QHM to give a unified analysis of several popular algorithms, covering their asymptotic convergence conditions, stability regions, and properties of their stationary distributions. In addition, by combining the results on convergence rates and stationary distributions, we obtain sometimes counter-intuitive practical guidelines for setting the learning rate and momentum parameters.

* 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada

View paper on

Share this with someone who'll enjoy it:

Title:Understanding the Role of Momentum in Stochastic Gradient Methods

Paper and Code