Tuesday, May 30, 2023, 3pm
This thesis quantitatively characterizes the optimization dynamics of neural network training, focusing mostly on the simplified scenario where the training algorithm is deterministic. First, we demonstrate that for gradient flow training (i.e. gradient descent in the limit of small learning rates), the curvature of the objective function tends to monotonically increase during training, which can perhaps be interpreted as continual growth in the complexity of the network. Next, we show that for gradient descent with a fixed step size, the curvature increases until the largest Hessian eigenvalue hits the critical threshold 2/(step size), at which point training enters a phase we refer to as "edge of stability" in which: (1) the learning algorithm oscillates in weight space along the largest eigenvalue of the Hessian; (2) the training loss decreases non-monotonically; (3) the largest Hessian eigenvalue equilibrates near the value 2/(step size). In this regime, gradient descent effectively performs constrained optimization - optimizing the training objective while keeping all Hessian eigenvalues approximately beneath the threshold 2/(step size). Subsequent work has explained these dynamics by noting that when gradient descent oscillates in a high-curvature region, the oscillations implicitly perform gradient descent on the curvature itself. Next, we demonstrate that the phenomenon generalizes to the setting of adaptive gradient methods (notably, Adam) as well, though here the curvature regularization has a different form. Finally, for future work, we propose to: (1) quantitatively analyze the dynamics of adaptive gradient methods at the edge of stability, by extending the simplified analysis of Damian et al (2022) from gradient descent to adaptive gradient methods; (2) release a benchmark set of small-scale neural network training problems to facilitate future research in this area.
Thesis Committee:
Zico Kolter (Co-advisor)
Ameet Talwalkar (Co-advisor)
Sebastien Bubeck (Microsoft)
Jason Lee (Princeton University)
Tom Goldstein (University of Maryland)
Additional Information
In Person and Zoom Participation. See announcement.
Event Type: Thesis Proposals
Room Number: In Person and Virtual - ET
Building: Reddy Conference Room, Gates Hillman 4405 and Zoom
Speaker's Name: JEREMY COHEN
Speaker Website: www.cs.cmu.edu…
Speaker's Professional Title: Ph.D. Student, Machine Learning Department, Carnegie Mellon University
Talk Title: The Dynamics of Neural Network Training
For More Information: stidle@andrew.cmu.edu
Affiliations: Machine Learning Department (MLD)
Organization(s): SCS