Tuesday, January 17, 2023, 10:30 – 11:30am

AI is undergoing a paradigm shift with the rise of models pre-trained with self-supervisions and then adapted to a wide range of downstream tasks. However, their working largely remains a mystery—classical learning theory does not apply to situations where training and test tasks are different. This talk will first investigate the role of pre-training losses, showing that contrastive loss extracts meaningful structural information from unlabeled data and the Euclidean distance between embeddings captures the manifold distance between raw datapoints (or, more generally, the graph distance of a so-called positive-pair graph). Moreover, directions in the embedding space correspond to relationships between clusters in the positive-pair graph. Then, I will discuss two other elements necessary for a sharp characterization of the practical pre-trained models: inductive bias of architectures and implicit bias of optimizers. I will introduce two recent projects, where we strengthen the previous theoretical framework by incorporating the inductive bias of architectures and analyze the role of implicit bias of optimizers in pre-training, empirically and theoretically.

Based on arxiv.org/abs/2106.04156, arxiv.org/abs/2204.02683, arxiv.org/abs/2211.14699, and arxiv.org/abs/2210.14199



Tengyu Ma is an assistant professor of Computer Science at Stanford University. He received his Ph.D. from Princeton University and B.E. from Tsinghua University. His research interests include topics in machine learning, algorithms, and their theory, such as deep learning, (deep) reinforcement learning, pre-training / foundation models, robustness, non-convex optimization, distributed optimization, and high-dimensional statistics. He is a recipient of the ACM Doctoral Dissertation Award Honorable Mention, the Sloan Fellowship, the NSF CAREER Award, NeurIPS 2016 best student paper award, and COLT 2018 best paper award.

In Person and Zoom Participation.


Event Type: Seminars
Room Number: In Person and Virtual - ET
Building: Newell-Simon 4305 and Zoom
Speaker's Name: TENGYU MA
Speaker Websiteai.stanford.edu…
Speaker's Professional Title: Assistant Professor, Computer Science Department, Stanford University
Talk Title: Three Facets of Understanding Pre-training: Loss, Inductive Bias of Architectures, and Implicit Bias of Optimizers
For More Informationsharonw@cs.cmu.edu
Affiliations: Computer Science Department (CSD), Machine Learning Department (MLD)