Tuesday, November 22, 2022, 12 – 1pm
We will introduce the problem of online model selection where a learner is to select among a set of online algorithms to solve a specific problem instance. We would like to design algorithms that allow such a learner to select in an online fashion the best algorithm without incurring much regret. This problem is challenging because in contrast with for example multi armed bandits, the algorithms' rewards -due to the algorithm's own learning process- may be non-stationary. We will introduce the principle of regret balancing, a simple, practical and effective model selection algorithmic design technique that allows for online selection of the best among multiple (base) algorithms in a fully blackbox fashion. Regret balancing solves the problem of non-stationarity by introducing an elegant `misspecification test' that can efficiently detect when a base algorithm is not appropriate for the problem at hand. Regret balancing techniques have also been used to provide clarity to some long-standing problems in online learning such as corruption learning in MDPs.
—
Aldo Pacchiano is a postdoctoral researcher at Microsoft Research NYC. He obtained his PhD at UC Berkeley where he was advised by Prof. Peter Bartlett and Prof. Michael Jordan. His research lies in the areas of Reinforcement Learning, Online Learning, Bandits and Algorithmic Fairness. He is particularly interested in furthering our statistical understanding of learning phenomena in adaptive environments and use these theoretical insights and techniques to design useful algorithms in (among other things) bandits, RL, and experimental design.
*The AI Seminar is generously sponsored by SambaNova Systems.
In Person and Zoom Participation. See announcement.*
Event Type: Seminars
Room Number: In Person and Virtual - ET
Building: Newell-Simon 3305 and Zoom
Speaker's Name: ALDO PACCHIANO
Speaker Website: www.aldopacchiano.ai…
Speaker's Professional Title: Postdoctoral Researcher, Microsoft Researcher NYC
Talk Title: Online Model Selection: the principle of regret balancing
For More Information: ashert@cs.cmu.edu
Affiliations: Computer Science Department (CSD), Language Technologies Institute (LTI), Machine Learning Department (MLD)
Event Website Title: Series Website
Event Website URL: www.cs.cmu.edu…