Monday, April 29, 2024, 10:30 – 11:30am
Distributional Reinforcement Learning (RL) fits Q-functions by learning the whole conditional distribution of rewards-to-go and then taking its mean (e.g., C51, IQN). Empirically it often improves on analogous approaches that learn the conditional mean directly (e.g., DQN) even in risk-neutral RL where we only care about the mean, but a principled understanding as to why and when has been elusive. We resolve this by showing that distributional RL enjoys first- and second-order regret bounds in both online and offline RL in general MDPs with function approximation. In some cases these are the first bounds of their kind for any RL algorithm. First-order bounds scale with the cost of the optimal policy and, for example, establish fast regret rates in goal-based MDPs when a policy exists reaching the goal reliably. Second-order bounds scale with the variance of the return and, for example, establish fast regret rates in nearly deterministic systems. We explain this phenomenon in terms of sensitivity to heteroskedasticity and demonstrate the predictions of the theory empirically on real-world tasks. Beyond distributional RL, I will discuss the implications for choosing the right loss function for decision making.
—
Nathan Kallus is an Associate Professor at the Cornell Tech campus of Cornell University in NYC and a Research Director at Netflix. Nathan's research interests include the statistics of optimization under uncertainty, causal inference especially when combined with machine learning, sequential and dynamic decision making, and algorithmic fairness. He holds a PhD in Operations Research from MIT as well as a BA in Mathematics and a BS in Computer Science from UC Berkeley. Before coming to Cornell, Nathan was a Visiting Scholar at USC's Department of Data Sciences and Operations and a Postdoctoral Associate at MIT's Operations Research and Statistics group.
Papers: