Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AXRP Episode 14 - Infra-Bayesian Physicalism with Vanessa Kosoy, published by DanielFilan on April 5, 2022 on The AI Alignment Forum. Google Podcasts link This podcast is called AXRP, pronounced axe-urp and short for the AI X-risk Research Podcast. Here, I (Daniel Filan) have conversations with researchers about their research. We discuss their work and hopefully get a sense of why it’s been written and how it might reduce the risk of artificial intelligence causing an existential catastrophe: that is, permanently and drastically curtailing humanity’s future potential. Late last year, Vanessa Kosoy and Alexander Appel published some research under the heading of “Infra-Bayesian physicalism”. But wait - what was infra-Bayesianism again? Why should we care? And what does any of this have to do with physicalism? In this episode, I talk with Vanessa Kosoy about these questions, and get a technical overview of how infra-Bayesian physicalism works and what its implications are. Topics we discuss: The basics of infra-Bayes An invitation to infra-Bayes What is naturalized induction? How infra-Bayesian physicalism helps with naturalized induction Bridge rules Logical uncertainty Open source game theory Logical counterfactuals Self-improvement How infra-Bayesian physicalism works World models Priors Counterfactuals Anthropics Loss functions The monotonicity principle How to care about various things Decision theory Follow-up research Infra-Bayesian physicalist quantum mechanics Infra-Bayesian physicalist agreement theorems The production of infra-Bayesianism research Bridge rules and malign priors Following Vanessa’s work Daniel Filan: Hello everybody. Today, I’m going to be talking with Vanessa Kosoy. She is a research associate at the Machine Intelligence Research Institute, and she’s worked for over 15 years in software engineering. About seven years ago, she started AI alignment research, and is now doing that full-time. Back in episode five, she was on the show to talk about a sequence of posts introducing Infra-Bayesianism. But today, we’re going to be talking about her recent post, Infra-Bayesian Physicalism: a Formal Theory of Naturalized Induction, co-authored with Alex Appel. For links to what we’re discussing, you can check the description of this episode, and you can read the transcript at axrp.net. Vanessa, welcome to AXRP. Vanessa Kosoy: Thank you for inviting me. The basics of infra-Bayes Daniel Filan: Cool. So, this episode is about Infra-Bayesian physicalism. Can you remind us of the basics of just what Infra-Bayesianism is? Vanessa Kosoy: Yes. Infra-Bayesianism is a theory we came up with to solve the problem of non-realizability, which is how to do theoretical analysis of reinforcement learning algorithms in situations where you cannot assume that the environment is in your hypothesis class, which is something that has not been studied much in the literature for reinforcement learning specifically. And the way we approach this is by bringing in concepts from so-called imprecise probability theory, which is something that’s mostly decision theorists and economists have been using. And the basic idea is, instead of thinking of a probability distribution, you could be working with a convex set of probability distributions. That’s what’s called a credal set in imprecise probability theory. And then, when you are making decisions, instead of just maximizing the expected value of your utility function, with respect to some probability distribution, you are maximizing the minimal expected value where you minimize over the set. That’s as if you imagine an adversary is selecting some distribution out of the set. Vanessa Kosoy: The nice thing about it is that you can start with this basic idea, and on the one hand, construct an entire theory analogous to classical pro...