Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: People care about each other even though they have imperfect motivational pointers?, published by Alex Turner on November 8, 2022 on The AI Alignment Forum. I wrote this essay in early August. I now consider the presentation to be somewhat confused, and now better understand where problems arise within the "standard alignment model." I'm publishing a somewhat edited version, on the grounds that something is better than nothing. Summary: Consider the argument: "Imperfect value representations will, in the limit of optimization power, be optimized into oblivion by the true goal we really wanted the AI to optimize." But... I think my brother cares about me in some human and "imperfect" way. I also think that the future would contain lots of value for me if he were a superintelligent dictator (this could be quite bad for other reasons, of course). Therefore, this argument seems to prove too much. It seems like one of the following must be true: My brother cares about me in the "perfect" way, or Dictator-brother would do valueless things according to my true values, or The argument is wrong. To explore these points, I dialogue with my model of Eliezer. Suppose you win a raffle and get to choose one of n prizes. The first prize is a book with true value 10, but your evaluation of it is noisy (drawn from the Gaussian N(10,1) with standard deviation 1). The other n-1 prizes are widgets with true value 1, but your evaluation is more noisy (drawn from N(1,16) with standard deviation 4). As n increases, you’re more probable to select a widget and lose out on 10-1=9 utility. By considering so many options, you’re selecting against your own ability to judge prizes by implicitly selecting for high noise. You end up “optimizing so hard” that you delude yourself. This is the Optimizer’s Curse. You’re probably already familiar with Goodhart’s Law, which applies when an agent optimizes a proxy U (e.g. how many nails are produced) for the true quantity V which we value (e.g. how profitable the nail factory is). Goodhart’s Curse is their combination. According to the article, optimizing a proxy measure U over a trillion plans can lead to high regret under the true values V, even if U is an unbiased but noisy estimator of the true values V. This seems like bad news, insofar as it suggests that even getting an AI which understands human values on average across possibilities can still produce bad outcomes. Below is a dialogue with my model of Eliezer. I wrote the dialogue to help me think about the question at hand. I put in work to model his counterarguments, but ultimately make no claim to have written him in a way he would endorse. An obvious next question is "Why not just define the AI such that the AI itself regards U [a proxy measure] as an estimate of V [true human values], causing the AI's [proxy measure] to more closely align with [true human values] as the AI gets a more accurate empirical picture of the world?" Reply: Of course this is the obvious thing that we'd want to do. But what if we make an error in exactly how we define "treat U as an estimate of V"? Goodhart's Curse will magnify and blow up any error in this definition as well. — Goodhart’s Curse Alex (A): Suppose that I can get smarter over time—that AGI just doesn’t happen, that the march of reason continues, that lifespans extend, and that I therefore gradually become very old and very smart. Goodhart’s Curse predicts that—even though I think I want to help my brother be happy, even though I think I value his having a good life by his own values—my desire to help him is not “error-free”, it is not perfect, and so Goodhart’s Curse will magnify and blow up these errors. And time unfolds, the Curse is realized, and I bring about a future which is high-regret by his “true values” V. Insofar as Goodhart’s Curse has teeth...