Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AGI safety from first principles: Alignment, published by Richard Ngo on the AI Alignment Forum. Parts of this section were rewritten in mid-October. In the previous section, I discussed the plausibility of ML-based agents developing the capability to seek influence for instrumental reasons. This would not be a problem if they do so only in the ways that are aligned with human values. Indeed, many of the benefits we expect from AGIs will require them to wield power to influence the world. And by default, AI researchers will apply their efforts towards making agents do whatever tasks those researchers desire, rather than learning to be disobedient. However, there are reasons to worry that despite such efforts by AI researchers, AIs will develop undesirable final goals which lead to conflict with humans. To start with, what does “aligned with human values” even mean? Following Gabriel and Christiano, I’ll distinguish between two types of interpretations. Minimalist (aka narrow) approaches focus on avoiding catastrophic outcomes. The best example is Christiano’s concept of intent alignment: “When I say an AI A is aligned with an operator H, I mean: A is trying to do what H wants it to do.” While there will always be some edge cases in figuring out a given human’s intentions, there is at least a rough commonsense interpretation. By contrast, maximalist (aka ambitious) approaches attempt to make AIs adopt or defer to a specific overarching set of values - like a particular moral theory, or a global democratic consensus, or a meta-level procedure for deciding between moral theories. My opinion is that defining alignment in maximalist terms is unhelpful, because it bundles together technical, ethical and political problems. While it may be the case that we need to make progress on all of these, assumptions about the latter two can significantly reduce clarity about technical issues. So from now on, when I refer to alignment, I’ll only refer to intent alignment. I’ll also define an AI A to be misaligned with a human H if H would want A not to do what A is trying to do (if H were aware of A’s intentions). This implies that AIs could potentially be neither aligned nor misaligned with an operator - for example, if they only do things which the operator doesn’t care about. Whether an AI qualifies as aligned or misaligned obviously depends a lot on who the operator is, but for the purposes of this report I’ll focus on AIs which are clearly misaligned with respect to most humans. One important feature of these definitions: by using the word “trying”, they focus on the AI’s intentions, not the actual outcomes achieved. I think this makes sense because we should expect AGIs to be very good at understanding the world, and so the key safety problem is setting their intentions correctly. In particular, I want to be clear that when I talk about misaligned AGI, the central example in my mind is not agents that misbehave just because they misunderstand what we want, or interpret our instructions overly literally (which Bostrom calls “perverse instantiation”). It seems likely that AGIs will understand the intentions of our instructions very well by default. This is because they will probably be trained on tasks involving humans, and human data - and understanding human minds is particularly important for acting competently in those tasks and the rest of the world.[1] Rather, my main concern is that AGIs will understand what we want, but just not care, because the motivations they acquired during training weren’t those we intended them to have. The idea that AIs won’t automatically gain the right motivations by virtue of being more intelligent is an implication of Bostrom’s orthogonality thesis, which states that “more or less any level of intelligence could in principle be combined with mo...