Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Alignment versus AI Alignment, published by Alex Flint on February 4, 2022 on The AI Alignment Forum. Financial status: This work is supported by individual donors and a grant from LTFF. Epistemic status: This post contains many inside-view stories about the difficulty of alignment. Thanks to Adam Shimi, John Wentworth, and Rob Miles for comments on this essay. What exactly is difficult about AI alignment that is not also difficult about alignment of governments, economies, companies, and other non-AI systems? Is it merely that the fast speed of AI makes the AI alignment problem quantitatively more acute than other alignment problems, or are there deeper qualitative differences? Is there a real connection between alignment of AI and non-AI systems at all? In this essay we attempt to clarify which difficulties of AI alignment show up similarly in non-AI systems, and which do not. Our goal is to provide a frame for importing and exporting insights from and to other fields without losing sight of the difficulties of AI alignment that are unique. Clarifying which difficulties are shared should clarify the difficulties that are truly unusual about AI alignment. We begin with a series of examples of aligning different kinds of systems, then we seek explanations for the relative difficulty of AI and non-AI alignment. Alignment in general In general, we take actions in the world in order to steer the future in a certain direction. One particular approach to steering the future is to take actions that influence the constitutions of some intelligent system in the world. A general property of intelligent systems seems to be that there are interventions one can execute on them that have robustly long-lasting effects, such as changing the genome of a bacterium, or the trade regulations of a market economy. These are the aspects of the respective intelligent systems that persist through time and dictate their equilibrium behavior. In contrast, although plucking a single hair from a human head or adding a single barrel of oil to a market does have an impact on the future, the self-correcting mechanisms of the respective intelligent systems negate rather than propagate such changes. Furthermore, we will take alignment in general to be about utilizing such interventions on intelligent systems to realize our true terminal values. Therefore we will adopt the following working definition of alignment: Successfully taking an action that steers the future in the direction of our true terminal values by influencing the part of an intelligent system that dictates its equilibrium behavior. Our question is: in what ways is the difficulty of alignment of AI systems different from that of non-AI systems? Example: Aligning an economic society by establishing property rights Suppose the thing we are trying to align is a human society and that we view that thing as a collection of households and firms making purchasing decisions that maximize their individual utilities. Suppose that we take as a working operationalization of our terminal values the maximization of the sum of individual utilities of the people in the society. Then we might proceed by creating the conditions for the free exchange of goods and services between the households and firms, perhaps by setting up a government that enforces property rights. This is one particular approach to aligning a thing (a human society) with an operationalization of The Good (maximization of the sum of individual utilities). This particular approach works by structuring the environment in which the humans live in such a way that the equilibrium behavior of the society brings about accomplishment of the goal. We have: An intelligent system being aligned, which in this case is a human society. A model of that system, which in this case is a collection o...