Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AGI safety from first principles: Control, published by Richard Ngo on the AI Alignment Forum. It’s important to note that my previous arguments by themselves do not imply that AGIs will end up in control of the world instead of us. As an analogy, scientific knowledge allows us to be much more capable than stone-age humans. Yet if dropped back in that time with just our current knowledge, I very much doubt that one modern human could take over the stone-age world. Rather, this last step of the argument relies on additional predictions about the dynamics of the transition from humans being the smartest agents on Earth to AGIs taking over that role. These will depend on technological, economic and political factors, as I’ll discuss in this section. One recurring theme will be the importance of our expectation that AGIs will be deployed as software that can be run on many different computers, rather than being tied to a specific piece of hardware as humans are.[1] I’ll start off by discussing two very high-level arguments. The first is that being more generally intelligent allows you to acquire more power, via large-scale coordination and development of novel technological capabilities. Both of these contributed to the human species taking control of the world; and they both contributed to other big shifts in the distribution of power (such as the industrial revolution). If the set of all humans and aligned AGIs is much less capable in these two ways than the set of all misaligned AGIs, then we should expect the latter to develop more novel technologies, and use them to amass more resources, unless strong constraints are placed on them, or they’re unable to coordinate well (I’ll discuss both possibilities shortly.) On the other hand, though, it’s also very hard to take over the world. In particular, if people in power see their positions being eroded, it’s generally a safe bet that they’ll take action to prevent that. Further, it’s always much easier to understand and reason about a problem when it’s more concrete and tangible; our track record at predicting large-scale future developments is pretty bad. And so even if the high-level arguments laid out above seem difficult to rebut, there may well be some solutions we missed which people will spot when their incentives to do so, and the range of approaches available to them, are laid out more clearly. How can we move beyond these high-level arguments? In the rest of this section I’ll lay out two types of disaster scenarios, and then four factors which will affect our ability to remain in control if we develop AGIs that are not fully aligned: Speed of AI development Transparency of AI systems Constrained deployment strategies Human political and economic coordination Disaster scenarios There have been a number of attempts to describe the catastrophic outcomes that might arise from misaligned superintelligences, although it has proven difficult to characterise them in detail. Broadly speaking, the most compelling scenarios fall into two categories. Christiano describes AGIs gaining influence within our current economic and political systems by taking or being given control of companies and institutions. Eventually “we reach the point where we could not recover from a correlated automation failure” - after which those AGIs are no longer incentivised to follow human laws. Hanson also lays out a scenario in which virtual minds come to dominate the economy (although he is less worried about misalignment, partly because he focuses on emulated human minds). In both scenarios, biological humans lose influence because they are less competitive at strategically important tasks, but no single AGI is able to seize control of the world. To some extent these scenarios are analogous to our current situation, in which large corporations...