Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Abstracting The Hardness of Alignment: Unbounded Atomic Optimization, published by Adam Shimi on July 29, 2022 on The AI Alignment Forum. This work has been done while at Conjecture Disagree to Agree (Practically-A-Book Review: Yudkowsky Contra Ngo On Agents, Scott Alexander, 2022) This is a weird dialogue to start with. It grants so many assumptions about the risk of future AI that most of you probably think both participants are crazy. (Personal Communication about a conversation with Evan Hubinger, John Wentworth, 2022) We'd definitely rank proposals very differently, within the "good" ones, but we both thought we'd basically agree on the divide between "any hope at all" and "no hope at all". The question dividing the "any hope at all" proposals from the "no hope at all" is something like... does this proposal have any theory of change? Any actual model of how it will stop humanity from being wiped out by AI? Or is it just sort of... vaguely mood-affiliating with alignment? If there's one thing alignment researchers excel at, it's disagreeing with each other. I dislike the term pre paradigmatic, but even I must admit that it captures one obvious feature of the alignment field: the constant debates about the what and the how and the value of different attempts. Recently, we even had a whole sequence of debates, and since I first wrote this post Nate shared his take on why he can’t see any current work in the field actually tackling the problem. More generally, the culture of disagreement and debate and criticism is obvious to anyone reading the AF. Yet Scott Alexander has a point: behind all these disagreements lies so much agreement! Not only in discriminating the "any hope at all" proposals from the "no hope at all", as in John's quote above; agreement also manifests itself in the common components of the different research traditions, for example in their favorite scenarios. When I look at Eliezer's FOOM, at Paul's What failure looks like, at Critch's RAAPs, and at Evan's Homogeneous takeoffs, the differences and incompatibilities jump to me — yet they still all point in the same general direction. So much so that one can wonder if a significant part of the problem lies outside of the fine details of these debates. In this post, I start from this hunch — deep commonalities — and craft an abstraction that highlights it: unbounded atomic optimization (abbreviated UAO and pronounced wow). That is, alignment as the problem of dealing with impact on the world (optimization) that is both of unknown magnitude (unbounded) and non-interruptible (atomic). As any model, it is necessarily mistaken in some way; I nonetheless believe it to be a productive mistake, because it reveals both what we can do without the details and what these details give us when they're filled in. As such, UAO strikes me as a great tool for epistemological vigilance. I first present UAO in more details; then I show its use as a mental tool by giving four applications: (Convergence of AI Risk) UAO makes clear that the worries about AI Risk don’t come from one particular form of technology or scenario, but from a general principle which we’re pushing towards in a myriad of convergent ways. (Exploration of Conditions for AI Risk) UAO is only a mechanism; but it’s abstraction makes it helpful to study what conditions about the world and how we apply optimization lead to AI Risk (Operationalization Pluralism) UAO, as an abstraction of the problem, admits many distinct operationalizations. It’s thus a great basis on which to build operationalization pluralism. (Distinguishing AI Alignment) Last but not least, UAO answers Alex Flint’s question about the difference between aligning AIs and aligning other entities (like a society). Thanks to TJ, Alex Flint, John Wentworth, Connor Leahy, Kyle McDonell, Lari...