Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Modeling Failure Modes of High-Level Machine Intelligence, published by Ben Cottier on December 6, 2021 on The AI Alignment Forum. This post, which deals with some widely discussed failure modes of transformative AI, is part 7 in our sequence on Modeling Transformative AI Risk. In this series of posts, we are presenting a preliminary model of the relationships between key hypotheses in debates about catastrophic risks from AI. Previous posts in this sequence explained how different subtopics of this project, such as mesa-optimization and safety research agendas, are incorporated into our model. We now come to the potential failure modes caused by High-Level Machine Intelligence (HLMI). These failure modes are the principal outputs of our model. Ultimately, we need to look at these outputs to analyze the effect that upstream parts of our model have on the risks, and in turn to guide decision-making. This post will first explain the relevant high-level components of our model before going through each component in detail. In the figure below, failure modes are represented by red-colored nodes. We do not model catastrophe itself in detail, because there seem to be a very wide range of ways it could play out. Rather, the outcomes are states from which existential catastrophe is very likely, or "points of no return" for a range of catastrophes. Catastrophically Misaligned HLMI covers scenarios where one HLMI, or a coalition of HLMIs and possibly humans, achieves a Decisive Strategic Advantage (DSA), and the DSA enables an existential catastrophe to occur. Loss of Control covers existential scenarios where a DSA does not occur; instead HLMI systems gain influence on the world in an incremental and distributed fashion, and either humanity's control over the future gradually fades away, or a sudden change in the world causes a sufficiently large and irreversibly damaging automation failure (see What failure looks like); alternatively, the use of HLMI leads to an extreme "Moloch" scenario where most of what we value is burned down through competition. Moving one step back, there are three major drivers of those outcomes. Firstly, Misaligned HLMI is the technical failure to align an HLMI. Second, the DSA module considers different ways DSA could be achieved by a leading HLMI project or coalition. Third, Influence-seeking Behavior considers whether an HLMI would pursue accumulation of power (e.g. political, financial) and/or manipulation of humans as an instrumental goal. This third consideration is based on whether HLMI will be agent-like and to what extent the Instrumental Convergence thesis applies. Outcomes of HLMI Development As discussed in the "impact of safety agendas" post, we are still in the process of understanding the theory of change behind many of the safety agendas, and similarly our attempts at translating the success/failure of these safety agendas into models of HLMI development are still only approximate. With that said, for each of the green modules shown below on the Success of [agenda] Research, we model their affect on the likelihood of an incorrigible HLMI - in particular, if the research fails to produce its intended effect. Meanwhile, the Mesa-optimisation module evaluates the likelihood of dangers from learned optimization which would tend to make HLMI incorrigible. Besides Incorrigibility, the question of HLMI being aligned may be influenced by the expected speed and level of discontinuity during takeoff (represented by the modules Takeoff Speed and Discontinuity around HLMI without self-improvement, covered in Takeoff Speeds and Discontinuities). The expectation is that misalignment is more likely the faster progress becomes, because attempts at control will be harder if less time is available. Those factors - incorrigibility and takeoff - are more...