Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Continuity Assumptions, published by Jan Kulveit on June 13, 2022 on The Effective Altruism Forum. This post will try to explain what I mean by continuity assumptions (and discontinuity assumptions), and why differences in these are upstream of many disagreements about AI safety. None of this is particularly new, but seemed worth reiterating in the form of a short post. What I mean by continuity Start with a cliff: This is sometimes called the Heaviside step function. This function is actually discontinuous, and so it can represent categorical change. In contrast, this cliff is actually continuous: From a distance or in low resolution, it looks like a step function; it gets really steep. Yet it is fundamentally different. How is this relevant Over the years, I've come to the view that intuitions about whether we live in a "continuous” or "discontinuous” world are one of the few top principal components underlying all disagreements about AI safety. This includes but goes beyond the classical continuous vs. discrete takeoff debates.A lot of models of what can or can't work in AI alignment depends on intuitions about whether to expect "true discontinuities" or just "steep bits". This holds not just in one, but many relevant variables (e.g. the generality of the AI’s reasoning, the speed or differentiability of its takeoff). The discrete intuition usually leads to sharp categories like: Before the cliffAfter the cliffNon-general systems. Lack the core of general reasoning, that which allows thought in domains far from training dataGeneral systems. Capabilities generalise farWeak systems - that won't kill you, but also won't help you solve alignmentStrong systems - that would help solve alignment, but unfortunately will kill you by default, if unalignedSystems which may be misaligned, but aren't competently deceptive about it System which is actively modelling you at a level where the deception is beyond your ability to noticeWeak actsPivotal acts.. In Discrete World, empirical trends, alignment techniques, etc usually don't generalise across the categorical boundary. The right is far from the training distribution on the left. Your solutions don't survive the generality cliff, there are no fire alarms - and so on. Note that while assumptions about continuity in different dimensions are in principle not necessarily related, and you could e.g. assume continuity in takeoff and discontinuity in generality - in practice, they seem strongly correlated. Deep cruxes Deep priors over continuity versus discontinuity seem to be a crux which is hard to resolve. My guess is intuitions about continuity/discreteness are actually quite deep-seated: based more on how people do maths, rather than specific observations about the world. In practice, for most researchers, the "intuition" is something like a deep net trained on a whole lifetime of STEM reasoning - they won't update much on individual datapoints, and if they are smart, they are often able to re-interpret observations to be in line with their continuity priors. (As an example, compare Paul Christiano's post on takeoff speeds from 2018, which is heavily about continuity, to the debate between Paul and Eliezer in late 2021. Despite the participants spending years in discussion, progress on bridging the continuous-discrete gap between them seems very limited.) How continuity helps In basically every case, continuity implies the existence of systems "somewhere in between". Systems which are moderately strong: maybe weakly superhuman in some relevant domain and quite general, but at same time maybe still bad with plans on how to kill everyone. Moderately general systems: maybe are able of general reasoning, but in a strongly bounded way Proto-deceptive systems which are bad at deception. The existence of such systems helps us with t...