Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Value extrapolation, concept extrapolation, model splintering, published by Stuart Armstrong on March 8, 2022 on The AI Alignment Forum. Post written with Rebecca Gorman. We've written before that model splintering, as we called it then, was a problem with almost all AI safety approaches. There's a converse to this: solving the problem would help with almost all AI safety approaches. But so far, we've been posting mainly about value extrapolation. In this post, we'll start looking at how other AI safety approaches could be helped. Definitions To clarify, let's make three definitions, distinguishing ideas that we'd previously been grouping together: Model splintering is when the features and concepts that are valid in one world-model, break down when transitioning to another world-model. Concept extrapolation is extrapolating a feature or concept from one world-model to another. Value extrapolation is concept extrapolation when the particular concept to extrapolate is a value, a preference, a reward function, an agent's goal, or something of that nature. Examples Consider for example Turner et al's attainable utility. It has a formal definition, but the reason for that definition is that preserving attainable utility is aimed at restricting the "power" of the agent, or at minimising its "side effects". And it succeeds, in the typical situation. If you measure the attainable utility of an agent, this will give you an idea of its power, and how many side effects it may be causing. However, when we move to general situations, this breaks down: attainable utility preservation no longer restricts power or reduces side effects. So the concepts of power and side effects have splintered when moving from typical situations to general situations. This is the model splintering[1]. If we solve concept extrapolation for this, then we could extend the concepts of power restriction or side effect minimisation, to the general situations. And thus successfully create low impact AIs. Another example is wireheading. We have a reward signal that corresponds to something we desire in the world; maybe the negative of the CO2 concentration in the atmosphere. This is measured by, say, a series of CO2 detectors spread over the Earth's surface. Typically, the reward signal does correspond to what we want. But if the AI hacks its own reward signal, that correspondence breaks down[2]: model splintering. If we can extend the reward properly to new situations, we get concept extrapolation - which, since this is a reward function, is value extrapolation. Helping with multiple methods Hence the concept extrapolation/value extrapolation ideas can help with many different approaches to AI safety, not just the value learning approaches. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.