Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Are there alternative to solving value transfer and extrapolation?, published by Stuart Armstrong on December 6, 2021 on The AI Alignment Forum. Thanks to Rebecca Gorman for the discussions that inspired these ideas. We recently argued that full understanding of value extrapolation[1] was necessary and almost sufficient for solving the AI alignment problem, due to the existence of morally underdefined situations. People seemed to generally agree that some knowledge of human morality was needed for a safe powerful AI. But it wasn't clear this AI needed almost-full knowledge of human morality, as well as value extrapolation. We got some high-quality alternative suggestions. In this post, we'll try and distil those ideas and analyse them. AI seems to behave reasonably Steven Byrnes's position, if we understand it correctly, is that the AI should learn to behave in non-dangerous seeming ways[2]. Our phrasing of the idea would be (apologies for any misinterpretations): There are many behaviours that agents can follow. Typical humans follow typical behaviours, to achieve typical consequences. And other typical humans can assess both behaviours and consequences, and grade them as (typical) good, (typical) bad, or weird. An AI should behave in ways that typical humans judge as good, achieving consequences that typical humans judge as good. If there is too high a risk of weirdness, the AI will output NOOP (no operation). This seems a sensible approach. But it has a crucial flaw. The AI must behave in typical ways to achieve typical consequences - but how it achieves these consequences doesn't have to be typical. For example, it might make an explanatory video to convince someone to follow some (reasonable) course of action. The video itself is not unusual, but the colour scheme happens to hypnotise that specific viewer into following the suggestion. At a gross level of description, everything is fine: convincing-but-typical video convinces. But this only works because the AI has inhuman levels of knowledge and predictive ability, and carefully selects the "typical behaviour" that is the most effective. Now, we might be able to make sure the AI doesn't have secret superhuman abilities (we've had an old underdeveloped idea that we could force the AI to use human models of the world to achieve its goals), but this is a very different problem, and a much harder one. Then, given that the AI has access to superhuman levels of ability, the "typical consequences" is no longer a safe goal. It reduces to "consequences that seem typical to an observing human", which is not a safe goal for the AI; it can now do whatever it wants, as long as its actions and consequences look ok. Still, it might be worth developing this idea; it might work well combined with some other safety approaches. Advanced human feedback Rohin's argument is that we don't need the AI to solve value transfer and extrapolation. Instead, we just need a "well-motivated" AI that asks us what we prefer. It must present the options and the consequences in an informative manner. If there are morally underdefined questions where we might give multiple answers depending on question phrasing, then the AI goes to a meta level and asks us how we would want it to ask us. This is a potentially powerful approach. It solves the value extrapolation problem indirectly, by deference to humans. It is similar to a scaled-up version of "informed consent". Would such an approach work to contain a malevolent AI? If a superintelligent AI wished to do us harm, but was constrained to following the above feedback approach, would it be safe? It seems clear that it wouldn't. The malevolent AI would deploy all its resourcefulness to undermine the spirit of the feedback requirement. Unless we had it perfectly defined, this would just be a speedbump ...