Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Stable self-improvement as a research problem, published by Paul Christiano on the AI Alignment Forum. “Stable self-improvement" seems to be a primary focus of MIRI’s work. As I understand it, the problem is "How do we build an agent which rationally pursues some goal, is willing to modify itself, and with very high probability continues to pursue the same goal after modification?" The key difficulty is that it is impossible for an agent to formally "trust" its own reasoning, i.e. to believe that "anything that I believe is true." Indeed, even the natural concept of "truth" is logically problematic. But without such a notion of trust, why should an agent even believe that its own continued existence is valuable? I agree that there are open philosophical questions concerning reasoning under logical uncertainty, and that reflective reasoning highlights some of the difficulties. But I am not yet convinced that stable self-improvement as an especially important problem; I think it would be handled correctly by a human-level reasoner as a special case of decision-making under logical uncertainty. This suggests that (1) it will probably be resolved en route to human-level AI, (2) it can probably be "safely" delegated to a human-level AI. I would prefer for energy to be used on other aspects of the AI safety problem. Consider an agent A which shares our values and is able to reason "as well as we are"---for any particular empirical or mathematical quantity, A's estimate of its expectation is as good as ours. For notational convenience, suppose that A's preferences are the same as "our" preferences, and let U be the associated utility function. Now suppose that A is thinking about an outcome including the existence of an agent B. (Perhaps B is a new AI that A is considering designing; perhaps B is a version of A that has made some further observations; whatever.) We'd like the agent to evaluate this outcome on its merits. It should think about how good the existence of B is. If B also maximizes U, then A should correctly understand that B's existence will tend to be good. The expected value of U conditioned on this outcome is just another empirical quantity. If A is as good at estimation as humans, then it won't predictably over- or under-estimate this quantity. And so it will weigh B's existence correctly when considering the consequences of its actions. So if we really had a “human-level” reasoner in the sense I assumed at the outset, our problem would be solved. There are a number of reasons to think the problem might be important anyway. I haven’t seen any of these arguments fleshed out in much detail, and for the most part I am skeptical.
If we anticipate a long sequence of ever-more-powerful AI's, then we might want to be very sure that each change is really an improvement. There are two sides to this concern. First is the idea that an AI might not exercise sufficient caution when designing a successor. But if the AI has well-calibrated beliefs and shares our values, then by construction it will make the appropriate tradeoffs between reliability and efficiency. So I don't take this concern very seriously. Second is the concern that, if the required confidence is very high, then it might be very difficult to be confident enough to go ahead with a proposed AI design. In this scenario, an AI might correctly realize that it should not make any risky changes; but this restriction might introduce unacceptable efficiency losses. While the "good guys" proceed cautiously, competitors will race ahead (allowing their systems' values to change over time). On this view, by working out these issues farther in advance we can save some time for the “good guys,” or push research in a direction which makes their task easier. But this problem c...