Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: My thoughts on OpenAI's Alignment plan, published by Donald Hobson on December 10, 2022 on The AI Alignment Forum.This is my attempt at Eliezer's challenge.My overall impression of this plan is that it is surprisingly good. Not that I am particularly surprised such a plan exists, if MIRI had created this plan it would have been surprisingly bad. I think it is somewhat plausible that this plan could actually work, or at least some steel-manned version could work if several unknown parameters are set to favorable values.Our alignment research aims to make artificial general intelligence (AGI) aligned with human values and follow human intent. We take an iterative, empirical approach: by attempting to align highly capable AI systems, we can learn what works and what doesn’t, thus refining our ability to make AI systems safer and more aligned. Using scientific experiments, we study how alignment techniques scale and where they will break.We tackle alignment problems both in our most capable AI systems as well as alignment problems that we expect to encounter on our path to AGI. Our main goal is to push current alignment ideas as far as possible, and to understand and document precisely how they can succeed or why they will fail.You stand in a vast minefield, one end armed with popping balloons, the other end armed with some quantum vacuum decay weapon. Between there is a huge range, fireworks, conventional mines, nukes, antimatter and more.Your plan is to wander the safer parts of the minefield, recording where mines go off, in the hope you spot a pattern that can lead you through the whole minefield.This plan is not hopelessly doomed. But it is risky. You definitely need a way to detect that dangerous terrain is approaching, and to stop before you get there. Your job isn't to charge ahead. It is to march back and forth over swaths of balloon filled ground, carefully recording every balloon that pops, and scrutinizing the data for a pattern. Maybe you need to venture a little further into the space of firecrackers. But have a plan for where you stop, and an idea what the danger signals would be.It may be that there is no pattern in the mines. Or at least none you can discern. In which case, you don't venture further. You don't keep on, hoping that you will spot a pattern with just a few more large fireworks. You go back to base. And you hope that someone somewhere has been working on an airplane to fly over the minefield, and you can ask to help out with that.We believe that even without fundamentally new alignment ideas, we can likely build sufficiently aligned AI systems to substantially advance alignment research itself.That is at least a possibility. A favorable bit not totally implausible setting of those hidden dials. You probably need some new ideas, and if someone shows you a paper on say conservative learning, how to learn a classification boundary that is big enough to fit the datapoints and no bigger, be ready to read that paper.Unaligned AGI could pose substantial risks to humanity and solving the AGI alignment problem could be so difficult that it will require all of humanity to work together. Therefore we are committed to openly sharing our alignment research when it’s safe to do so: We want to be transparent about how well our alignment techniques actually work in practice and we want every AGI developer to use the world’s best alignment techniques.I do hope you have some procedure for deciding "when it's safe to do so". And ideally a way to share your results with a few other top labs, if you deem something safe enough to share with Deepmind or MIRI, but not safe enough to make public.At a high-level, our approach to alignment research focuses on engineering a scalable training signal for very smart AI systems that is aligned with human inten...