Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Ngo's view on alignment difficulty, published by Richard Ngo on December 14, 2021 on The AI Alignment Forum. This post features a write-up by Richard Ngo on his views, with inline comments. Color key: Chat Google Doc content Inline comments 13. Follow-ups to the Ngo/Yudkowsky conversation 13.1. Alignment difficulty debate: Richard Ngo's case [Ngo][9:31] (Sep. 25) As promised, here's a write-up of some thoughts from my end. In particular, since I've spent a lot of the debate poking Eliezer about his views, I've tried here to put forward more positive beliefs of my own in this doc (along with some more specific claims): [GDocs link] [Soares: ✨] [Ngo] (Sep. 25 Google Doc) We take as a starting observation that a number of “grand challenges” in AI have been solved by AIs that are very far from the level of generality which people expected would be needed. Chess, once considered to be the pinnacle of human reasoning, was solved by an algorithm that’s essentially useless for real-world tasks. Go required more flexible learning algorithms, but policies which beat human performance are still nowhere near generalising to anything else; the same for StarCraft, DOTA, and the protein folding problem. Now it seems very plausible that AIs will even be able to pass (many versions of) the Turing Test while still being a long way from AGI. [Yudkowsky][11:26] (Sep. 25 comment) Now it seems very plausible that AIs will even be able to pass (many versions of) the Turing Test while still being a long way from AGI. I remark: Restricted versions of the Turing Test. Unrestricted passing of the Turing Test happens after the world ends. Consider how smart you'd have to be to pose as an AGI to an AGI; you'd need all the cognitive powers of an AGI as well as all of your human powers. [Ngo][11:24] (Sep. 29 comment) Perhaps we can quantify the Turing test by asking something like: What percentile of competence is the judge? What percentile of competence are the humans who the AI is meant to pass as? How much effort does the judge put in (measured in, say, hours of strategic preparation)? Does this framing seem reasonable to you? And if so, what are the highest numbers for each of these metrics that correspond to a Turing test which an AI could plausibly pass before the world ends? [Ngo] (Sep. 25 Google Doc) I expect this trend to continue until after we have AIs which are superhuman at mathematical theorem-proving, programming, many other white-collar jobs, and many types of scientific research. It seems like Eliezer doesn't. I’ll highlight two specific disagreements which seem to play into this. [Yudkowsky][11:28] (Sep. 25 comment) doesn't Eh? I'm pretty fine with something proving the Riemann Hypothesis before the world ends. It came up during my recent debate with Paul, in fact. Not so fine with something designing nanomachinery that can be built by factories built by proteins. They're legitimately different orders of problem, and it's no coincidence that the second one has a path to pivotal impact, and the first does not. [Ngo] (Sep. 25 Google Doc) A first disagreement is related to Eliezer’s characterisation of GPT-3 as a shallow pattern-memoriser. I think there’s a continuous spectrum between pattern-memorisation and general intelligence. In order to memorise more and more patterns, you need to start understanding them at a high level of abstraction, draw inferences about parts of the patterns based on other parts, and so on. When those patterns are drawn from the real world, then this process leads to the gradual development of a world-model. This position seems more consistent with the success of deep learning so far than Eliezer’s position (although my advocacy of it loses points for being post-hoc; I was closer to Eliezer’s position before the GPTs). It also predicts that deep learning ...