Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Mediocre AI safety as existential risk, published by Gavin on March 16, 2022 on The Effective Altruism Forum. Epistemic status: Written in one 2 hour session for a deadline. Probably ill-conceptualised in some way I can't quite make out. Broader impacts: Could underwrite unfair cynicism. Has been read by a couple careful alignment people who didn't hate it. I propose a premortem. The familiar sense of ‘dual-use’ technology (when a civilian technology has military implications) receives a gratifying amount of EA, popular, and government attention. But consider a different sense: AI alignment (AIA) or AI governance work which actually increases existential risk. This tragic line of inquiry has some basic theory under the names ‘differential progress’, ‘accidental harm’, and a generalised sense of ‘dual-use’. Some missing extensions: there is almost no public evaluation of the downside risks of particular AIA agendas, projects, or organisations. (Some evaluations exist in private but I, a relative insider, have access to only one, a private survey by a major org. I understand why the contents might not be public, but the existence of the documents seems important to publicise.) In the (permanent) absence of concrete feedback on these projects, we are trading in products which neither producer nor consumer know the quality of. We should model this. Tools from economics could help us reason about our situation (see Methods). As David Krueger noted some years ago, there is little serious public thought regarding how much AI capabilities work it is wise to do for e.g. career capital or research training for young alignment researchers. (There’s a trivial sense that mediocre projects increase existential risk: they represent an opportunity cost, by nominally taking resources from good projects. I instead mean the nontrivial sense that the work could actively increase risk.) Example: Reward learning Some work in ML safety will enable the deployment of new systems. Ben Garfinkel gives the example of a robot cleaner: Let’s say you’re trying to develop a robotic system that can clean a house as well as a human house-cleaner can... This is essentially an alignment problem... until we actually develop these techniques, probably we’re not in a position to develop anything that even really looks like it’s trying to clean a house, or anything that anyone would ever really want to deploy in the real world. He sees this as positive: it implies massive economic incentives to do some alignment, and a block on capabilities until this is done. But it could be a liability as well, if the alignment of weak systems is correspondingly weak, and if mid-term safety work fed into a capabilities feedback loop with greater amplification. (That is, successful deployment means profit, which means reinvestment and induced investment in AI capabilities.) More generally, human modelling approaches to alignment risk improving the capability of deceiving operators, and invite beyond-catastrophic ‘alignment near-misses’ (i.e. S-risks). Methods 1. Private audits plus canaries. Interview members of AIA projects under an NDA, or with the interviewee anonymous to me. The resulting public writeup then merely reports 5 bits of information about each project: 1) whether the organisation has a process for managing accidental harm, 2) whether this has been vetted by any independent party, 3) whether any project has in fact been curtailed as a result, 4) because of potentially dangerous capabilities or not, and 5) whether we are persuaded that they are net-positive. Refusal to engage is also noted. This process has problems (e.g. positivity bias from employees, or the audit team's credibility) but seems the best thing we can do with private endeavours, short of soliciting whistleblowers. Audit the auditors too, why ...