Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: (My understanding of) What Everyone in Technical Alignment is Doing and Why, published by Thomas Larsen on August 29, 2022 on The AI Alignment Forum. Epistemic Status: My best guess Epistemic Effort: ~50 hours of work put into this document Contributions: Thomas wrote ~85% of this, Eli wrote ~15% and helped edit + structure it. Unless specified otherwise, writing in the first person is by Thomas and so are the opinions. Thanks to Miranda Zhang, Caleb Parikh, and Akash Wasil for comments. Thanks to many others for relevant conversations. Introduction Despite a clear need for it, a good source explaining who is doing what and why in technical AI alignment doesn't exist. This is our attempt to produce such a resource. We expect to be inaccurate in some ways, but it seems great to get out there and let Cunningham’s Law do its thing. The main body contains our understanding of what everyone is doing in technical alignment and why, as well as at least one of our opinions on each approach. We include supplements visualizing differences between approaches and Thomas’s big picture view on alignment. The opinions written are Thomas and Eli’s independent impressions, many of which have low resilience. Our all-things-considered views are significantly more uncertain. A summary of our understanding of each approach: Problem FocusCurrent Approach SummaryModel splinteringSolve extrapolation problems. Inaccessible informationELK + LLM power-seeking evaluationLack of good interpretability tools (?)Interpretability + HHH + augmenting alignment research with LLMsBrain-like AGI SafetyUse brains as a model for how AGI will be developed, think about alignment in this contextEngaging the ML community, many technical problems Technical research, Infrastructure, and ML community field-building for safetyOuter alignment, though CHAI is diverseImprove CIRL + many other independent approaches. Suffering risksFoundational game theory researchInner alignmentInterpretability + automating alignment research with LLMsMany including scalable oversight and goal misgeneralizationMany including Debate, ERO, and discovering agents. Multipolar failure from lack of coordinationVideo gameDeceptionGet the reasoning of the AGI to happen in natural language, then oversee that reasoningMany (?)Incubate new, scalable alignment research agendasMany including deception, the sharp left turn, corrigibility is anti-naturalMathematical research to resolve fundamental confusion about the nature of goals/agency/optimizationScalable oversightRLHF / Recursive Reward Modeling, then automate alignment researchScalable oversightSupervise process rather than outcomes + augment alignment researchersInner alignment (?)Interpretability + Adversarial Training Being able to robustly point at objects in the worldSelection Theorems based on natural abstractionsInstilling inner values from an outer training loopFind patterns of values given by current RL setups and humans, then create quantitative rules to do thisDeceptionCreate standards and datasets to evaluate model truthfulness Approach Aligned AI ARC Anthropic Brain-like-AGI Safety CAIS CHAI CLR Conjecture DeepMind Encultured Externalized Reasoning Oversight FAR MIRI OpenAI Ought Redwood Selection Theorems Team Shard Truthful AI Previous related overviews include: Neel Nanda's My Overview of the AI Alignment Landscape Evan Hubinger's An overview of 11 proposals for building safe advanced AI Larks' yearly Alignment Literature Review and Charity Comparison Nate Soares' On how various plans miss the hard bits of the alignment challenge Andrew Critch's Some AI research areas and their relevance to existential safety 80,000 Hours’ list of organizations working in the area Aligned AI / Stuart Armstrong One of the key problems in AI safety is that there are many ways for an AI to gener...