AXRP (pronounced axe-urp) is the AI X-risk Research Podcast where I, Daniel Filan, have conversations with researchers about their papers. We discuss the paper, and hopefully get a sense of why it's been written and how it might reduce the risk of AI causing an existential catastrophe: that is, permanently and drastically curtailing humanity's future potential. You can visit the website and read transcripts at axrp.net.
Remember AI 2027? Not AI 2040, the newest coolest thing AI Futures Project has done, but AI 2027, their OG product? At long last, we have an AXRP episode about it. Enjoy!
Transcript: https://axrp.net/episode/2026/08/03/episode-50-eli-lifland-ai-2027.html
Topics we discuss, and timestamps:
0:00:12 What is AI 2027?
0:10:17 What happens in AI 2027?
0:18:08 Why two endings?
0:21:15 Who did what?
0:23:59 Why superhuman AI in 2027?
0:32:53 Forecasting time horizon growth
0:48:34 When do time horizons go infinite?
1:02:13 Forecasting effective compute growth
1:07:09 From superhuman coders to superintelligence
1:20:58 How many AI companies?
1:26:29 What AGI will want
1:39:48 What misaligned AI does
1:52:20 Will AIs be able to align their successors
1:57:21 Why so long until AI takeover?
2:03:06 Would misaligned AGI kill us?
2:04:53 Will there just be one AGI?
2:13:25 The reception of AI 2027
2:18:57 What do you now think about takeoff?
2:26:31 What's next for AI Futures Project
2:31:36 How to work on AI forecasting
2:38:01 Following Eli's and AI Futures Project's work
Links to AI 2027 and related research:
AI 2027: https://ai-2027.com/
AI Futures Project blog: https://blog.ai-futures.org
AI Futures Research Notes: https://aifuturesnotes.substack.com/
AI Futures Model: https://www.aifuturesmodel.com/
X/Twitter links:
Eli Lifland: https://x.com/eli_lifland
Daniel Kokotajlo: https://x.com/dkokotajlo
AI Futures Project: https://x.com/AI_futures_
Research we discuss:
Task-Completion Time Horizons of Frontier AI Models: https://metr.org/time-horizons/
What Happens When Superhuman AIs Compete for Control?: https://blog.ai-futures.org/p/what-happens-when-superhuman-ais
How AI Takeover Might Happen In 2 Years: https://www.alignmentforum.org/posts/KFJ2LFogYqzfGB3uX/how-ai-takeover-might-happen-in-2-years
Guive Assadi on AI Property Rights (AXRP): https://axrp.net/episode/2026/02/15/episode-48-guive-assadi-ai-property-rights.html
Titotal: A deep critique of AI 2027's bad timeline models: https://titotal.substack.com/p/a-deep-critique-of-ai-2027s-bad-timeline
Response to titotal's critique of our AI 2027 timelines model: https://www.lesswrong.com/posts/G7MmNkYADKkmCiumj/response-to-titotal-s-critique-of-our-ai-2027-timelines
What you can do about AI 2027: https://blog.aifutures.org/p/what-you-can-do-about-ai-2027
Episode art by Hamish Doodles: hamishdoodles.com
How does game theory work when everyone is a computer program who can read everyone else's source code? This is the problem of 'program equilibria'. In this episode, I talk with Caspar Oesterheld on work he's done on equilibria of programs that simulate each other, and how robust these equilibria are.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2026/02/18/episode-49-caspar-oesterheld-program-equilibrium.html
Note from Caspar on 2:00:06: At least given my current interpretation of what you say here, my answer is wrong. What actually happens is that we're just back in the uncorrelated case. Basically my simulations will be a simulated repeated game in which everything is correlated _because I feed you my random sequence_ and your simulations will be a repeated game where everything is correlated. Halting works the same as usual. But of course what we end up actually playing will be uncorrelated. We discuss something like this later in the episode.
Topics we discuss, and timestamps:
0:00:44 Program equilibrium basics
0:14:20 Desiderata for program equilibria
0:24:35 Why program equilibrium matters
0:33:35 Prior work: reachable equilibria and proof-based approaches
0:53:26 The basic idea of Robust Program Equilibrium
1:07:47 Are ϵGroundedπBots inefficient?
1:15:06 Compatibility of proof-based and simulation-based program equilibria
1:18:32 Cooperating against CooperateBot, and how to avoid it
1:44:43 Making better simulation-based bots
2:01:22 Characterizing simulation-based program equilibria
2:21:24 Follow-up work
2:29:49 Following Caspar's research
Links for Caspar:
Academic website: https://www.andrew.cmu.edu/user/coesterh/
Google Scholar: https://scholar.google.com/citations?user=xeEcRjkAAAAJ&hl=en
Blog: https://casparoesterheld.com/
X / Twitter: https://x.com/c_oesterheld
Research we discuss:
Robust program equilibrium: https://link.springer.com/article/10.1007/s11238-018-9679-3
Characterising Simulation-Based Program Equilibria: https://arxiv.org/abs/2412.14570
Manifold open-source prisoner's dilemma tournament: https://manifold.markets/IsaacKing/which-240-character-program-wins-th
Results of Alex Mennen's open source prisoner's dilemma tournament: https://www.lesswrong.com/posts/QP7Ne4KXKytj4Krkx/prisoner-s-dilemma-tournament-results-0
A General Counterexample to Any Decision Theory and Some Responses: https://arxiv.org/abs/2101.00280
Cooperative and uncooperative institution designs: Surprises and problems in open-source game theory: https://arxiv.org/abs/2208.07006
Parametric Bounded Löb's Theorem and Robust Cooperation of Bounded Agents: https://arxiv.org/abs/1602.04184
A Note on the Compatibility of Different Robust Program Equilibria of the Prisoner's Dilemma: https://arxiv.org/abs/2211.05057
Episode art by Hamish Doodles: hamishdoodles.com
In this episode, Guive Assadi argues that we should give AIs property rights, so that they are integrated in our system of property and come to rely on it. The claim is that this means that AIs would not kill or steal from humans, because that would undermine the whole property system, which would be extremely valuable to them.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2026/02/15/episode-48-guive-assadi-ai-property-rights.html
Topics we discuss, and timestamps:
0:00:28 AI property rights
0:08:01 Why not steal from and kill humans
0:15:25 Why AIs may fear it could be them next
0:20:56 AI retirement
0:23:28 Could humans be upgraded to stay useful?
0:26:41 Will AI progress continue?
0:30:00 Why non-obsoletable AIs may still not end human property rights
0:38:35 Why make AIs with property rights?
0:48:01 Do property rights incentivize alignment?
0:50:09 Humans and non-human property rights
1:02:18 Humans and non-human bodily autonomy
1:16:59 Step changes in coordination ability
1:24:39 Acausal coordination
1:32:37 AI, humans, and civilizations with different technology levels
1:41:39 The case of British settlers and Tasmanians
1:47:22 Non-total expropriation
1:53:47 How Guive thinks x-risk could happen, and other loose ends
2:03:46 Following Guive's work
Guive on Substack: https://guive.substack.com/
Guive on X/Twitter: https://x.com/GuiveAssadi
Research we discuss:
The Case for AI Property Rights: https://guive.substack.com/p/the-case-for-ai-property-rights
AXRP Episode 44 - Peter Salib on AI Rights for Human Safety: https://axrp.net/episode/2025/06/28/episode-44-peter-salib-ai-rights-human-safety.html
AI Rights for Human Safety (by Salib and Goldstein): https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4913167
We don't trade with ants: https://worldspiritsockpuppet.substack.com/p/we-dont-trade-with-ants
Alignment Fine-tuning is Character Writing (on Claude as a techy philosophy SF-dwelling type): https://guive.substack.com/p/alignment-fine-tuning-is-character
Claude's charater (Anthropic post on character training): https://www.anthropic.com/research/claude-character
Git Re-Basin: Merging Models modulo Permutation Symmetries: https://arxiv.org/abs/2209.04836
The Filan Cabinet: Caspar Oesterheld on Evidential Cooperation in Large Worlds: https://thefilancabinet.com/episodes/2025/08/03/caspar-oesterheld-on-evidential-cooperation-in-large-worlds-ecl.html
Episode art by Hamish Doodles: hamishdoodles.com
When METR says something like "Claude Opus 4.5 has a 50% time horizon of 4 hours and 50 minutes", what does that mean? In this episode David Rein, METR researcher and co-author of the paper "Measuring AI ability to complete long tasks", talks about METR's work on measuring time horizons, the methodology behind those numbers, and what work remains to be done in this domain.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2026/01/03/episode-47-david-rein-metr-time-horizons.html
Topics we discuss, and timestamps:
0:00:32 Measuring AI Ability to Complete Long Tasks
0:10:54 The meaning of "task length"
0:19:27 Examples of intermediate and hard tasks
0:25:12 Why the software engineering focus
0:32:17 Why task length as difficulty measure
0:46:32 Is AI progress going superexponential?
0:50:58 Is AI progress due to increased cost to run models?
0:54:45 Why METR measures model capabilities
1:04:10 How time horizons relate to recursive self-improvement
1:12:58 Cost of estimating time horizons
1:16:23 Task realism vs mimicking important task features
1:19:50 Excursus on "Inventing Temperature"
1:25:46 Return to task realism discussion
1:33:53 Open questions on time horizons
Links for METR:
Main website: https://metr.org/
X/Twitter account: https://x.com/METR_Evals/
Research we discuss:
Measuring AI Ability to Complete Long Tasks: https://arxiv.org/abs/2503.14499
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts: https://arxiv.org/abs/2411.15114
HCAST: Human-Calibrated Autonomy Software Tasks: https://arxiv.org/abs/2503.17354
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity: https://arxiv.org/abs/2507.09089
Anthropic Economic Index: Tracking AI's role in the US and global economy: https://www.anthropic.com/research/anthropic-economic-index-september-2025-report
Bridging RL Theory and Practice with the Effective Horizon (i.e. the Cassidy Laidlaw paper): https://arxiv.org/abs/2304.09853
How Does Time Horizon Vary Across Domains?: https://metr.org/blog/2025-07-14-how-does-time-horizon-vary-across-domains/
Inventing Temperature: https://global.oup.com/academic/product/inventing-temperature-9780195337389
Is there a Half-Life for the Success Rates of AI Agents? (by Toby Ord): https://www.tobyord.com/writing/half-life
Lawrence Chan's response to the above: https://nitter.net/justanotherlaw/status/1920254586771710009
AI Task Length Horizons in Offensive Cybersecurity: https://sean-peters-au.github.io/2025/07/02/ai-task-length-horizons-in-offensive-cybersecurity.html
Episode art by Hamish Doodles: hamishdoodles.com
Could AI enable a small group to gain power over a large country, and lock in their power permanently? Often, people worried about catastrophic risks from AI have been concerned with misalignment risks. In this episode, Tom Davidson talks about a risk that could be comparably important: that of AI-enabled coups.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/08/07/episode-46-tom-davidson-ai-enabled-coups.html
Topics we discuss, and timestamps:
0:00:35 How to stage a coup without AI
0:16:17 Why AI might enable coups
0:33:29 How bad AI-enabled coups are
0:37:28 Executive coups with singularly loyal AIs
0:48:35 Executive coups with exclusive access to AI
0:54:41 Corporate AI-enabled coups
0:57:56 Secret loyalty and misalignment in corporate coups
1:11:39 Likelihood of different types of AI-enabled coups
1:25:52 How to prevent AI-enabled coups
1:33:43 Downsides of AIs loyal to the law
1:41:06 Cultural shifts vs individual action
1:45:53 Technical research to prevent AI-enabled coups
1:51:40 Non-technical research to prevent AI-enabled coups
1:58:17 Forethought
2:03:03 Following Tom's and Forethought's research
Links for Tom and Forethought:
Tom on X / Twitter: https://x.com/tomdavidsonx
Tom on LessWrong: https://www.lesswrong.com/users/tom-davidson-1
Forethought Substack: https://newsletter.forethought.org/
Will MacAskill on X / Twitter: https://x.com/willmacaskill
Will MacAskill on LessWrong: https://www.lesswrong.com/users/wdmacaskill
Research we discuss:
AI-Enabled Coups: How a Small Group Could Use AI to Seize Power: https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-power
Seizing Power: The Strategic Logic of Military Coups, by Naunihal Singh: https://muse.jhu.edu/book/31450
Experiment using AI-generated posts on Reddit draws fire for ethics concerns: https://retractionwatch.com/2025/04/28/experiment-using-ai-generated-posts-on-reddit-draws-fire-for-ethics-concerns/
Episode art by Hamish Doodles: hamishdoodles.com
In this episode, I chat with Samuel Albanie about the Google DeepMind paper he co-authored called "An Approach to Technical AGI Safety and Security". It covers the assumptions made by the approach, as well as the types of mitigations it outlines.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/07/06/episode-45-samuel-albanie-deepminds-agi-safety-approach.html
Topics we discuss, and timestamps:
0:00:37 DeepMind's Approach to Technical AGI Safety and Security
0:04:29 Current paradigm continuation
0:19:13 No human ceiling
0:21:22 Uncertain timelines
0:23:36 Approximate continuity and the potential for accelerating capability improvement
0:34:29 Misuse and misalignment
0:39:34 Societal readiness
0:43:58 Misuse mitigations
0:52:57 Misalignment mitigations
1:05:20 Samuel's thinking about technical AGI safety
1:14:02 Following Samuel's work
Samuel on Twitter/X: x.com/samuelalbanie
Research we discuss:
An Approach to Technical AGI Safety and Security: https://arxiv.org/abs/2504.01849
Levels of AGI for Operationalizing Progress on the Path to AGI: https://arxiv.org/abs/2311.02462
The Checklist: What Succeeding at AI Safety Will Involve: https://sleepinyourhat.github.io/checklist/
Measuring AI Ability to Complete Long Tasks: https://arxiv.org/abs/2503.14499
Episode art by Hamish Doodles: hamishdoodles.com
In this episode, I talk with Peter Salib about his paper "AI Rights for Human Safety", arguing that giving AIs the right to contract, hold property, and sue people will reduce the risk of their trying to attack humanity and take over. He also tells me how law reviews work, in the face of my incredulity.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/06/28/episode-44-peter-salib-ai-rights-human-safety.html
Topics we discuss, and timestamps:
0:00:40 Why AI rights
0:18:34 Why not reputation
0:27:10 Do AI rights lead to AI war?
0:36:42 Scope for human-AI trade
0:44:25 Concerns with comparative advantage
0:53:42 Proxy AI wars
0:57:56 Can companies profitably make AIs with rights?
1:09:43 Can we have AI rights and AI safety measures?
1:24:31 Liability for AIs with rights
1:38:29 Which AIs get rights?
1:43:36 AI rights and stochastic gradient descent
1:54:54 Individuating "AIs"
2:03:28 Social institutions for AI safety
2:08:20 Outer misalignment and trading with AIs
2:15:27 Why statutes of limitations should exist
2:18:39 Starting AI x-risk research in legal academia
2:24:18 How law reviews and AI conferences work
2:41:49 More on Peter moving to AI x-risk research
2:45:37 Reception of the paper
2:53:24 What publishing in law reviews does
3:04:48 Which parts of legal academia focus on AI
3:18:03 Following Peter's research
Links for Peter:
Personal website: https://www.peternsalib.com/
Writings at Lawfare: https://www.lawfaremedia.org/contributors/psalib
CLAIR: https://clair-ai.org/
Research we discuss:
AI Rights for Human Safety: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4913167
Will humans and AIs go to war? https://philpapers.org/rec/GOLWAA
Infrastructure for AI agents: https://arxiv.org/abs/2501.10114
Governing AI Agents: https://arxiv.org/abs/2501.07913
Episode art by Hamish Doodles: hamishdoodles.com
In this episode, I talk with David Lindner about Myopic Optimization with Non-myopic Approval, or MONA, which attempts to address (multi-step) reward hacking by myopically optimizing actions against a human's sense of whether those actions are generally good. Does this work? Can we get smarter-than-human AI this way? How does this compare to approaches like conservativism? Listen to find out.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/06/15/episode-43-david-lindner-mona.html
Topics we discuss, and timestamps:
0:00:29 What MONA is
0:06:33 How MONA deals with reward hacking
0:23:15 Failure cases for MONA
0:36:25 MONA's capability
0:55:40 MONA vs other approaches
1:05:03 Follow-up work
1:10:17 Other MONA test cases
1:33:47 When increasing time horizon doesn't increase capability
1:39:04 Following David's research
Links for David:
Website: https://www.davidlindner.me
Twitter / X: https://x.com/davlindner
DeepMind Medium: https://deepmindsafetyresearch.medium.com
David on the Alignment Forum: https://www.alignmentforum.org/users/david-lindner
Research we discuss:
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking: https://arxiv.org/abs/2501.13011
Arguments Against Myopic Training: https://www.alignmentforum.org/posts/GqxuDtZvfgL2bEQ5v/arguments-against-myopic-training
Episode art by Hamish Doodles: hamishdoodles.com
Earlier this year, the paper "Emergent Misalignment" made the rounds on AI x-risk social media for seemingly showing LLMs generalizing from 'misaligned' training data of insecure code to acting comically evil in response to innocuous questions. In this episode, I chat with one of the authors of that paper, Owain Evans, about that research as well as other work he's done to understand the psychology of large language models.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/06/06/episode-42-owain-evans-llm-psychology.html
Topics we discuss, and timestamps:
0:00:37 Why introspection?
0:06:24 Experiments in "Looking Inward"
0:15:11 Why fine-tune for introspection?
0:22:32 Does "Looking Inward" test introspection, or something else?
0:34:14 Interpreting the results of "Looking Inward"
0:44:56 Limitations to introspection?
0:49:54 "Tell me about yourself", and its relation to other papers
1:05:45 Backdoor results
1:12:01 Emergent Misalignment
1:22:13 Why so hammy, and so infrequently evil?
1:36:31 Why emergent misalignment?
1:46:45 Emergent misalignment and other types of misalignment
1:53:57 Is emergent misalignment good news?
2:00:01 Follow-up work to "Emergent Misalignment"
2:03:10 Reception of "Emergent Misalignment" vs other papers
2:07:43 Evil numbers
2:12:20 Following Owain's research
Links for Owain:
Truthful AI: https://www.truthfulai.org
Owain's website: https://owainevans.github.io/
Owain's twitter/X account: https://twitter.com/OwainEvans_UK
Research we discuss:
Looking Inward: Language Models Can Learn About Themselves by Introspection: https://arxiv.org/abs/2410.13787
Tell me about yourself: LLMs are aware of their learned behaviors: https://arxiv.org/abs/2501.11120
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data: https://arxiv.org/abs/2406.14546
Emergent Misalignment: Narrow fine-tuning can produce broadly misaligned LLMs: https://arxiv.org/abs/2502.17424
X/Twitter thread of GPT-4.1 emergent misalignment results: https://x.com/OwainEvans_UK/status/1912701650051190852
Taken out of context: On measuring situational awareness in LLMs: https://arxiv.org/abs/2309.00667
Episode art by Hamish Doodles: hamishdoodles.com
What's the next step forward in interpretability? In this episode, I chat with Lee Sharkey about his proposal for detecting computational mechanisms within neural networks: Attribution-based Parameter Decomposition, or APD for short.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/06/03/episode-41-lee-sharkey-attribution-based-parameter-decomposition.html
Topics we discuss, and timestamps:
0:00:41 APD basics
0:07:57 Faithfulness
0:11:10 Minimality
0:28:44 Simplicity
0:34:50 Concrete-ish examples of APD
0:52:00 Which parts of APD are canonical
0:58:10 Hyperparameter selection
1:06:40 APD in toy models of superposition
1:14:40 APD and compressed computation
1:25:43 Mechanisms vs representations
1:34:41 Future applications of APD?
1:44:19 How costly is APD?
1:49:14 More on minimality training
1:51:49 Follow-up work
2:05:24 APD on giant chain-of-thought models?
2:11:27 APD and "features"
2:14:11 Following Lee's work
Lee links (Leenks):
X/Twitter: https://twitter.com/leedsharkey
Alignment Forum: https://www.alignmentforum.org/users/lee_sharkey
Research we discuss:
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-Based Parameter Decomposition: https://arxiv.org/abs/2501.14926
Toy Models of Superposition: https://transformer-circuits.pub/2022/toy_model/index.html
Towards a unified and verified understanding of group-operation networks: https://arxiv.org/abs/2410.07476
Feature geometry is outside the superposition hypothesis: https://www.alignmentforum.org/posts/MFBTjb2qf3ziWmzz6/sae-feature-geometry-is-outside-the-superposition-hypothesis
Episode art by Hamish Doodles: hamishdoodles.com
How do we figure out whether interpretability is doing its job? One way is to see if it helps us prove things about models that we care about knowing. In this episode, I speak with Jason Gross about his agenda to benchmark interpretability in this way, and his exploration of the intersection of proofs and modern machine learning.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/03/28/episode-40-jason-gross-compact-proofs-interpretability.html
Topics we discuss, and timestamps:
0:00:40 - Why compact proofs
0:07:25 - Compact Proofs of Model Performance via Mechanistic Interpretability
0:14:19 - What compact proofs look like
0:32:43 - Structureless noise, and why proofs
0:48:23 - What we've learned about compact proofs in general
0:59:02 - Generalizing 'symmetry'
1:11:24 - Grading mechanistic interpretability
1:43:34 - What helps compact proofs
1:51:08 - The limits of compact proofs
2:07:33 - Guaranteed safe AI, and AI for guaranteed safety
2:27:44 - Jason and Rajashree's start-up
2:34:19 - Following Jason's work
Links to Jason:
Github: https://github.com/jasongross
Website: https://jasongross.github.io
Alignment Forum: https://www.alignmentforum.org/users/jason-gross
Links to work we discuss:
Compact Proofs of Model Performance via Mechanistic Interpretability: https://arxiv.org/abs/2406.11779
Unifying and Verifying Mechanistic Interpretability: A Case Study with Group Operations: https://arxiv.org/abs/2410.07476
Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration: https://arxiv.org/abs/2412.03773
Stage-Wise Model Diffing: https://transformer-circuits.pub/2024/model-diffing/index.html
Causal Scrubbing: a method for rigorously testing interpretability hypotheses: https://www.lesswrong.com/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition (aka the Apollo paper on APD): https://arxiv.org/abs/2501.14926
Towards Guaranteed Safe AI: https://www2.eecs.berkeley.edu/Pubs/TechRpts/2024/EECS-2024-45.pdf
Episode art by Hamish Doodles: hamishdoodles.com
In this episode, I chat with David Duvenaud about two topics he's been thinking about: firstly, a paper he wrote about evaluating whether or not frontier models can sabotage human decision-making or monitoring of the same models; and secondly, the difficult situation humans find themselves in in a post-AGI future, even if AI is aligned with human intentions.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/03/01/episode-38_8-david-duvenaud-sabotage-evaluations-post-agi-future.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: @FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
01:42 - The difficulty of sabotage evaluations
05:23 - Types of sabotage evaluation
08:45 - The state of sabotage evaluations
12:26 - What happens after AGI?
Links:
Sabotage Evaluations for Frontier Models: https://arxiv.org/abs/2410.21514
Gradual Disempowerment: https://gradual-disempowerment.ai/
Episode art by Hamish Doodles: hamishdoodles.com
The Future of Life Institute is one of the oldest and most prominant organizations in the AI existential safety space, working on such topics as the AI pause open letter and how the EU AI Act can be improved. Metaculus is one of the premier forecasting sites on the internet. Behind both of them lie one man: Anthony Aguirre, who I talk with in this episode.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/02/09/episode-38_7-anthony-aguirre-future-of-life-institute.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
00:33 - Anthony, FLI, and Metaculus
06:46 - The Alignment Workshop
07:15 - FLI's current activity
11:04 - AI policy
17:09 - Work FLI funds
Links:
Future of Life Institute: https://futureoflife.org/
Metaculus: https://www.metaculus.com/
Future of Life Foundation: https://www.flf.org/
Episode art by Hamish Doodles: hamishdoodles.com
Typically this podcast talks about how to avert destruction from AI. But what would it take to ensure AI promotes human flourishing as well as it can? Is alignment to individuals enough, and if not, where do we go form here? In this episode, I talk with Joel Lehman about these questions.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/01/24/episode-38_6-joel-lehman-positive-visions-of-ai.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
01:12 - Why aligned AI might not be enough
04:05 - Positive visions of AI
08:27 - Improving recommendation systems
Links:
Why Greatness Cannot Be Planned: https://www.amazon.com/Why-Greatness-Cannot-Planned-Objective/dp/3319155237
We Need Positive Visions of AI Grounded in Wellbeing: https://thegradientpub.substack.com/p/beneficial-ai-wellbeing-lehman-ngo
Machine Love: https://arxiv.org/abs/2302.09248
AI Alignment with Changing and Influenceable Reward Functions: https://arxiv.org/abs/2405.17713
Episode art by Hamish Doodles: hamishdoodles.com
Suppose we're worried about AIs engaging in long-term plans that they don't tell us about. If we were to peek inside their brains, what should we look for to check whether this was happening? In this episode Adrià Garriga-Alonso talks about his work trying to answer this question.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Transcript: https://axrp.net/episode/2025/01/20/episode-38_5-adria-garriga-alonso-detecting-ai-scheming.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
01:04 - The Alignment Workshop
02:49 - How to detect scheming AIs
05:29 - Sokoban-solving networks taking time to think
12:18 - Model organisms of long-term planning
19:44 - How and why to study planning in networks
Links:
Adrià's website: https://agarri.ga/
An investigation of model-free planning: https://arxiv.org/abs/1901.03559
Model-Free Planning: https://tuphs28.github.io/projects/interpplanning/
Planning in a recurrent neural network that plays Sokoban: https://arxiv.org/abs/2407.15421
Episode art by Hamish Doodles: hamishdoodles.com
AI researchers often complain about the poor coverage of their work in the news media. But why is this happening, and how can it be fixed? In this episode, I speak with Shakeel Hashim about the resource constraints facing AI journalism, the disconnect between journalists' and AI researchers' views on transformative AI, and efforts to improve the state of AI journalism, such as Tarbell and Shakeel's newsletter, Transformer.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2025/01/05/episode-38_4-shakeel-hashim-ai-journalism.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
01:31 - The AI media ecosystem
02:34 - Why not more AI news?
07:18 - Disconnects between journalists and the AI field
12:42 - Tarbell
18:44 - The Transformer newsletter
Links:
Transformer (Shakeel's substack): https://www.transformernews.ai/
Tarbell: https://www.tarbellfellowship.org/
Episode art by Hamish Doodles: hamishdoodles.com
Lots of people in the AI safety space worry about models being able to make deliberate, multi-step plans. But can we already see this in existing neural nets? In this episode, I talk with Erik Jenner about his work looking at internal look-ahead within chess-playing neural networks.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/12/12/episode-38_3-erik-jenner-learned-look-ahead.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
00:57 - How chess neural nets look into the future
04:29 - The dataset and basic methodology
05:23 - Testing for branching futures?
07:57 - Which experiments demonstrate what
10:43 - How the ablation experiments work
12:38 - Effect sizes
15:23 - X-risk relevance
18:08 - Follow-up work
21:29 - How much planning does the network do?
Research we mention:
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network: https://arxiv.org/abs/2406.00877
Understanding the learned look-ahead behavior of chess neural networks (a development of the follow-up research Erik mentioned): https://openreview.net/forum?id=Tl8EzmgsEp
Linear Latent World Models in Simple Transformers: A Case Study on Othello-GPT: https://arxiv.org/abs/2310.07582
Episode art by Hamish Doodles: hamishdoodles.com
The 'model organisms of misalignment' line of research creates AI models that exhibit various types of misalignment, and studies them to try to understand how the misalignment occurs and whether it can be somehow removed. In this episode, Evan Hubinger talks about two papers he's worked on at Anthropic under this agenda: "Sleeper Agents" and "Sycophancy to Subterfuge".
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/12/01/episode-39-evan-hubinger-model-organisms-misalignment.html
Topics we discuss, and timestamps:
0:00:36 - Model organisms and stress-testing
0:07:38 - Sleeper Agents
0:22:32 - Do 'sleeper agents' properly model deceptive alignment?
0:38:32 - Surprising results in "Sleeper Agents"
0:57:25 - Sycophancy to Subterfuge
1:09:21 - How models generalize from sycophancy to subterfuge
1:16:37 - Is the reward editing task valid?
1:21:46 - Training away sycophancy and subterfuge
1:29:22 - Model organisms, AI control, and evaluations
1:33:45 - Other model organisms research
1:35:27 - Alignment stress-testing at Anthropic
1:43:32 - Following Evan's work
Main papers:
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training: https://arxiv.org/abs/2401.05566
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models: https://arxiv.org/abs/2406.10162
Anthropic links:
Anthropic's newsroom: https://www.anthropic.com/news
Careers at Anthropic: https://www.anthropic.com/careers
Other links:
Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research: https://www.alignmentforum.org/posts/ChDH335ckdvpxXaXX/model-organisms-of-misalignment-the-case-for-a-new-pillar-of-1
Simple probes can catch sleeper agents: https://www.anthropic.com/research/probes-catch-sleeper-agents
Studying Large Language Model Generalization with Influence Functions: https://arxiv.org/abs/2308.03296
Stress-Testing Capability Elicitation With Password-Locked Models [aka model organisms of sandbagging]: https://arxiv.org/abs/2405.19550
Episode art by Hamish Doodles: hamishdoodles.com
You may have heard of singular learning theory, and its "local learning coefficient", or LLC - but have you heard of the refined LLC? In this episode, I chat with Jesse Hoogland about his work on SLT, and using the refined LLC to find a new circuit in language models.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/11/27/38_2-jesse-hoogland-singular-learning-theory.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
00:34 - About Jesse
01:49 - The Alignment Workshop
02:31 - About Timaeus
05:25 - SLT that isn't developmental interpretability
10:41 - The refined local learning coefficient
14:06 - Finding the multigram circuit
Links:
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient: https://arxiv.org/abs/2410.02984
Investigating the learning coefficient of modular addition: hackathon project: https://www.lesswrong.com/posts/4v3hMuKfsGatLXPgt/investigating-the-learning-coefficient-of-modular-addition
Episode art by Hamish Doodles: hamishdoodles.com
Road lines, street lights, and licence plates are examples of infrastructure used to ensure that roads operate smoothly. In this episode, Alan Chan talks about using similar interventions to help avoid bad outcomes from the deployment of AI agents.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/11/16/episode-38_1-alan-chan-agent-infrastructure.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
The Alignment Workshop: https://www.alignment-workshop.com/
Topics we discuss, and timestamps:
01:02 - How the Alignment Workshop is
01:32 - Agent infrastructure
04:57 - Why agent infrastructure
07:54 - A trichotomy of agent infrastructure
13:59 - Agent IDs
18:17 - Agent channels
20:29 - Relation to AI control
Links:
Alan on Google Scholar: https://scholar.google.com/citations?user=lmQmYPgAAAAJ&hl=en&oi=ao
IDs for AI Systems: https://arxiv.org/abs/2406.12137
Visibility into AI Agents: https://arxiv.org/abs/2401.13138
Episode art by Hamish Doodles: hamishdoodles.com
Do language models understand the causal structure of the world, or do they merely note correlations? And what happens when you build a big AI society out of them? In this brief episode, recorded at the Bay Area Alignment Workshop, I chat with Zhijing Jin about her research on these questions.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/11/14/episode-38_0-zhijing-jin-llms-causality-multi-agent-systems.html
FAR.AI: https://far.ai/
FAR.AI on X (aka Twitter): https://x.com/farairesearch
FAR.AI on YouTube: https://www.youtube.com/@FARAIResearch
Topics we discuss, and timestamps:
00:35 - How the Alignment Workshop is
00:47 - How Zhijing got interested in causality and natural language processing
03:14 - Causality and alignment
06:21 - Causality without randomness
10:07 - Causal abstraction
11:42 - Why LLM causal reasoning?
13:20 - Understanding LLM causal reasoning
16:33 - Multi-agent systems
Links:
Zhijing's website: https://zhijing-jin.com/fantasy/
Zhijing on X (aka Twitter): https://x.com/zhijingjin
Can Large Language Models Infer Causation from Correlation?: https://arxiv.org/abs/2306.05836
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents: https://arxiv.org/abs/2404.16698
Episode art by Hamish Doodles: hamishdoodles.com
Epoch AI is the premier organization that tracks the trajectory of AI - how much compute is used, the role of algorithmic improvements, the growth in data used, and when the above trends might hit an end. In this episode, I speak with the director of Epoch AI, Jaime Sevilla, about how compute, data, and algorithmic improvements are impacting AI, and whether continuing to scale can get us AGI.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/10/04/episode-37-jaime-sevilla-forecasting-ai.html
Topics we discuss, and timestamps:
0:00:38 - The pace of AI progress
0:07:49 - How Epoch AI tracks AI compute
0:11:44 - Why does AI compute grow so smoothly?
0:21:46 - When will we run out of computers?
0:38:56 - Algorithmic improvement
0:44:21 - Algorithmic improvement and scaling laws
0:56:56 - Training data
1:04:56 - Can scaling produce AGI?
1:16:55 - When will AGI arrive?
1:21:20 - Epoch AI
1:27:06 - Open questions in AI forecasting
1:35:21 - Epoch AI and x-risk
1:41:34 - Following Epoch AI's research
Links for Jaime and Epoch AI:
Epoch AI: https://epochai.org/
Machine Learning Trends dashboard: https://epochai.org/trends
Epoch AI on X / Twitter: https://x.com/EpochAIResearch
Jaime on X / Twitter: https://x.com/Jsevillamol
Research we discuss:
Training Compute of Frontier AI Models Grows by 4-5x per Year: https://epochai.org/blog/training-compute-of-frontier-ai-models-grows-by-4-5x-per-year
Optimally Allocating Compute Between Inference and Training: https://epochai.org/blog/optimally-allocating-compute-between-inference-and-training
Algorithmic Progress in Language Models [blog post]: https://epochai.org/blog/algorithmic-progress-in-language-models
Algorithmic progress in language models [paper]: https://arxiv.org/abs/2403.05812
Training Compute-Optimal Large Language Models [aka the Chinchilla scaling law paper]: https://arxiv.org/abs/2203.15556
Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data [blog post]: https://epochai.org/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data
Will we run out of data? Limits of LLM scaling based on human-generated data [paper]: https://arxiv.org/abs/2211.04325
The Direct Approach: https://epochai.org/blog/the-direct-approach
Episode art by Hamish Doodles: hamishdoodles.com
Sometimes, people talk about transformers as having "world models" as a result of being trained to predict text data on the internet. But what does this even mean? In this episode, I talk with Adam Shai and Paul Riechers about their work applying computational mechanics, a sub-field of physics studying how to predict random processes, to neural networks.
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/09/29/episode-36-adam-shai-paul-riechers-computational-mechanics.html
Topics we discuss, and timestamps:
0:00:42 - What computational mechanics is
0:29:49 - Computational mechanics vs other approaches
0:36:16 - What world models are
0:48:41 - Fractals
0:57:43 - How the fractals are formed
1:09:55 - Scaling computational mechanics for transformers
1:21:52 - How Adam and Paul found computational mechanics
1:36:16 - Computational mechanics for AI safety
1:46:05 - Following Adam and Paul's research
Simplex AI Safety: https://www.simplexaisafety.com/
Research we discuss:
Transformers represent belief state geometry in their residual stream: https://arxiv.org/abs/2405.15943
Transformers represent belief state geometry in their residual stream [LessWrong post]: https://www.lesswrong.com/posts/gTZ2SxesbHckJ3CkF/transformers-represent-belief-state-geometry-in-their
Why Would Belief-States Have A Fractal Structure, And Why Would That Matter For Interpretability? An Explainer: https://www.lesswrong.com/posts/mBw7nc4ipdyeeEpWs/why-would-belief-states-have-a-fractal-structure-and-why
Episode art by Hamish Doodles: hamishdoodles.com
Patreon: https://www.patreon.com/axrpodcast
MATS: https://www.matsprogram.org
Note: I'm employed by MATS, but they're not paying me to make this video.
How do we figure out what large language models believe? In fact, do they even have beliefs? Do those beliefs have locations, and if so, can we edit those locations to change the beliefs? Also, how are we going to get AI to perform tasks so hard that we can't figure out if they succeeded at them? In this episode, I chat to Peter Hase about his research into these questions.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/08/24/episode-35-peter-hase-llm-beliefs-easy-to-hard-generalization.html
Topics we discuss, and timestamps:
0:00:36 - NLP and interpretability
0:10:20 - Interpretability lessons
0:32:22 - Belief interpretability
1:00:12 - Localizing and editing models' beliefs
1:19:18 - Beliefs beyond language models
1:27:21 - Easy-to-hard generalization
1:47:16 - What do easy-to-hard results tell us?
1:57:33 - Easy-to-hard vs weak-to-strong
2:03:50 - Different notions of hardness
2:13:01 - Easy-to-hard vs weak-to-strong, round 2
2:15:39 - Following Peter's work
Peter on Twitter: https://x.com/peterbhase
Peter's papers:
Foundational Challenges in Assuring Alignment and Safety of Large Language Models: https://arxiv.org/abs/2404.09932
Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs: https://arxiv.org/abs/2111.13654
Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models: https://arxiv.org/abs/2301.04213
Are Language Models Rational? The Case of Coherence Norms and Belief Revision: https://arxiv.org/abs/2406.03442
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks: https://arxiv.org/abs/2401.06751
Other links:
Toy Models of Superposition: https://transformer-circuits.pub/2022/toy_model/index.html
Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV): https://arxiv.org/abs/1711.11279
Locating and Editing Factual Associations in GPT (aka the ROME paper): https://arxiv.org/abs/2202.05262
Of nonlinearity and commutativity in BERT: https://arxiv.org/abs/2101.04547
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model: https://arxiv.org/abs/2306.03341
Editing a classifier by rewriting its prediction rules: https://arxiv.org/abs/2112.01008
Discovering Latent Knowledge Without Supervision (aka the Collin Burns CCS paper): https://arxiv.org/abs/2212.03827
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision: https://arxiv.org/abs/2312.09390
Concrete problems in AI safety: https://arxiv.org/abs/1606.06565
Rissanen Data Analysis: Examining Dataset Characteristics via Description Length: https://arxiv.org/abs/2103.03872
Episode art by Hamish Doodles: hamishdoodles.com
How can we figure out if AIs are capable enough to pose a threat to humans? When should we make a big effort to mitigate risks of catastrophic AI misbehaviour? In this episode, I chat with Beth Barnes, founder of and head of research at METR, about these questions and more.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/07/28/episode-34-ai-evaluations-beth-barnes.html
Topics we discuss, and timestamps:
0:00:37 - What is METR?
0:02:44 - What is an "eval"?
0:14:42 - How good are evals?
0:37:25 - Are models showing their full capabilities?
0:53:25 - Evaluating alignment
1:01:38 - Existential safety methodology
1:12:13 - Threat models and capability buffers
1:38:25 - METR's policy work
1:48:19 - METR's relationships with labs
2:04:12 - Related research
2:10:02 - Roles at METR, and following METR's work
Links for METR:
METR: https://metr.org
METR Task Development Guide - Bounty: https://taskdev.metr.org/bounty/
METR - Hiring: https://metr.org/hiring
Autonomy evaluation resources: https://metr.org/blog/2024-03-13-autonomy-evaluation-resources/
Other links:
Update on ARC's recent eval efforts (contains GPT-4 taskrabbit captcha story) https://metr.org/blog/2023-03-18-update-on-recent-evals/
Password-locked models: a stress case for capabilities evaluation: https://www.alignmentforum.org/posts/rZs6ddqNnW8LXuJqA/password-locked-models-a-stress-case-for-capabilities
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training: https://arxiv.org/abs/2401.05566
Untrusted smart models and trusted dumb models: https://www.alignmentforum.org/posts/LhxHcASQwpNa3mRNk/untrusted-smart-models-and-trusted-dumb-models
AI companies aren't really using external evaluators: https://www.lesswrong.com/posts/WjtnvndbsHxCnFNyc/ai-companies-aren-t-really-using-external-evaluators
Nobody Knows How to Safety-Test AI (Time): https://time.com/6958868/artificial-intelligence-safety-evaluations-risks/
ChatGPT can talk, but OpenAI employees sure can’t: https://www.vox.com/future-perfect/2024/5/17/24158478/openai-departures-sam-altman-employees-chatgpt-release
Leaked OpenAI documents reveal aggressive tactics toward former employees: https://www.vox.com/future-perfect/351132/openai-vested-equity-nda-sam-altman-documents-employees
Beth on her non-disparagement agreement with OpenAI: https://www.lesswrong.com/posts/yRWv5kkDD4YhzwRLq/non-disparagement-canaries-for-openai?commentId=MrJF3tWiKYMtJepgX
Sam Altman's statement on OpenAI equity: https://x.com/sama/status/1791936857594581428
Episode art by Hamish Doodles: hamishdoodles.com
Reinforcement Learning from Human Feedback, or RLHF, is one of the main ways that makers of large language models make them 'aligned'. But people have long noted that there are difficulties with this approach when the models are smarter than the humans providing feedback. In this episode, I talk with Scott Emmons about his work categorizing the problems that can show up in this setting.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
The transcript: https://axrp.net/episode/2024/06/12/episode-33-rlhf-problems-scott-emmons.html
Topics we discuss, and timestamps:
0:00:33 - Deceptive inflation
0:17:56 - Overjustification
0:32:48 - Bounded human rationality
0:50:46 - Avoiding these problems
1:14:13 - Dimensional analysis
1:23:32 - RLHF problems, in theory and practice
1:31:29 - Scott's research program
1:39:42 - Following Scott's research
Scott's website: https://www.scottemmons.com
Scott's X/twitter account: https://x.com/emmons_scott
When Your AIs Deceive You: Challenges With Partial Observability of Human Evaluators in Reward Learning: https://arxiv.org/abs/2402.17747
Other works we discuss:
AI Deception: A Survey of Examples, Risks, and Potential Solutions: https://arxiv.org/abs/2308.14752
Uncertain decisions facilitate better preference learning: https://arxiv.org/abs/2106.10394
Invariance in Policy Optimisation and Partial Identifiability in Reward Learning: https://arxiv.org/abs/2203.07475
The Humble Gaussian Distribution (aka principal component analysis and dimensional analysis): http://www.inference.org.uk/mackay/humble.pdf
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!: https://arxiv.org/abs/2310.03693
Episode art by Hamish Doodles: hamishdoodles.com
What's the difference between a large language model and the human brain? And what's wrong with our theories of agency? In this episode, I chat about these questions with Jan Kulveit, who leads the Alignment of Complex Systems research group.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
The transcript: axrp.net/episode/2024/05/30/episode-32-understanding-agency-jan-kulveit.html
Topics we discuss, and timestamps:
0:00:47 - What is active inference?
0:15:14 - Preferences in active inference
0:31:33 - Action vs perception in active inference
0:46:07 - Feedback loops
1:01:32 - Active inference vs LLMs
1:12:04 - Hierarchical agency
1:58:28 - The Alignment of Complex Systems group
Website of the Alignment of Complex Systems group (ACS): acsresearch.org
ACS on X/Twitter: x.com/acsresearchorg
Jan on LessWrong: lesswrong.com/users/jan-kulveit
Predictive Minds: Large Language Models as Atypical Active Inference Agents: arxiv.org/abs/2311.10215
Other works we discuss:
Active Inference: The Free Energy Principle in Mind, Brain, and Behavior: https://www.goodreads.com/en/book/show/58275959
Book Review: Surfing Uncertainty: https://slatestarcodex.com/2017/09/05/book-review-surfing-uncertainty/
The self-unalignment problem: https://www.lesswrong.com/posts/9GyniEBaN3YYTqZXn/the-self-unalignment-problem
Mitigating generative agent social dilemmas (aka language models writing contracts for Minecraft): https://social-dilemmas.github.io/
Episode art by Hamish Doodles: hamishdoodles.com
What's going on with deep learning? What sorts of models get learned, and what are the learning dynamics? Singular learning theory is a theory of Bayesian statistics broad enough in scope to encompass deep neural networks that may help answer these questions. In this episode, I speak with Daniel Murfet about this research program and what it tells us.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
0:00:26 - What is singular learning theory?
0:16:00 - Phase transitions
0:35:12 - Estimating the local learning coefficient
0:44:37 - Singular learning theory and generalization
1:00:39 - Singular learning theory vs other deep learning theory
1:17:06 - How singular learning theory hit AI alignment
1:33:12 - Payoffs of singular learning theory for AI alignment
1:59:36 - Does singular learning theory advance AI capabilities?
2:13:02 - Open problems in singular learning theory for AI alignment
2:20:53 - What is the singular fluctuation?
2:25:33 - How geometry relates to information
2:30:13 - Following Daniel Murfet's work
The transcript: https://axrp.net/episode/2024/05/07/episode-31-singular-learning-theory-dan-murfet.html
Daniel Murfet's twitter/X account: https://twitter.com/danielmurfet
Developmental interpretability website: https://devinterp.com
Developmental interpretability YouTube channel: https://www.youtube.com/@Devinterp
Main research discussed in this episode:
Developmental Landscape of In-Context Learning: https://arxiv.org/abs/2402.02364
Estimating the Local Learning Coefficient at Scale: https://arxiv.org/abs/2402.03698
Simple versus Short: Higher-order degeneracy and error-correction: https://www.lesswrong.com/posts/nWRj6Ey8e5siAEXbK/simple-versus-short-higher-order-degeneracy-and-error-1
Other links:
Algebraic Geometry and Statistical Learning Theory (the grey book): https://www.cambridge.org/core/books/algebraic-geometry-and-statistical-learning-theory/9C8FD1BDC817E2FC79117C7F41544A3A
Mathematical Theory of Bayesian Statistics (the green book): https://www.routledge.com/Mathematical-Theory-of-Bayesian-Statistics/Watanabe/p/book/9780367734817 In-context learning and induction heads: https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity: https://arxiv.org/abs/2106.15933
A mathematical theory of semantic development in deep neural networks: https://www.pnas.org/doi/abs/10.1073/pnas.1820226116
Consideration on the Learning Efficiency Of Multiple-Layered Neural Networks with Linear Units: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4404877
Neural Tangent Kernel: Convergence and Generalization in Neural Networks: https://arxiv.org/abs/1806.07572
The Interpolating Information Criterion for Overparameterized Models: https://arxiv.org/abs/2307.07785
Feature Learning in Infinite-Width Neural Networks: https://arxiv.org/abs/2011.14522
A central AI alignment problem: capabilities generalization, and the sharp left turn: https://www.lesswrong.com/posts/GNhMPAWcfBCASy8e6/a-central-ai-alignment-problem-capabilities-generalization
Quantifying degeneracy in singular models via the learning coefficient: https://arxiv.org/abs/2308.12108
Episode art by Hamish Doodles: hamishdoodles.com
Top labs use various forms of "safety training" on models before their release to make sure they don't do nasty stuff - but how robust is that? How can we ensure that the weights of powerful AIs don't get leaked or stolen? And what can AI even do these days? In this episode, I speak with Jeffrey Ladish about security and AI.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
0:00:38 - Fine-tuning away safety training
0:13:50 - Dangers of open LLMs vs internet search
0:19:52 - What we learn by undoing safety filters
0:27:34 - What can you do with jailbroken AI?
0:35:28 - Security of AI model weights
0:49:21 - Securing against attackers vs AI exfiltration
1:08:43 - The state of computer security
1:23:08 - How AI labs could be more secure
1:33:13 - What does Palisade do?
1:44:40 - AI phishing
1:53:32 - More on Palisade's work
1:59:56 - Red lines in AI development
2:09:56 - Making AI legible
2:14:08 - Following Jeffrey's research
The transcript: axrp.net/episode/2024/04/30/episode-30-ai-security-jeffrey-ladish.html
Palisade Research: palisaderesearch.org
Jeffrey's Twitter/X account: twitter.com/JeffLadish
Main papers we discussed:
LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B: arxiv.org/abs/2310.20624
BadLLaMa: Cheaply Removing Safety Fine-tuning From LLaMa 2-Chat 13B: arxiv.org/abs/2311.00117
Securing Artificial Intelligence Model Weights: rand.org/pubs/working_papers/WRA2849-1.html
Other links:
Llama 2: Open Foundation and Fine-Tuned Chat Models: https://arxiv.org/abs/2307.09288
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!: https://arxiv.org/abs/2310.03693
Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models: https://arxiv.org/abs/2310.02949
On the Societal Impact of Open Foundation Models (Stanford paper on marginal harms from open-weight models): https://crfm.stanford.edu/open-fms/
The Operational Risks of AI in Large-Scale Biological Attacks (RAND): https://www.rand.org/pubs/research_reports/RRA2977-2.html
Preventing model exfiltration with upload limits: https://www.alignmentforum.org/posts/rf66R4YsrCHgWx9RG/preventing-model-exfiltration-with-upload-limits
A deep dive into an NSO zero-click iMessage exploit: Remote Code Execution: https://googleprojectzero.blogspot.com/2021/12/a-deep-dive-into-nso-zero-click.html
In-browser transformer inference: https://aiserv.cloud/
Anatomy of a rental phishing scam: https://jeffreyladish.com/anatomy-of-a-rental-phishing-scam/
Causal Scrubbing: a method for rigorously testing interpretability hypotheses: https://www.alignmentforum.org/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing
Episode art by Hamish Doodles: hamishdoodles.com
In 2022, it was announced that a fairly simple method can be used to extract the true beliefs of a language model on any given topic, without having to actually understand the topic at hand. Earlier, in 2021, it was announced that neural networks sometimes 'grok': that is, when training them on certain tasks, they initially memorize their training data (achieving their training goal in a way that doesn't generalize), but then suddenly switch to understanding the 'real' solution in a way that generalizes. What's going on with these discoveries? Are they all they're cracked up to be, and if so, how are they working? In this episode, I talk to Vikrant Varma about his research getting to the bottom of these questions.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
0:00:36 - Challenges with unsupervised LLM knowledge discovery, aka contra CCS
0:00:36 - What is CCS?
0:09:54 - Consistent and contrastive features other than model beliefs
0:20:34 - Understanding the banana/shed mystery
0:41:59 - Future CCS-like approaches
0:53:29 - CCS as principal component analysis
0:56:21 - Explaining grokking through circuit efficiency
0:57:44 - Why research science of deep learning?
1:12:07 - Summary of the paper's hypothesis
1:14:05 - What are 'circuits'?
1:20:48 - The role of complexity
1:24:07 - Many kinds of circuits
1:28:10 - How circuits are learned
1:38:24 - Semi-grokking and ungrokking
1:50:53 - Generalizing the results
1:58:51 - Vikrant's research approach
2:06:36 - The DeepMind alignment team
2:09:06 - Follow-up work
The transcript: axrp.net/episode/2024/04/25/episode-29-science-of-deep-learning-vikrant-varma.html
Vikrant's Twitter/X account: twitter.com/vikrantvarma_
Main papers:
Challenges with unsupervised LLM knowledge discovery: arxiv.org/abs/2312.10029
Explaining grokking through circuit efficiency: arxiv.org/abs/2309.02390
Other works discussed:
Discovering latent knowledge in language models without supervision (CCS): arxiv.org/abs/2212.03827
Eliciting Latent Knowledge: How to Tell if your Eyes Deceive You: https://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/edit
Discussion: Challenges with unsupervised LLM knowledge discovery: lesswrong.com/posts/wtfvbsYjNHYYBmT3k/discussion-challenges-with-unsupervised-llm-knowledge-1
Comment thread on the banana/shed results: lesswrong.com/posts/wtfvbsYjNHYYBmT3k/discussion-challenges-with-unsupervised-llm-knowledge-1?commentId=hPZfgA3BdXieNfFuY
Fabien Roger, What discovering latent knowledge did and did not find: lesswrong.com/posts/bWxNPMy5MhPnQTzKz/what-discovering-latent-knowledge-did-and-did-not-find-4
Scott Emmons, Contrast Pairs Drive the Performance of Contrast Consistent Search (CCS): lesswrong.com/posts/9vwekjD6xyuePX7Zr/contrast-pairs-drive-the-empirical-performance-of-contrast
Grokking: Generalizing Beyond Overfitting on Small Algorithmic Datasets: arxiv.org/abs/2201.02177
Keeping Neural Networks Simple by Minimizing the Minimum Description Length of the Weights (Hinton 1993 L2): dl.acm.org/doi/pdf/10.1145/168304.168306
Progress measures for grokking via mechanistic interpretability: arxiv.org/abs/2301.0521
Episode art by Hamish Doodles: hamishdoodles.com
How should the law govern AI? Those concerned about existential risks often push either for bans or for regulations meant to ensure that AI is developed safely - but another approach is possible. In this episode, Gabriel Weil talks about his proposal to modify tort law to enable people to sue AI companies for disasters that are "nearly catastrophic".
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
0:00:35 - The basic idea
0:20:36 - Tort law vs regulation
0:29:10 - Weil's proposal vs Hanson's proposal
0:37:00 - Tort law vs Pigouvian taxation
0:41:16 - Does disagreement on AI risk make this proposal less effective?
0:49:53 - Warning shots - their prevalence and character
0:59:17 - Feasibility of big changes to liability law
1:29:17 - Interactions with other areas of law
1:38:59 - How Gabriel encountered the AI x-risk field
1:42:41 - AI x-risk and the legal field
1:47:44 - Technical research to help with this proposal
1:50:47 - Decisions this proposal could influence
1:55:34 - Following Gabriel's research
The transcript: axrp.net/episode/2024/04/17/episode-28-tort-law-for-ai-risk-gabriel-weil.html
Links for Gabriel:
SSRN page: papers.ssrn.com/sol3/cf_dev/AbsByAuth.cfm?per_id=1648032
Twitter/X account: twitter.com/gabriel_weil
Tort Law as a Tool for Mitigating Catastrophic Risk from Artificial Intelligence: papers.ssrn.com/sol3/papers.cfm?abstract_id=4694006
Other links:
Foom liability: overcomingbias.com/p/foom-liability
Punitive Damages: An Economic Analysis: law.harvard.edu/faculty/shavell/pdf/111_Harvard_Law_Rev_869.pdf
Efficiency, Fairness, and the Externalization of Reasonable Risks: The Problem With the Learned Hand Formula: papers.ssrn.com/sol3/papers.cfm?abstract_id=4466197
Tort Law Can Play an Important Role in Mitigating AI Risk: forum.effectivealtruism.org/posts/epKBmiyLpZWWFEYDb/tort-law-can-play-an-important-role-in-mitigating-ai-risk
How Technical AI Safety Researchers Can Help Implement Punitive Damages to Mitigate Catastrophic AI Risk: forum.effectivealtruism.org/posts/yWKaBdBygecE42hFZ/how-technical-ai-safety-researchers-can-help-implement
Can the courts save us from dangerous AI? [Vox]: vox.com/future-perfect/2024/2/7/24062374/ai-openai-anthropic-deepmind-legal-liability-gabriel-weil
Episode art by Hamish Doodles: hamishdoodles.com
A lot of work to prevent AI existential risk takes the form of ensuring that AIs don't want to cause harm or take over the world---or in other words, ensuring that they're aligned. In this episode, I talk with Buck Shlegeris and Ryan Greenblatt about a different approach, called "AI control": ensuring that AI systems couldn't take over the world, even if they were trying to.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
0:00:31 - What is AI control?
0:16:16 - Protocols for AI control
0:22:43 - Which AIs are controllable?
0:29:56 - Preventing dangerous coded AI communication
0:40:42 - Unpredictably uncontrollable AI
0:58:01 - What control looks like
1:08:45 - Is AI control evil?
1:24:42 - Can red teams match misaligned AI?
1:36:51 - How expensive is AI monitoring?
1:52:32 - AI control experiments
2:03:50 - GPT-4's aptitude at inserting backdoors
2:14:50 - How AI control relates to the AI safety field
2:39:25 - How AI control relates to previous Redwood Research work
2:49:16 - How people can work on AI control
2:54:07 - Following Buck and Ryan's research
The transcript: axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.html
Links for Buck and Ryan:
Buck's twitter/X account: twitter.com/bshlgrs
Ryan on LessWrong: lesswrong.com/users/ryan_greenblatt
You can contact both Buck and Ryan by electronic mail at [firstname] [at-sign] rdwrs.com
Main research works we talk about:
The case for ensuring that powerful AIs are controlled: lesswrong.com/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlled
AI Control: Improving Safety Despite Intentional Subversion: arxiv.org/abs/2312.06942
Other things we mention:
The prototypical catastrophic AI action is getting root access to its datacenter (aka "Hacking the SSH server"): lesswrong.com/posts/BAzCGCys4BkzGDCWR/the-prototypical-catastrophic-ai-action-is-getting-root
Preventing language models from hiding their reasoning: arxiv.org/abs/2310.18512
Improving the Welfare of AIs: A Nearcasted Proposal: lesswrong.com/posts/F6HSHzKezkh6aoTr2/improving-the-welfare-of-ais-a-nearcasted-proposal
Measuring coding challenge competence with APPS: arxiv.org/abs/2105.09938
Causal Scrubbing: a method for rigorously testing interpretability hypotheses lesswrong.com/posts/JvZhhzycHu2Yd57RN/causal-scrubbing-a-method-for-rigorously-testing
Episode art by Hamish Doodles: hamishdoodles.com
The events of this year have highlighted important questions about the governance of artificial intelligence. For instance, what does it mean to democratize AI? And how should we balance benefits and dangers of open-sourcing powerful AI systems such as large language models? In this episode, I speak with Elizabeth Seger about her research on these questions.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
0:00:40 - What kinds of AI?
0:01:30 - Democratizing AI
0:04:44 - How people talk about democratizing AI
0:09:34 - Is democratizing AI important?
0:13:31 - Links between types of democratization
0:22:43 - Democratizing profits from AI
0:27:06 - Democratizing AI governance
0:29:45 - Normative underpinnings of democratization
0:44:19 - Open-sourcing AI
0:50:47 - Risks from open-sourcing
0:56:07 - Should we make AI too dangerous to open source?
1:00:33 - Offense-defense balance
1:03:13 - KataGo as a case study
1:09:03 - Openness for interpretability research
1:15:47 - Effectiveness of substitutes for open sourcing
1:20:49 - Offense-defense balance, part 2
1:29:49 - Making open-sourcing safer?
1:40:37 - AI governance research
1:41:05 - The state of the field
1:43:33 - Open questions
1:49:58 - Distinctive governance issues of x-risk
1:53:04 - Technical research to help governance
1:55:23 - Following Elizabeth's research
The transcript: https://axrp.net/episode/2023/11/26/episode-26-ai-governance-elizabeth-seger.html
Links for Elizabeth:
Personal website: elizabethseger.com
Centre for the Governance of AI (AKA GovAI): governance.ai
Main papers:
Democratizing AI: Multiple Meanings, Goals, and Methods: arxiv.org/abs/2303.12642
Open-sourcing highly capable foundation models: an evaluation of risks, benefits, and alternative methods for pursuing open source objectives: papers.ssrn.com/sol3/papers.cfm?abstract_id=4596436
Other research we discuss:
What Do We Mean When We Talk About "AI democratisation"? (blog post): governance.ai/post/what-do-we-mean-when-we-talk-about-ai-democratisation
Democratic Inputs to AI (OpenAI): openai.com/blog/democratic-inputs-to-ai
Collective Constitutional AI: Aligning a Language Model with Public Input (Anthropic): anthropic.com/index/collective-constitutional-ai-aligning-a-language-model-with-public-input
Against "Democratizing AI": johanneshimmelreich.net/papers/against-democratizing-AI.pdf
Adversarial Policies Beat Superhuman Go AIs: goattack.far.ai
Structured access: an emerging paradigm for safe AI deployment: arxiv.org/abs/2201.05159
Universal and Transferable Adversarial Attacks on Aligned Language Models (aka Adversarial Suffixes): arxiv.org/abs/2307.15043
Episode art by Hamish Doodles: hamishdoodles.com
Imagine a world where there are many powerful AI systems, working at cross purposes. You could suppose that different governments use AIs to manage their militaries, or simply that many powerful AIs have their own wills. At any rate, it seems valuable for them to be able to cooperatively work together and minimize pointless conflict. How do we ensure that AIs behave this way - and what do we need to learn about how rational agents interact to make that more clear? In this episode, I'll be speaking with Caspar Oesterheld about some of his research on this very topic.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Episode art by Hamish Doodles: hamishdoodles.com
Topics we discuss, and timestamps:
0:00:34 - Cooperative AI
0:06:21 - Cooperative AI vs standard game theory
0:19:45 - Do we need cooperative AI if we get alignment?
0:29:29 - Cooperative AI and agent foundations
0:34:59 - A Theory of Bounded Inductive Rationality
0:50:05 - Why it matters
0:53:55 - How the theory works
1:01:38 - Relationship to logical inductors
1:15:56 - How fast does it converge?
1:19:46 - Non-myopic bounded rational inductive agents?
1:24:25 - Relationship to game theory
1:30:39 - Safe Pareto Improvements
1:30:39 - What they try to solve
1:36:15 - Alternative solutions
1:40:46 - How safe Pareto improvements work
1:51:19 - Will players fight over which safe Pareto improvement to adopt?
2:06:02 - Relationship to program equilibrium
2:11:25 - Do safe Pareto improvements break themselves?
2:15:52 - Similarity-based Cooperation
2:23:07 - Are similarity-based cooperators overly cliqueish?
2:27:12 - Sensitivity to noise
2:29:41 - Training neural nets to do similarity-based cooperation
2:50:25 - FOCAL, Caspar's research lab
2:52:52 - How the papers all relate
2:57:49 - Relationship to functional decision theory
2:59:45 - Following Caspar's research
The transcript: axrp.net/episode/2023/10/03/episode-25-cooperative-ai-caspar-oesterheld.html
Links for Caspar:
FOCAL at CMU: www.cs.cmu.edu/~focal/
Caspar on X, formerly known as Twitter: twitter.com/C_Oesterheld
Caspar's blog: casparoesterheld.com/
Caspar on Google Scholar: scholar.google.com/citations?user=xeEcRjkAAAAJ&hl=en&oi=ao
Research we discuss:
A Theory of Bounded Inductive Rationality: arxiv.org/abs/2307.05068
Safe Pareto improvements for delegated game playing: link.springer.com/article/10.1007/s10458-022-09574-6
Similarity-based Cooperation: arxiv.org/abs/2211.14468
Logical Induction: arxiv.org/abs/1609.03543
Program Equilibrium: citeseerx.ist.psu.edu/document?repid=rep1&type=pdf&doi=e1a060cda74e0e3493d0d81901a5a796158c8410
Formalizing Objections against Surrogate Goals: www.alignmentforum.org/posts/K4FrKRTrmyxrw5Dip/formalizing-objections-against-surrogate-goals
Learning with Opponent-Learning Awareness: arxiv.org/abs/1709.04326
Recently, OpenAI made a splash by announcing a new "Superalignment" team. Lead by Jan Leike and Ilya Sutskever, the team would consist of top researchers, attempting to solve alignment for superintelligent AIs in four years by figuring out how to build a trustworthy human-level AI alignment researcher, and then using it to solve the rest of the problem. But what does this plan actually involve? In this episode, I talk to Jan Leike about the plan and the challenges it faces.
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
The transcript: axrp.net/episode/2023/07/27/episode-24-superalignment-jan-leike.html
Links for Jan and OpenAI:
Links to research and other writings we discuss:
Is there some way we can detect bad behaviour in our AI system without having to know exactly what it looks like? In this episode, I speak with Mark Xu about mechanistic anomaly detection: a research direction based on the idea of detecting strange things happening in neural networks, in the hope that that will alert us of potential treacherous turns. We both talk about the core problems of relating these mechanistic anomalies to bad behaviour, as well as the paper "Formalizing the presumption of independence", which formulates the problem of formalizing heuristic mathematical reasoning, in the hope that this will let us mathematically define "mechanistic anomalies".
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
The transcript: axrp.net/episode/2023/07/24/episode-23-mechanistic-anomaly-detection-mark-xu.html
ARC links:
Research we discuss:
Very brief survey: bit.ly/axrpsurvey2023
Store is closing in a week! Link: store.axrp.net/
Patreon: patreon.com/axrpodcast
Ko-fi: ko-fi.com/axrpodcast
What can we learn about advanced deep learning systems by understanding how humans learn and form values over their lifetimes? Will superhuman AI look like ruthless coherent utility optimization, or more like a mishmash of contextually activated desires? This episode's guest, Quintin Pope, has been thinking about these questions as a leading researcher in the shard theory community. We talk about what shard theory is, what it says about humans and neural networks, and what the implications are for making AI safe.
Patreon: patreon.com/axrpodcast
Store: store.axrp.net
Ko-fi: ko-fi.com/axrpodcast
Episode art by Hamish Doodles
Topics we discuss, and timestamps:
The transcript
Shard theorist links:
Research we discuss:
Links for the addendum on mesa-optimization skepticism:
Lots of people in the field of machine learning study 'interpretability', developing tools that they say give us useful information about neural networks. But how do we know if meaningful progress is actually being made? What should we want out of these tools? In this episode, I speak to Stephen Casper about these questions, as well as about a benchmark he's co-developed to evaluate whether interpretability tools can find 'Trojan horses' hidden inside neural nets.
Patreon: patreon.com/axrpodcast
Store: store.axrp.net
Ko-fi: ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
The transcript
Links for Casper:
Research we discuss:
How should we scientifically think about the impact of AI on human civilization, and whether or not it will doom us all? In this episode, I speak with Scott Aaronson about his views on how to make progress in AI alignment, as well as his work on watermarking the output of language models, and how he moved from a background in quantum complexity theory to working on AI.
Note: this episode was recorded before this story emerged of a man committing suicide after discussions with a language-model-based chatbot, that included discussion of the possibility of him killing himself.
Patreon: https://www.patreon.com/axrpodcast
Store: https://store.axrp.net/
Ko-fi: https://ko-fi.com/axrpodcast
Topics we discuss, and timestamps:
The transcript
Links to Scott's things:
Writings we discuss:
Store: https://store.axrp.net/
Patreon: https://www.patreon.com/axrpodcast
Ko-fi: https://ko-fi.com/axrpodcast
Video: https://www.youtube.com/watch?v=kmPFjpEibu0
How good are we at understanding the internal computation of advanced machine learning models, and do we have a hope at getting better? In this episode, Neel Nanda talks about the sub-field of mechanistic interpretability research, as well as papers he's contributed to that explore the basics of transformer circuits, induction heads, and grokking.
Topics we discuss, and timestamps:
The transcript
Links to Neel's things:
Writings we discuss:
I have a new podcast, where I interview whoever I want about whatever I want. It's called "The Filan Cabinet", and you can find it wherever you listen to podcasts. The first three episodes are about pandemic preparedness, God, and cryptocurrency. For more details, check out the podcast website, or search "The Filan Cabinet" in your podcast app.
Concept extrapolation is the idea of taking concepts an AI has about the world - say, "mass" or "does this picture contain a hot dog" - and extending them sensibly to situations where things are different - like learning that the world works via special relativity, or seeing a picture of a novel sausage-bread combination. For a while, Stuart Armstrong has been thinking about concept extrapolation and how it relates to AI alignment. In this episode, we discuss where his thoughts are at on this topic, what the relationship to AI alignment is, and what the open questions are.
Topics we discuss, and timestamps:
The transcript
Stuart's startup, Aligned AI
Research we discuss:
Concept extrapolation is the idea of taking concepts an AI has about the world - say, "mass" or "does this picture contain a hot dog" - and extending them sensibly to situations where things are different - like learning that the world works via special relativity, or seeing a picture of a novel sausage-bread combination. For a while, Stuart Armstrong has been thinking about concept extrapolation and how it relates to AI alignment. In this episode, we discuss where his thoughts are at on this topic, what the relationship to AI alignment is, and what the open questions are.
Topics we discuss, and timestamps:
The transcript
Stuart's startup, Aligned AI
Research we discuss:
Sometimes, people talk about making AI systems safe by taking examples where they fail and training them to do well on those. But how can we actually do this well, especially when we can't use a computer program to say what a 'failure' is? In this episode, I speak with Daniel Ziegler about his research group's efforts to try doing this with present-day language models, and what they learned.
Listeners beware: this episode contains a spoiler for the Animorphs franchise around minute 41 (in the 'Fanfiction' section of the transcript).
Topics we discuss, and timestamps:
The transcript
Daniel Ziegler on Google Scholar
Research we discuss:
Many people in the AI alignment space have heard of AI safety via debate - check out AXRP episode 6 if you need a primer. But how do we get language models to the stage where they can usefully implement debate? In this episode, I talk to Geoffrey Irving about the role of language models in AI safety, as well as three projects he's done that get us closer to making debate happen: using language models to find flaws in themselves, getting language models to back up claims they make with citations, and figuring out how uncertain language models should be about the quality of various answers.
Topics we discuss, and timestamps:
The transcript
Geoffrey's twitter
Research we discuss:
Why does anybody care about natural abstractions? Do they somehow relate to math, or value learning? How do E. coli bacteria find sources of sugar? All these questions and more will be answered in this interview with John Wentworth, where we talk about his research plan of understanding agency via natural abstractions. Topics we discuss, and timestamps:
The transcript
John on LessWrong
Research that we discuss:
Errata:
Late last year, Vanessa Kosoy and Alexander Appel published some research under the heading of "Infra-Bayesian physicalism". But wait - what was infra-Bayesianism again? Why should we care? And what does any of this have to do with physicalism? In this episode, I talk with Vanessa Kosoy about these questions, and get a technical overview of how infra-Bayesian physicalism works and what its implications are.
Topics we discuss, and timestamps:
The transcript
Vanessa on the Alignment Forum
Research that we discuss:
How should we think about artificial general intelligence (AGI), and the risks it might pose? What constraints exist on technical solutions to the problem of aligning superhuman AI systems with human intentions? In this episode, I talk to Richard Ngo about his report analyzing AGI safety from first principles, and recent conversations he had with Eliezer Yudkowsky about the difficulty of AI alignment.
Topics we discuss, and timestamps:
The transcript
Richard on the Alignment Forum
Richard on Twitter
The AGI Safety Fundamentals course
Materials that we mention:
Why would advanced AI systems pose an existential risk, and what would it look like to develop safer systems? In this episode, I interview Paul Christiano about his views of how AI could be so dangerous, what bad AI scenarios could look like, and what he thinks about various techniques to reduce this risk.
Topics we discuss, and timestamps (due to mp3 compression, the timestamps may be tens of seconds off):
The transcript
Paul's blog posts on AI alignment
Material that we mention:
Many scary stories about AI involve an AI system deceiving and subjugating humans in order to gain the ability to achieve its goals without us stopping it. This episode's guest, Alex Turner, will tell us about his research analyzing the notions of "attainable utility" and "power" that underlie these stories, so that we can better evaluate how likely they are and how to prevent them.
Topics we discuss:
The transcript
Alex on the AI Alignment Forum
Alex's Google Scholar page
Conservative Agency via Attainable Utility Preservation
Optimal Policies Tend to Seek Power
Other works discussed:
When going about trying to ensure that AI does not cause an existential catastrophe, it's likely important to understand how AI will develop in the future, and why exactly it might or might not cause such a catastrophe. In this episode, I interview Katja Grace, researcher at AI Impacts, who's done work surveying AI researchers about when they expect superhuman AI to be reached, collecting data about how rapidly AI tends to progress, and thinking about the weak points in arguments that AI could be catastrophic for humanity.
Topics we discuss:
The transcript
"When Will AI Exceed Human Performance? Evidence from AI Experts"
AI Impacts page of more complete survey results
Likelihood of discontinuous progress around the development of AGI
Discontinuous progress investigation
The range of human intelligence
Being an agent can get loopy quickly. For instance, imagine that we're playing chess and I'm trying to decide what move to make. Your next move influences the outcome of the game, and my guess of that influences my move, which influences your next move, which influences the outcome of the game. How can we model these dependencies in a general way, without baking in primitive notions of 'belief' or 'agency'? Today, I talk with Scott Garrabrant about his recent work on finite factored sets that aims to answer this question.
Topics we discuss:
Link to the transcript
Link to a transcript of Scott's talk on finite factored sets
Scott's LessWrong account
Other work mentioned in the discussion:
How should we think about the technical problem of building smarter-than-human AI that does what we want? When and how should AI systems defer to us? Should they have their own goals, and how should those goals be managed? In this episode, Dylan Hadfield-Menell talks about his work on assistance games that formalizes these questions. The first couple years of my PhD program included many long conversations with Dylan that helped shape how I view AI x-risk research, so it was great to have another one in the form of a recorded interview.
Link to the transcript
Link to the paper "Cooperative Inverse Reinforcement Learning"
Link to the paper "The Off-Switch Game"
Link to the paper "Inverse Reward Design"
Dylan's twitter account
Link to apply to the MIT EECS graduate program
Other work mentioned in the discussion:
If you want to shape the development and forecast the consequences of powerful AI technology, it's important to know when it might appear. In this episode, I talk to Ajeya Cotra about her draft report "Forecasting Transformative AI from Biological Anchors" which aims to build a probabilistic model to answer this question. We talk about a variety of topics, including the structure of the model, what the most important parts are to get right, how the estimates should shape our behaviour, and Ajeya's current work at Open Philanthropy and perspective on the AI x-risk landscape.
Unfortunately, there was a problem with the recording of our interview, so we weren't able to release it in audio form, but you can read a transcript of the whole conversation.
Link to the transcript
Link to the draft report "Forecasting Transformative AI from Biological Anchors"
One way of thinking about how AI might pose an existential threat is by taking drastic actions to maximize its achievement of some objective function, such as taking control of the power supply or the world's computers. This might suggest a mitigation strategy of minimizing the degree to which AI systems have large effects on the world that are not absolutely necessary for achieving their objective. In this episode, Victoria Krakovna talks about her research on quantifying and minimizing side effects. Topics discussed include how one goes about defining side effects and the difficulties in doing so, her work using relative reachability and the ability to achieve future tasks as side effects measures, and what she thinks the open problems and difficulties are.
Link to the transcript
Link to the paper "Penalizing Side Effects Using Stepwise Relative Reachability"
Link to the paper "Avoiding Side Effects by Considering Future Tasks"
Victoria Krakovna's website
Victoria Krakovna's Alignment Forum profile
Work mentioned in the episode:
One proposal to train AIs that can be useful is to have ML models debate each other about the answer to a human-provided question, where the human judges which side has won. In this episode, I talk with Beth Barnes about her thoughts on the pros and cons of this strategy, what she learned from seeing how humans behaved in debate protocols, and how a technique called imitative generalization can augment debate. Those who are already quite familiar with the basic proposal might want to skip past the explanation of debate to 13:00, "what problems does it solve and does it not solve".
Link to Beth's posts on the Alignment Forum
Link to the transcript
The theory of sequential decision-making has a problem: how can we deal with situations where we have some hypotheses about the environment we're acting in, but its exact form might be outside the range of possibilities we can possibly consider? Relatedly, how do we deal with situations where the environment can simulate what we'll do in the future, and put us in better or worse situations now depending on what we'll do then? Today's episode features Vanessa Kosoy talking about infra-Bayesianism, the mathematical framework she developed with Alex Appel that modifies Bayesian decision theory to succeed in these types of situations. Link to the listener survey
Link to the sequence of posts - Infra-Bayesianism
Link to the transcript
Vanessa Kosoy's Alignment Forum profile
In machine learning, typically optimization is done to produce a model that performs well according to some metric. Today's episode features Evan Hubinger talking about what happens when the learned model itself is doing optimization in order to perform well, how the goals of the learned model could differ from the goals we used to select the learned model, and what would happen if they did differ.
Link to the paper - Risks from Learned Optimization in Advanced Machine Learning Systems
Link to the transcript
Evan Hubinger's Alignment Forum profile
In this episode, I talk with Andrew Critch about negotiable reinforcement learning: what happens when two people (or organizations, or what have you) who have different beliefs and preferences jointly build some agent that will take actions in the real world. In the paper we discuss, it's proven that the only way to make such an agent Pareto optimal - that is, have it not be the case that there's a different agent that both people would prefer to use instead - is to have it preferentially optimize the preferences of whoever's beliefs were more accurate. We discuss his motivations for working on the problem and what he thinks about it.
Link to the paper - Negotiable Reinforcement Learning for Pareto Optimal Sequential Decision-Making
Link to the transcript
Critch's Google Scholar profile
One approach to creating useful AI systems is to watch humans doing a task, infer what they're trying to do, and then try to do that well. The simplest way to infer what the humans are trying to do is to assume there's one goal that they share, and that they're optimally achieving the goal. This has the problem that humans aren't actually optimal at achieving the goals they pursue. We could instead code in the exact way in which humans behave suboptimally, except that we don't know that either. In this episode, I talk with Rohin Shah about his paper about learning the ways in which humans are suboptimal at the same time as learning what goals they pursue: why it's hard, how he tried to do it, how well he did, and why it matters.
Link to the paper - On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference
Link to the transcript
The Alignment Newsletter
Rohin's contributions to the AI alignment forum
Rohin's website
In this episode, Adam Gleave and I talk about adversarial policies. Basically, in current reinforcement learning, people train agents that act in some kind of environment, sometimes an environment that contains other agents. For instance, you might train agents that play sumo with each other, with the objective of making them generally good at sumo. Adam's research looks at the case where all you're trying to do is make an agent that defeats one specific other agents: how easy is it, and what happens? He discovers that often, you can do it pretty easily, and your agent can behave in a very silly-seeming way that nevertheless happens to exploit some 'bug' in the opponent. We talk about the experiments he ran, the results, and what they say about how we do reinforcement learning.
Link to the paper - Adversarial Policies: Attacking Deep Reinforcement Learning
Link to the transcript
Adam's website
Adam's twitter account