A weekly podcast updating you with the latest research in AI and machine learning safety from people such as DeepMind, Anthropic, and MIRI.
5 years ago, the Google AlphaGo beated reigning world number 1 in Go, Ke Jie, but if you think the board game playing AI's have stopped evolving since, think twice! Today we will look into the new language model, Cicero's, deceptive abilities along with considerations on what board-game playing AI's teach us about AI-development.
Table of contents:
Sources
AlphaGo beating Ke Jie in GO (5 years ago)
https://www.bbc.com/news/technology-40042581
Considerations on the funding situation for AI Safety, exciting projects from Apart's interpretability hackathon, Meta AI-math transformer interpretability and considerations on what to spend time on in AI Safety.
00:25 Thoughts of the FTX-Crash and the Future of EA
2:20 Alignment Hackathon Interpretability Research
4:20 Mathematical Transformer Interpretability
5:10 Should We Focus on Buying Time instead of Technical Solutions
Sources:
Evan Hubinger on why Fraud must never happen in Effective Altruism: https://www.lesswrong.com/posts/8wYH4WggxFqT9yhzJ/we-must-be-very-clear-fraud-in-the-service-of-effective
Strawberry Calm on what we can consider in EA: https://forum.effectivealtruism.org/posts/yMKmCbmL8ekDcJhQd/ (in the comments)
Emergency funding: https://www.lesswrong.com/posts/SkoxdYCAozBPfcJrP/announcing-nonlinear-emergency-funding
Apart's interpretability hackathon, https://itch.io/jam/interpretability
1st price: Investigating Neuron Behaviour via Dataset Example Pruning and Local Search, https://alexfoote.itch.io/investigating-neuron-behaviour-via-dataset-example-pruning-and-local-search
2nd price: Backup Head Behaviour is Robust to the Distribution Used to Perform the Ablation, https://satojk.itch.io/backup-transformer-heads-are-robust
3rd price: Model editing hazards at the example of ROME, https://jas-ho.itch.io/model-editing-hazards-at-the-example-of-rome
Thoughts on buying time for AI Safety
https://www.lesswrong.com/posts/bkpZHXMJx3dG5waA7/ways-to-buy-time
Martin Soto criticizes Vannessa Kosoy's PreDCA-protocol for interpretability https://www.lesswrong.com/posts/FhKkFcojhKZt7nHzG/a-short-critique-of-vanessa-kosoy-s-predca-1
Will we run out of ML-data? https://www.lesswrong.com/posts/Couhhp4pPHbbhJ2Mg/will-we-run-out-of-ml-data-evidence-from-projecting-dataset
Black box agnostic models are worth still considering? https://www.lesswrong.com/posts/uXGLciramzNfb8Hvz/why-i-m-working-on-model-agnostic-interpretability
Instrumental convergence explains why general intelligence is possible https://www.lesswrong.com/posts/GZgLa5Xc4HjwketWe/instrumental-convergence-is-what-makes-general-intelligence
Opportunities:
The crypto giant FTX crashes, introducing massive uncertainty in the funding space for AI safety, humans cooperate better with lying AI, and interpretability is promising but also not.
Content
00:25 Uncertainty for AI Safety Funding
1:20 Human-AI cooperation
2:30 Interpretability in the wild
3:25 Other news
Opportunities
Sources:
FTX fallout continues to roll out markets, https://www.youtube.com/watch?v=IgpbdnXOpEk
FTX has probably collapsed, https://forum.effectivealtruism.org/posts/tdLRvYHpfYjimwhyL/ftx-com-has-probably-collapsed
Human-AI cooperation is better when non-calibrated confidence, https://arxiv.org/pdf/2202.05983.pdf
Google shows inverse scaling can become U-shaped, https://www.alignmentforum.org/posts/LvKmjKMvozpdmiQhP/inverse-scaling-can-become-u-shaped
Ethan Perez calls them out on their methods, https://twitter.com/EthanJPerez/status/1588352204540235776
Interpretability in the wild,
https://arxiv.org/pdf/2211.00593.pdf
Eric Drexler discusses superintelligences with Eliezer (in the comments), https://www.alignmentforum.org/posts/HByDKLLdaWEcA2QQD/applying-superintelligence-without-collusion
Janus (an alias for several people) shows which places GPT-3 davinci-text-002 selects very specific outputs - favorite number, no answers, etc, https://www.alignmentforum.org/posts/t9svvNPNmFf5Qa3TA/mysteries-of-mode-collapse-due-to-rlhf
Mesa-optimizer implemented, https://www.alignmentforum.org/posts/b44zed5fBWyyQwBHL/trying-to-make-a-treacherous-mesa-optimizer
David Krueger argues against mechanistic interpretability, https://www.alignmentforum.org/posts/kjRGMdRxXb9c5bWq5/mechanistic-interpretability-as-reverse-engineering-follow
Nate Soares overviews strategies for “knowing AGI is safe”, https://www.alignmentforum.org/posts/iDFTmb8HSGtL4zTvf/how-could-we-know-that-an-agi-system-will-have-good
Interpretability starter resources, https://ais.pub/alignmentjam
This week, we look at broken scaling laws, surgical fine-tuning, interpretability in the wild, and threat models of AI.
Opportunities
Redwood Research is inviting 30-50 researchers to join them in Berkeley for a very interesting mechanistic interpretability research programme. https://ais.pub/remix
Anthropic is looking for operations managers, recruiters, researchers, engineers, and product managers: https://ais.pub/anthropic
AI Safety Ideas: https://ais.pub/aisi
Hackathons: https://ais.pub/jam
00:20 Broken scaling laws & surgical fine-tuning
01:34 Debate & interpretability
02:55 Threat models in ML safety
04:30 Other news
05:30 REMIX, Anthropic, and hackathons
Sources
Interpreting DeepMind’s in-context RL paper: https://www.lesswrong.com/posts/avvXAvGhhGgkJDDso/caution-when-interpreting-deepmind-s-in-context-rl-paper
Interpretability in the Wild: https://twitter.com/kevrowan/status/1587601532639494146
Precision machine learning, Max Tegmark: https://arxiv.org/pdf/2210.13447.pdf
Broken natural scaling laws: https://arxiv.org/pdf/2210.14891.pdf
Debate does not work for question-answering: https://arxiv.org/pdf/2210.10860.pdf
Vision for meta-science: https://scienceplusplus.org/metascience/index.html#how-does-the-culture-of-science-change
AI X-risk 35% mostly based on a recent peer-reviewed argument - LessWrong: https://www.lesswrong.com/posts/XtBJTFszs8oP3vXic/ai-x-risk-greater-than-35-mostly-based-on-a-recent-peer
Clarifying X-risk: https://www.lesswrong.com/posts/GctJD5oCDRxCspEaZ/clarifying-ai-x-risk
Threat model literature review: https://www.lesswrong.com/posts/wnnkD6P2k2TfHnNmt/threat-model-literature-review
Frames vs. boundaries: https://www.alignmentforum.org/posts/SZjHimszxqjNJzQWK/boundaries-vs-frames
Surgical fine-tuning improves robustness against distributional shifts: https://arxiv.org/pdf/2210.11466.pdf
This week, we look at how we can safeguard against AGI, look at new research on Goodhart’s law, see an open source dataset with 60,000 emotional videos, and share new opportunities in ML and AI safety.
Sources
🎉 You can now subscribe to our newsletter and listen to these updates in your favorite podcasting app. Check out newsletter.apartresearch.com and podcast.apartresearch.com.
This week, we’re looking at counterarguments to the basic case for why AI is an existential risk to humanity, looking at how strong AI might come very soon, and sharing interesting papers.
Sources:
✉️ Read the ML Safety Newsletter: https://ais.pub/mlsn6
📈 Help Center for AI Safety come up with new benchmarks: https://ais.pub/benchmarks
💡 Help Redwood Research find interesting heuristics: https://ais.pub/gptheuristics
👩🔬 Join the hackathon: https://ais.pub/interpretability
🙋♀️ Register as a local organizer: https://ais.pub/localorganizer
🤓 Read the Alignment 201 curriculum: https://ais.pub/alignment201
[Corrections] Marcus = Marius. AGI Safety Fundamentals is not taking applicants but you can read the Alignment 201 curriculum here:
https://ais.pub/alignment201.
00:00 Intro
00:19 Law defines alignment principles
00:40 Out-of-distribution alignment
01:27 Reward hacking defined
02:38 Inductive biases in learning algorithms
03:34 Warning shots are not enough
04:04 State of AI safety
04:45 Announcements
Sources:
Legal informatics for AI safety, robust specification https://arxiv.org/abs/2209.13020
Out-of-distribution GAN examples https://arxiv.org/abs/2209.11960
Formal definition of ‘reward hacking’ https://arxiv.org/abs/2209.13085
DeepMind: Why correct goals are not enough https://arxiv.org/abs/2210.01790
QAPR 4, inductive biases of learning processes https://www.alignmentforum.org/posts/...
QAPR 3: Training NNs from interpretability priors https://www.alignmentforum.org/s/5omS...
Neural tangent Kernel distillation https://www.alignmentforum.org/posts/...
OG paper: https://arxiv.org/pdf/1806.07572.pdf
Gaussian processes https://distill.pub/2019/visual-explo...
Soares’ critique of warning shots https://www.alignmentforum.org/posts/...
~300 people in AIS https://forum.effectivealtruism.org/p...
Statistics of machine learning https://financesonline.com/machine-le...
Chatting AI safety with 100+ researchers https://www.alignmentforum.org/posts/...
Smaller news
MLSN https://www.alignmentforum.org/posts/...
Safety benchmarks prize https://benchmarking.mlsafety.org/
Finding heuristics of GPT-2 small https://www.lesswrong.com/posts/LkBmA...
Alignment 201 curriculum: https://www.agisafetyfundamentals.com...
[correction] Eliezer kindly reached out to us to note that this is not the correct criticism from MIRI of CHAI's position. Here it is in his own words:
The basic argument goes that an AI whose utility function has been made dependent on hidden information, even if that information is inside humans, won't defer to humans because of that; it gets all the information that's obtainable and then ignores the humans (and kills them). There's never a point where "let the humans shut me off and build another AI" looks like a better strategy than "get all the info out of the humans and then stop listening".
Welcome to this week's Safe AI Progress Report where Thomas describes the scary developments in AI, a theoretical dispute, and risks from power-seeking AI.
00:00 Intro
00:18 Carmack's AGI
00:36 Video generation
00:57 AlphaTensor
01:26 PyTorch to Linux Foundation
02:00 Risks from power-seeking AI
03:09 MIRI criticizes CHAI's approach
03:55 Smaller news
04:35 Interpretability hackathon
Sources:
Capabilities updates
$20 million to develop AGI: https://www.insiderintelligence.com/c...
John Carmack doesn’t care about safety: https://twitter.com/ID_AA_Carmack/sta...
Meta’s GAN: https://ai.facebook.com/blog/generati...
Meta’s video gan is actually bad: https://phenaki.video/#interactive:~:...
Whisper: https://cdn.openai.com/papers/whisper...
People give the AI access to the internet: https://twitter.com/sergeykarayev/sta...
DeepMind’s AlphaMath matrix multiplication: https://www.deepmind.com/blog/discove...
Quanta article: https://www.quantamagazine.org/mathem...
Alignment researcher chatter: https://www.alignmentforum.org/posts/...
Grace
PyTorch to Linux: https://pytorch.org/blog/PyTorchfound...
Neutrality first (still better than Meta): https://www.linuxfoundation.org/blog/...
Yann LeCunn developing AGI: https://openreview.net/pdf?id=BZ5a1r-...
Paperclip clicker games: https://paperclips.tech/
Review of AI risk outlook: https://www.alignmentforum.org/posts/...
Joe Carlsmith’s report: https://arxiv.org/abs/2206.13353
Video lecture of the report: https://forum.effectivealtruism.org/p...
Endgames: https://www.alignmentforum.org/posts/...
Machine alignment Monday by Scott Alexander: https://astralcodexten.substack.com/p...
Article discussed: https://arbital.com/p/updated_deference/
Tamsin Leake’s outlook on AI safety: https://carado.moe/outlook-ai-risk-mi...
Alex’s loss functions: https://www.alignmentforum.org/posts/...
Physics-based deep learning: https://physicsbaseddeeplearning.org/...
Andrej Karpathy’s first video: https://www.youtube.com/watch?v=VMj-3...
Amazing interpretability tool from Redwood: http://interp-tools.redwoodresearch.org/
Tutorial for use: https://docs.google.com/document/d/1E...
Anthropic’s Svelte work: https://github.com/anthropics/PySvelte
Microscope: https://openai.com/blog/microscope/
Join the hackathon, Esben’s interpretability talk: https://itch.io/jam/interpretability
Join the hackathon GatherTown here: https://app.gather.town/app/iEPL2kx1L...
Or read more about the event here: https://itch.io/jam/llm-hackathon
👩🔬 This week's Safe AI Progress Report puts the focus on the Future Fund's new prize, Conjecture's new research, speed priors, and much more. Come along!
Chapters:
00:00 Intro
00:16 $1.5 million prize
01:00 Conjecture's interpretability research
01:50 Anthropic's superposition experiments
02:30 Inverse Scaling first round prizes
03:08 Forwarding speed regularizers
04:00 Reward is not the optimization target
04:25 Other news
04:52 🎉 Apart Hackathon
05:08 Israel AI Safety
05:20 Outro
Sources:
Future Fund world view competition: https://ftxfuturefund.org/
Strong general AGI soon: https://forum.effectivealtruism.org/p...
Polytopes lens: https://www.alignmentforum.org/posts/...
Anthropic’s papers: https://www.anthropic.com/research
Toy models of superposition: https://transformer-circuits.pub/2022...
Inverse scaling prize round 1: https://www.alignmentforum.org/posts/...
Inverse scaling prize: https://github.com/inverse-scaling/prize
Speed prior and forwarding speed priors: https://www.alignmentforum.org/posts/...
Are minimal circuits deceptive?: https://www.lesswrong.com/posts/fM5ZW...
Musings on the speed prior: https://www.alignmentforum.org/posts/...
Deconfusing wireheading: https://www.alignmentforum.org/posts/...
Reward is not the optimization target: https://www.alignmentforum.org/posts/...
Nearcasting AGI: https://www.alignmentforum.org/posts/...
7 traps new alignment researchers drop into: https://www.lesswrong.com/posts/h5CGM...
Language model hackathon: https://itch.io/jam/llm-hackathon
AI Safety Israel conference: https://aisic2022.net.technion.ac.il/
Apart Research: https://apartresearch.com
AI Safety Ideas: https://aisi.ai
👩🔬 Hosted by Sabrina Zaki
This Safe AI Progress Report describes the past week's developments in ML and AI safety. Follow along to get regular updates for the scientific field to stay safe from artificial intelligence.
Citations
AI Safety Ideas: https://aisafetyideas.com
🏛 Citations:
Circuits: https://distill.pub/2020/circuits/zoom-in/
Interpretability survey: https://arxiv.org/abs/2207.13243, see Twitter summary: https://twitter.com/StephenLCasper/status/1569401262558576642, and PDF: https://arxiv.org/pdf/2207.13243.pdf
Activation atlas: https://distill.pub/2019/activation-atlas/
Changing training data: https://arxiv.org/pdf/1811.12231.pdf
Editing factual associations in GPT: https://arxiv.org/pdf/2202.05262.pdf
Natural language descriptions of deep visual features: https://arxiv.org/pdf/2201.11114.pdf
Robust feature-level adversaries are interpretability tools: https://arxiv.org/pdf/2110.03605.pdf
Samotsvety’s AI risk forecast: https://forum.effectivealtruism.org/posts/EG9xDM8YRz4JN4wMN/samotsvety-s-ai-risk-forecasts
Date of AGI: https://www.metaculus.com/questions/5121/date-of-artificial-general-intelligence/
(June) Forecasting TAI with biological anchors summary: https://www.lesswrong.com/s/B9Qc8ifidAtDpsuu8/p/wgio8E758y9XWsi8j
Monitoring for deceptive alignment: https://www.alignmentforum.org/posts/Km9sHjHTsBdbgwKyi/monitoring-for-deceptive-alignment
Deceptive alignment: https://www.alignmentforum.org/posts/zthDPAjh9w6Ytbeks/deceptive-alignment
Quintin’s alignment paper lineup: https://www.lesswrong.com/posts/7cHgjJR2H5e4w4rxT/quintin-s-alignment-papers-roundup-week-1
Most people start with the same few bad ideas: https://www.lesswrong.com/posts/Afdohjyt6gESu4ANf/most-people-start-with-the-same-few-bad-ideas
Beth Barnes starting a risks and development evaluations group at ARC: https://www.alignmentforum.org/posts/svhQMdsefdYFDq5YM/evaluations-project-arc-is-hiring-a-researcher-and-a-webdev-1
Cognitive biases in LLMs: https://arxiv.org/pdf/2206.14576.pdf
Academia vs. industry: https://www.alignmentforum.org/posts/HXxHcRCxR4oHrAsEr/an-update-on-academia-vs-industry-one-year-into-my-faculty
🥳 Welcome to this Safe AI Progress Report! It is the first of a regular series of videos that summarize what has happened in AI safety since the last report. They accompany a different progress report that focuses on measurable metrics for the field of AI safety.
📃Script
00:00 - Intro
00:13 - OpenAI alignment strategy
00:53 - Strategy explanation
01:50 - John Wentworth's criticism
02:37 - Sam Bowman's NLP survey
03:17 - New perspectives | Simulators
03:43 - New perspectives | Shard theory
04:10 - The Small Side
04:36 - Outro | Learn more
🏛 Chronological references:
- OpenAI dangerous: https://techcrunch.com/2019/02/17/ope...
- OpenAI = dangerous: https://www.lesswrong.com/posts/Nqn2t...
- OpenAI alignment: https://openai.com/blog/our-approach-...
- Jacob Hilton: https://www.lesswrong.com/posts/3S4ny...
- RLHF: https://arxiv.org/abs/2009.01325
- IDA: https://forum.effectivealtruism.org/p...
- Alignment research tools: https://www.lesswrong.com/posts/ebYio...
- Elicit: https://elicit.org/search?q=AI+safety...
- Eleuther network: https://arxiv.org/pdf/2206.02841.pdf
- Not everyone happy: https://www.lesswrong.com/posts/3S4ny...
- Grokking: https://arxiv.org/pdf/2201.02177.pdf
- Deception: https://www.lesswrong.com/posts/zthDP...
- RLHF is terrible: https://www.lesswrong.com/posts/xFotX...
- Survey of NLP researchers: https://nlpsurvey.net/nlp-metasurvey-...
- Copilot good: https://github.blog/2022-09-07-resear...
- Programs for itself: https://arxiv.org/abs/2207.14502
- Simulators: https://www.lesswrong.com/posts/vJFdj...
- Shard theory: https://www.lesswrong.com/posts/iCfdc...
- Richard’s list: https://www.lesswrong.com/posts/27AWR...
- Thomas & Eli’s list: https://www.lesswrong.com/posts/QBAjn...
- Philosophy fellowship: https://philosophy.safe.ai/
- ML Safety course material: https://course.mlsafety.org/
- ML safety competitions: https://safe.ai/competitions
- Apart Research: https://apartresearch.com
- AI safety ideas: https://aisafetyideas.com/