ML Safety Report: Recent Episodes

Apart Research

A weekly podcast updating you with the latest research in AI and machine learning safety from people such as DeepMind, Anthropic, and MIRI.

View Details

5 years ago, the Google AlphaGo beated reigning world number 1 in Go, Ke Jie, but if you think the board game playing AI's have stopped evolving since, think twice! Today we will look into the new language model, Cicero's, deceptive abilities along with considerations on what board-game playing AI's teach us about AI-development.

Table of contents:

  • Language model plays Diplomacy better than humans
  • 3-dimensional chess-playing AI's might not be that dangerous
  • Presuming independence to formalise interpretability work
  • Monosemanticity engineering of toy models
  • Minor news

Sources

  • Meta AI announces Cicero able to play the board game Diplomacy better than humans - https://www.science.org/doi/10.1126/science.ade9097
  • Clarifications on the 'wireheading term'
    https://www.alignmentforum.org/posts/REesy8nqvknFFKywm/clarifying-wireheading-terminology
  • The “loss of control” scenario rests on a few key assumptions that are not justified by our current understanding of artificial intelligence research
    https://windowsontheory.org/2022/11/22/ai-will-change-the-world-but-wont-take-it-over-by-playing-3-dimensional-chess/
  • Alignment Research Center: When deduction suddenly becomes deceptive. Formalizing presumptions of independence
    https://arxiv.org/abs/2211.06738
  • Monosemanticity in neurons responding is great for interpretability
    https://arxiv.org/abs/2211.09169 (https://www.alignmentforum.org/posts/LvznjZuygoeoTpSE6/engineering-monosemanticity-in-toy-models)
  • In case we need some more thoughts on EA's relation to funders
    https://www.lesswrong.com/posts/p4XpZWcQksSiCPG72/sadly-ftx#The_Future_of_Effective_Altruist_Ethics &
    https://forum.effectivealtruism.org/posts/NeK9XYY2mDsH5bJdD/our-recommendations-for-giving-in-2022
  • Comparing AI Alignment research to orthodox and reform religions
    https://www.lesswrong.com/posts/XKraEJrQRfzbCtzKN/distillation-of-how-likely-is-deceptive-alignment
  • Conjecture report
    https://www.lesswrong.com/posts/bXTNKjsD4y3fabhwR/conjecture-a-retrospective-after-8-months-of-work-1

AlphaGo beating Ke Jie in GO (5 years ago)
https://www.bbc.com/news/technology-40042581

View Details

Considerations on the funding situation for AI Safety, exciting projects from Apart's interpretability hackathon, Meta AI-math transformer interpretability and considerations on what to spend time on in AI Safety.

00:25 Thoughts of the FTX-Crash and the Future of EA

2:20 Alignment Hackathon Interpretability Research

4:20 Mathematical Transformer Interpretability

5:10 Should We Focus on Buying Time instead of Technical Solutions

Sources:

Evan Hubinger on why Fraud must never happen in Effective Altruism: https://www.lesswrong.com/posts/8wYH4WggxFqT9yhzJ/we-must-be-very-clear-fraud-in-the-service-of-effective
Strawberry Calm on what we can consider in EA: https://forum.effectivealtruism.org/posts/yMKmCbmL8ekDcJhQd/ (in the comments)
Emergency funding: https://www.lesswrong.com/posts/SkoxdYCAozBPfcJrP/announcing-nonlinear-emergency-funding

Apart's interpretability hackathon, https://itch.io/jam/interpretability
1st price: Investigating Neuron Behaviour via Dataset Example Pruning and Local Search, https://alexfoote.itch.io/investigating-neuron-behaviour-via-dataset-example-pruning-and-local-search
2nd price: Backup Head Behaviour is Robust to the Distribution Used to Perform the Ablation, https://satojk.itch.io/backup-transformer-heads-are-robust
3rd price: Model editing hazards at the example of ROME, https://jas-ho.itch.io/model-editing-hazards-at-the-example-of-rome

Thoughts on buying time for AI Safety
https://www.lesswrong.com/posts/bkpZHXMJx3dG5waA7/ways-to-buy-time

Martin Soto criticizes Vannessa Kosoy's PreDCA-protocol for interpretability https://www.lesswrong.com/posts/FhKkFcojhKZt7nHzG/a-short-critique-of-vanessa-kosoy-s-predca-1
Will we run out of ML-data? https://www.lesswrong.com/posts/Couhhp4pPHbbhJ2Mg/will-we-run-out-of-ml-data-evidence-from-projecting-dataset
Black box agnostic models are worth still considering? https://www.lesswrong.com/posts/uXGLciramzNfb8Hvz/why-i-m-working-on-model-agnostic-interpretability
Instrumental convergence explains why general intelligence is possible https://www.lesswrong.com/posts/GZgLa5Xc4HjwketWe/instrumental-convergence-is-what-makes-general-intelligence

Opportunities:

  • AI impacts is still looking for a senior Research Analyst
  • And Anthropic is still looking for a senior software engineer
  • While Center of AI Safety is looking for a chief of staff
  • David Krueger’s lab is looking for collaborators

ml #opportunities #safety #engineers

View Details

The crypto giant FTX crashes, introducing massive uncertainty in the funding space for AI safety, humans cooperate better with lying AI, and interpretability is promising but also not.

Content

00:25 Uncertainty for AI Safety Funding

1:20 Human-AI cooperation

2:30 Interpretability in the wild

3:25 Other news

Opportunities

  • CHAI is offering an AI Research Internship under one of their mentors, https://ais.pub/d0de20
  • Today is the day the interpretability hackathon starts, open to all, https://ais.pub/alignmentjam
  • AI impacts is looking for a senior Research Analyst, https://ais.pub/aiimpactresearcher

Sources:

FTX fallout continues to roll out markets, https://www.youtube.com/watch?v=IgpbdnXOpEk

FTX has probably collapsed, https://forum.effectivealtruism.org/posts/tdLRvYHpfYjimwhyL/ftx-com-has-probably-collapsed

Human-AI cooperation is better when non-calibrated confidence, https://arxiv.org/pdf/2202.05983.pdf

Google shows inverse scaling can become U-shaped, https://www.alignmentforum.org/posts/LvKmjKMvozpdmiQhP/inverse-scaling-can-become-u-shaped
Ethan Perez calls them out on their methods, https://twitter.com/EthanJPerez/status/1588352204540235776

Interpretability in the wild,
https://arxiv.org/pdf/2211.00593.pdf

Eric Drexler discusses superintelligences with Eliezer (in the comments), https://www.alignmentforum.org/posts/HByDKLLdaWEcA2QQD/applying-superintelligence-without-collusion

Janus (an alias for several people) shows which places GPT-3 davinci-text-002 selects very specific outputs - favorite number, no answers, etc, https://www.alignmentforum.org/posts/t9svvNPNmFf5Qa3TA/mysteries-of-mode-collapse-due-to-rlhf

Mesa-optimizer implemented, https://www.alignmentforum.org/posts/b44zed5fBWyyQwBHL/trying-to-make-a-treacherous-mesa-optimizer

David Krueger argues against mechanistic interpretability, https://www.alignmentforum.org/posts/kjRGMdRxXb9c5bWq5/mechanistic-interpretability-as-reverse-engineering-follow

Nate Soares overviews strategies for “knowing AGI is safe”, https://www.alignmentforum.org/posts/iDFTmb8HSGtL4zTvf/how-could-we-know-that-an-agi-system-will-have-good

Interpretability starter resources, https://ais.pub/alignmentjam

View Details

This week, we look at broken scaling laws, surgical fine-tuning, interpretability in the wild, and threat models of AI.

Opportunities
Redwood Research is inviting 30-50 researchers to join them in Berkeley for a very interesting mechanistic interpretability research programme. https://ais.pub/remix
Anthropic is looking for operations managers, recruiters, researchers, engineers, and product managers: https://ais.pub/anthropic
AI Safety Ideas: https://ais.pub/aisi
Hackathons: https://ais.pub/jam

00:20 Broken scaling laws & surgical fine-tuning
01:34 Debate & interpretability
02:55 Threat models in ML safety
04:30 Other news
05:30 REMIX, Anthropic, and hackathons

Sources
Interpreting DeepMind’s in-context RL paper: https://www.lesswrong.com/posts/avvXAvGhhGgkJDDso/caution-when-interpreting-deepmind-s-in-context-rl-paper
Interpretability in the Wild: https://twitter.com/kevrowan/status/1587601532639494146
Precision machine learning, Max Tegmark: https://arxiv.org/pdf/2210.13447.pdf
Broken natural scaling laws: https://arxiv.org/pdf/2210.14891.pdf
Debate does not work for question-answering: https://arxiv.org/pdf/2210.10860.pdf
Vision for meta-science: https://scienceplusplus.org/metascience/index.html#how-does-the-culture-of-science-change
AI X-risk 35% mostly based on a recent peer-reviewed argument - LessWrong: https://www.lesswrong.com/posts/XtBJTFszs8oP3vXic/ai-x-risk-greater-than-35-mostly-based-on-a-recent-peer
Clarifying X-risk: https://www.lesswrong.com/posts/GctJD5oCDRxCspEaZ/clarifying-ai-x-risk
Threat model literature review: https://www.lesswrong.com/posts/wnnkD6P2k2TfHnNmt/threat-model-literature-review
Frames vs. boundaries: https://www.alignmentforum.org/posts/SZjHimszxqjNJzQWK/boundaries-vs-frames
Surgical fine-tuning improves robustness against distributional shifts: https://arxiv.org/pdf/2210.11466.pdf

View Details

This week, we look at how we can safeguard against AGI, look at new research on Goodhart’s law, see an open source dataset with 60,000 emotional videos, and share new opportunities in ML and AI safety.

  • $1.5 million in prizes for changing FTX’s mind: https://ais.pub/6wg
  • Interpretability research hackathon: https://ais.pub/s5m
  • The AI Safety Ideas platform: https://ais.pub/aisi
  • Apply to the long-term future fund for projects: https://ais.pub/ltff
  • Anthropic technical engineer: https://ais.pub/5f956b
  • AI Impacts research assistant: https://ais.pub/bjv
  • AI Impacts senior researcher: https://ais.pub/l9x
  • AI Impacts research analyst: https://ais.pub/xoz
  • Berkeley Existential Risks Initiative research assistant: https://ais.pub/jt9
  • Ought machine learning engineering internship: https://ais.pub/wns

Sources

  • Scaling laws for reward model overoptimization: https://arxiv.org/abs/2210.10760
  • Estimating wellbeing from video: https://arxiv.org/abs/2210.10039
  • Protecting world from out-of-control AGI: https://www.alignmentforum.org/posts/LFNXiQuGrar3duBzJ/what-does-it-take-to-defend-the-world-against-out-of-control
    • Pivotal acts: https://arbital.com/p/pivotal/
    • Strategy-stealing assumption: https://www.alignmentforum.org/posts/nRAMpjnb6Z4Qv3imF/the-strategy-stealing-assumption
  • Neel Nanda mechanistic interpretability learning guide: https://www.alignmentforum.org/posts/AaABQpuoNC8gpHf2n/a-barebones-guide-to-mechanistic-interpretability
  • Computational mechanics to understand mechanistic interpretability: https://www.alignmentforum.org/posts/kqxEJkq5Big9nNKxy/beyond-kolmogorov-and-shannon
  • Distilled representations agenda: https://www.alignmentforum.org/posts/wjQkQ8bgWWFym8zF9/distilled-representations-research-agenda-1
  • POWERplay RL environment for agentic power-seeking and modeling instrumental value of positions: https://github.com/gladstoneai/POWERplay

View Details

🎉 You can now subscribe to our newsletter and listen to these updates in your favorite podcasting app. Check out newsletter.apartresearch.com and podcast.apartresearch.com.

This week, we’re looking at counterarguments to the basic case for why AI is an existential risk to humanity, looking at how strong AI might come very soon, and sharing interesting papers.

  • Apply for the SERI MATS programme: https://ais.pub/serimats
  • Apply for a fellowship from FLI for your PhD or postdoc: https://ais.pub/fli
  • Join Redwood’s technical staff: https://ais.pub/redwoodjob
  • Join CHAI’s research staff: https://ais.pub/chaiintern
  • Visit the Alignment Jam page: https://ais.pub/jam
  • Sign up for our newsletters: https://ais.pub/news
  • Visit our podcast: https://ais.pub/pod

Sources:

  • X-risk counter arguments: https://www.alignmentforum.org/posts/LDRQ5Zfqwi8GjzPYG/counterarguments-to-the-basic-ai-x-risk-case
    • Response 1: https://www.alignmentforum.org/posts/GQat3Nrd9CStHyGaq/response-to-katja-grace-s-ai-x-risk-counterarguments
    • Conversation: https://www.alignmentforum.org/posts/iXuJLARFBZbaBGxW3/a-conversation-about-katja-s-counterarguments-to-ai-risk
    • Comprehensive AI services: https://www.lesswrong.com/posts/x3fNwSe5aWZb5yXEG
  • Strong general AI coming soon: https://www.lesswrong.com/posts/K4urTDkBbtNuLivJx/why-i-think-strong-general-ai-is-coming-soon
  • Nate’s decision theory != nice things: https://www.alignmentforum.org/posts/rP66bz34crvDudzcJ/decision-theory-does-not-imply-that-we-get-to-have-nice
  • Moral RL agent: https://openreview.net/pdf?id=CtS2Rs_aYk
  • Non-experts’ perspective on interpretability (high-stakes, otherwise accuracy): https://www.nature.com/articles/s41467-022-33417-3
  • Neel Nanda’s interpretability paper list: https://www.lesswrong.com/posts/SfPrNY45kQaBozwmu/an-extremely-opinionated-annotated-list-of-my-favourite
  • John’s bid for prediction markets: https://www.lesswrong.com/posts/oxSX9XDQHLu5YLpaD/how-to-make-prediction-markets-useful-for-alignment-work
  • Physics on LLMs: https://arxiv.org/pdf/2210.05359.pdf
  • Deep learning theory: https://www.lesswrong.com/posts/bumgqvRjTadFFkoAd/science-of-deep-learning-a-technical-agenda

View Details

✉️ Read the ML Safety Newsletter: https://ais.pub/mlsn6
📈 Help Center for AI Safety come up with new benchmarks: https://ais.pub/benchmarks
💡 Help Redwood Research find interesting heuristics: https://ais.pub/gptheuristics
👩‍🔬 Join the hackathon: https://ais.pub/interpretability
🙋‍♀️ Register as a local organizer: https://ais.pub/localorganizer
🤓 Read the Alignment 201 curriculum: https://ais.pub/alignment201

[Corrections] Marcus = Marius. AGI Safety Fundamentals is not taking applicants but you can read the Alignment 201 curriculum here:
https://ais.pub/alignment201.

00:00 Intro
00:19 Law defines alignment principles
00:40 Out-of-distribution alignment
01:27 Reward hacking defined
02:38 Inductive biases in learning algorithms
03:34 Warning shots are not enough
04:04 State of AI safety
04:45 Announcements

Sources:
Legal informatics for AI safety, robust specification https://arxiv.org/abs/2209.13020
Out-of-distribution GAN examples https://arxiv.org/abs/2209.11960
Formal definition of ‘reward hacking’ https://arxiv.org/abs/2209.13085
DeepMind: Why correct goals are not enough https://arxiv.org/abs/2210.01790
QAPR 4, inductive biases of learning processes https://www.alignmentforum.org/posts/...
QAPR 3: Training NNs from interpretability priors https://www.alignmentforum.org/s/5omS...
Neural tangent Kernel distillation https://www.alignmentforum.org/posts/...
OG paper: https://arxiv.org/pdf/1806.07572.pdf
Gaussian processes https://distill.pub/2019/visual-explo...
Soares’ critique of warning shots https://www.alignmentforum.org/posts/...
~300 people in AIS https://forum.effectivealtruism.org/p...
Statistics of machine learning https://financesonline.com/machine-le...
Chatting AI safety with 100+ researchers https://www.alignmentforum.org/posts/...
Smaller news
MLSN https://www.alignmentforum.org/posts/...
Safety benchmarks prize https://benchmarking.mlsafety.org/
Finding heuristics of GPT-2 small https://www.lesswrong.com/posts/LkBmA...
Alignment 201 curriculum: https://www.agisafetyfundamentals.com...

View Details

[correction] Eliezer kindly reached out to us to note that this is not the correct criticism from MIRI of CHAI's position. Here it is in his own words:

The basic argument goes that an AI whose utility function has been made dependent on hidden information, even if that information is inside humans, won't defer to humans because of that; it gets all the information that's obtainable and then ignores the humans (and kills them). There's never a point where "let the humans shut me off and build another AI" looks like a better strategy than "get all the info out of the humans and then stop listening".

Welcome to this week's Safe AI Progress Report where Thomas describes the scary developments in AI, a theoretical dispute, and risks from power-seeking AI.

00:00 Intro
00:18 Carmack's AGI
00:36 Video generation
00:57 AlphaTensor
01:26 PyTorch to Linux Foundation
02:00 Risks from power-seeking AI
03:09 MIRI criticizes CHAI's approach
03:55 Smaller news
04:35 Interpretability hackathon

Sources:
Capabilities updates
$20 million to develop AGI: https://www.insiderintelligence.com/c...
John Carmack doesn’t care about safety: https://twitter.com/ID_AA_Carmack/sta...
Meta’s GAN: https://ai.facebook.com/blog/generati...
Meta’s video gan is actually bad: https://phenaki.video/#interactive:~:...
Whisper: https://cdn.openai.com/papers/whisper...
People give the AI access to the internet: https://twitter.com/sergeykarayev/sta...
DeepMind’s AlphaMath matrix multiplication: https://www.deepmind.com/blog/discove...
Quanta article: https://www.quantamagazine.org/mathem...
Alignment researcher chatter: https://www.alignmentforum.org/posts/...
Grace
PyTorch to Linux: https://pytorch.org/blog/PyTorchfound...
Neutrality first (still better than Meta): https://www.linuxfoundation.org/blog/...
Yann LeCunn developing AGI: https://openreview.net/pdf?id=BZ5a1r-...
Paperclip clicker games: https://paperclips.tech/
Review of AI risk outlook: https://www.alignmentforum.org/posts/...
Joe Carlsmith’s report: https://arxiv.org/abs/2206.13353
Video lecture of the report: https://forum.effectivealtruism.org/p...
Endgames: https://www.alignmentforum.org/posts/...
Machine alignment Monday by Scott Alexander: https://astralcodexten.substack.com/p...
Article discussed: https://arbital.com/p/updated_deference/
Tamsin Leake’s outlook on AI safety: https://carado.moe/outlook-ai-risk-mi...
Alex’s loss functions: https://www.alignmentforum.org/posts/...
Physics-based deep learning: https://physicsbaseddeeplearning.org/...
Andrej Karpathy’s first video: https://www.youtube.com/watch?v=VMj-3...
Amazing interpretability tool from Redwood: http://interp-tools.redwoodresearch.org/
Tutorial for use: https://docs.google.com/document/d/1E...
Anthropic’s Svelte work: https://github.com/anthropics/PySvelte
Microscope: https://openai.com/blog/microscope/
Join the hackathon, Esben’s interpretability talk: https://itch.io/jam/interpretability

View Details

Join the hackathon GatherTown here: https://app.gather.town/app/iEPL2kx1L...
Or read more about the event here: https://itch.io/jam/llm-hackathon

👩‍🔬 This week's Safe AI Progress Report puts the focus on the Future Fund's new prize, Conjecture's new research, speed priors, and much more. Come along!

Chapters:
00:00 Intro
00:16 $1.5 million prize
01:00 Conjecture's interpretability research
01:50 Anthropic's superposition experiments
02:30 Inverse Scaling first round prizes
03:08 Forwarding speed regularizers
04:00 Reward is not the optimization target
04:25 Other news
04:52 🎉 Apart Hackathon
05:08 Israel AI Safety
05:20 Outro

Sources:
Future Fund world view competition: https://ftxfuturefund.org/
Strong general AGI soon: https://forum.effectivealtruism.org/p...
Polytopes lens: https://www.alignmentforum.org/posts/...
Anthropic’s papers: https://www.anthropic.com/research
Toy models of superposition: https://transformer-circuits.pub/2022...
Inverse scaling prize round 1: https://www.alignmentforum.org/posts/...
Inverse scaling prize: https://github.com/inverse-scaling/prize
Speed prior and forwarding speed priors: https://www.alignmentforum.org/posts/...
Are minimal circuits deceptive?: https://www.lesswrong.com/posts/fM5ZW...
Musings on the speed prior: https://www.alignmentforum.org/posts/...
Deconfusing wireheading: https://www.alignmentforum.org/posts/...
Reward is not the optimization target: https://www.alignmentforum.org/posts/...
Nearcasting AGI: https://www.alignmentforum.org/posts/...
7 traps new alignment researchers drop into: https://www.lesswrong.com/posts/h5CGM...
Language model hackathon: https://itch.io/jam/llm-hackathon
AI Safety Israel conference: https://aisic2022.net.technion.ac.il/
Apart Research: https://apartresearch.com
AI Safety Ideas: https://aisi.ai

View Details

👩‍🔬 Hosted by Sabrina Zaki

This Safe AI Progress Report describes the past week's developments in ML and AI safety. Follow along to get regular updates for the scientific field to stay safe from artificial intelligence.

Citations

  • Reinforcement learning from human feedback: https://twitter.com/anthropicai/status/1514277273070825476?lang=en
  • First SAIPR: https://www.youtube.com/watch?v=ETknJbbL3PY&t=5s&ab_channel=ApartResearch
  • Red teaming LLMs: https://arxiv.org/pdf/2202.03286.pdf
  • Adversarial training [Redwood]: https://arxiv.org/abs/2205.01663
  • Robust injury classifier [Redwood]: https://www.alignmentforum.org/posts/n3LAgnHg6ashQK3fF/takeaways-from-our-robust-injury-classifier-project-redwood
    • Original intent: https://www.alignmentforum.org/posts/k7oxdbNaGATZbtEg3/redwood-research-s-current-project
    • Original paper: https://arxiv.org/abs/2205.01663
    • Surge AI: https://www.surgehq.ai/case-study/adversarial-testing-redwood-research
  • Red teaming language models to reduce harms: Review [Anthropic]: https://arxiv.org/abs/2209.07858
  • Aligning language models: https://arxiv.org/abs/2209.00731
  • Refine’s third blog post battery: https://www.alignmentforum.org/posts/PhKSe9BT4h5peqrHL/refine-s-third-blog-post-day-week
    • Refine as a concept: https://www.alignmentforum.org/posts/5uiQkyKdejX3aEHLM/how-to-diversify-conceptual-alignment-the-model-behind
    • Ordering capability thresholds: https://www.alignmentforum.org/posts/ttRyu8u9vqX3jZFjr/ordering-capability-thresholds
    • Levels of goals and alignment: https://www.alignmentforum.org/posts/rzkCTPnkydQxfkZsX/levels-of-goals-and-alignment
    • Representational tether: https://www.alignmentforum.org/posts/h7BA7TQTo3dxvYrek/representational-tethers-tying-ai-latents-to-human-ones
  • Coordinate-free interpretability theory: https://www.alignmentforum.org/posts/sxhfSBej6gdAwcn7X/coordinate-free-interpretability-theory
    • Privileged basis: https://www.alignmentforum.org/posts/sxhfSBej6gdAwcn7X/coordinate-free-interpretability-theory?commentId=TiCE2Ai3LCdD7mvAw
    • Softmax Linear Units: https://transformer-circuits.pub/2022/solu/index.html
  • Backdoor bench: https://arxiv.org/abs/2206.12654
  • Leon Lang’s summary of AGISF readings: https://www.alignmentforum.org/posts/eymFwwc6jG9gPx5Zz/summaries-alignment-fundamentals-curriculum
  • Vanessa Kosoy’s ALTER prize for learning-theoretic progress in alignment: https://www.alignmentforum.org/posts/8BL7w55PS4rWYmrmv/prize-and-fast-track-to-alignment-research-at-alter
  • Apart Research: https://apartresearch.com

AI Safety Ideas: https://aisafetyideas.com

View Details

🏛 Citations:
Circuits: https://distill.pub/2020/circuits/zoom-in/ 
Interpretability survey: https://arxiv.org/abs/2207.13243, see Twitter summary: https://twitter.com/StephenLCasper/status/1569401262558576642, and PDF: https://arxiv.org/pdf/2207.13243.pdf 
Activation atlas: https://distill.pub/2019/activation-atlas/
Changing training data: https://arxiv.org/pdf/1811.12231.pdf
Editing factual associations in GPT: https://arxiv.org/pdf/2202.05262.pdf 
Natural language descriptions of deep visual features: https://arxiv.org/pdf/2201.11114.pdf
Robust feature-level adversaries are interpretability tools: https://arxiv.org/pdf/2110.03605.pdf
Samotsvety’s AI risk forecast: https://forum.effectivealtruism.org/posts/EG9xDM8YRz4JN4wMN/samotsvety-s-ai-risk-forecasts
Date of AGI: https://www.metaculus.com/questions/5121/date-of-artificial-general-intelligence/ 
(June) Forecasting TAI with biological anchors summary: https://www.lesswrong.com/s/B9Qc8ifidAtDpsuu8/p/wgio8E758y9XWsi8j
Monitoring for deceptive alignment: https://www.alignmentforum.org/posts/Km9sHjHTsBdbgwKyi/monitoring-for-deceptive-alignment 
Deceptive alignment: https://www.alignmentforum.org/posts/zthDPAjh9w6Ytbeks/deceptive-alignment 
Quintin’s alignment paper lineup: https://www.lesswrong.com/posts/7cHgjJR2H5e4w4rxT/quintin-s-alignment-papers-roundup-week-1
Most people start with the same few bad ideas: https://www.lesswrong.com/posts/Afdohjyt6gESu4ANf/most-people-start-with-the-same-few-bad-ideas
Beth Barnes starting a risks and development evaluations group at ARC: https://www.alignmentforum.org/posts/svhQMdsefdYFDq5YM/evaluations-project-arc-is-hiring-a-researcher-and-a-webdev-1
Cognitive biases in LLMs: https://arxiv.org/pdf/2206.14576.pdf
Academia vs. industry: https://www.alignmentforum.org/posts/HXxHcRCxR4oHrAsEr/an-update-on-academia-vs-industry-one-year-into-my-faculty

View Details

🥳 Welcome to this Safe AI Progress Report! It is the first of a regular series of videos that summarize what has happened in AI safety since the last report. They accompany a different progress report that focuses on measurable metrics for the field of AI safety.

📃Script
00:00 - Intro
00:13 - OpenAI alignment strategy
00:53 - Strategy explanation
01:50 - John Wentworth's criticism
02:37 - Sam Bowman's NLP survey
03:17 - New perspectives | Simulators
03:43 - New perspectives | Shard theory
04:10 - The Small Side
04:36 - Outro | Learn more

🏛 Chronological references:
- OpenAI dangerous: https://techcrunch.com/2019/02/17/ope...
- OpenAI = dangerous: https://www.lesswrong.com/posts/Nqn2t...
- OpenAI alignment: https://openai.com/blog/our-approach-...
- Jacob Hilton: https://www.lesswrong.com/posts/3S4ny...
- RLHF: https://arxiv.org/abs/2009.01325
- IDA: https://forum.effectivealtruism.org/p...
- Alignment research tools: https://www.lesswrong.com/posts/ebYio...
- Elicit: https://elicit.org/search?q=AI+safety...
- Eleuther network: https://arxiv.org/pdf/2206.02841.pdf
- Not everyone happy: https://www.lesswrong.com/posts/3S4ny...
- Grokking: https://arxiv.org/pdf/2201.02177.pdf
- Deception: https://www.lesswrong.com/posts/zthDP...
- RLHF is terrible: https://www.lesswrong.com/posts/xFotX...
- Survey of NLP researchers: https://nlpsurvey.net/nlp-metasurvey-...
- Copilot good: https://github.blog/2022-09-07-resear...
- Programs for itself: https://arxiv.org/abs/2207.14502
- Simulators: https://www.lesswrong.com/posts/vJFdj...
- Shard theory: https://www.lesswrong.com/posts/iCfdc...
- Richard’s list: https://www.lesswrong.com/posts/27AWR...
- Thomas & Eli’s list: https://www.lesswrong.com/posts/QBAjn...
- Philosophy fellowship: https://philosophy.safe.ai/
- ML Safety course material: https://course.mlsafety.org/
- ML safety competitions: https://safe.ai/competitions
- Apart Research: https://apartresearch.com
- AI safety ideas: https://aisafetyideas.com/