https://forum.effectivealtruism.org/posts/zoWypGfXLmYsDFivk/counterarguments-to-the-basic-ai-risk-case

This is cross-posted from the AI Impacts blog
This is going to be a list of holes I see in the basic argument for existential risk from superhuman AI systems1.

To start, here’s an outline of what I take to be the basic case2:

I. If superhuman AI systems are built, any given system is likely to be ‘goal-directed’
Reasons to expect this:

  1. Goal-directed behavior is likely to be valuable, e.g. economically.
  2. Goal-directed entities may tend to arise from machine learning training processes not intending to create them (at least via the methods that are likely to be used).
  3. ‘Coherence arguments’ may imply that systems with some goal-directedness will become more strongly goal-directed over time.

II. If goal-directed superhuman AI systems are built, their desired outcomes will probably be about as bad as an empty universe by human lights---
narrator_time: TBC
editing_time: TBC
narrator: pw
editor: pw
qa: km
client: ea_forum
project_id: TBC