client: 80000_hours
project_id: articles
narrator: pw
qa: km
qa_time: 0h30m
As the 2016 US presidential campaign was entering a fractious round of primaries, Hillary Clinton’s campaign chair, John Podesta, opened a disturbing email. The March 19 message warned that his Gmail password had been compromised and that he urgently needed to change it.
The email was a lie. It wasn’t trying to help him protect his account — it was a phishing attack trying to gain illicit access.
Podesta was suspicious, but the campaign’s IT team erroneously wrote the email was “legitimate” and told him to change his password. The IT team provided a safe link for Podesta to use, but it seems he or one of his staffers instead clicked the link in the forged email. That link was used by Russian intelligence hackers known as “Fancy Bear,” and they used their access to leak private campaign emails for public consumption in the final weeks of the 2016 race, embarrassing the Clinton team.
While there are plausibly many critical factors in any close election, it’s possible that the controversy around the leaked emails played a non-trivial role in Clinton’s subsequent loss to Donald Trump. This would mean the failure of the campaign’s security team to prevent the hack — which might have come down to a mere typo — was extraordinarily consequential.
Source:https://80000hours.org/career-reviews/information-security
Narrated for 80,000 Hours by TYPE III AUDIO.
Share feedback on this narration.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
TL;DR: To have a fulfilling career, get good at something and then use it to tackle pressing global problems.
Rather than expect to discover your passion in a flash of insight, your job satisfaction will grow over time as you learn more about what kind of work fits you, master valuable skills, and use them to find engaging work that helps others.
Source:https://80000hours.org/career-guide/summary/
Narrated for 80,000 Hours by Perrin Walker of TYPE III AUDIO.
Share feedback on this narration.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
You’ll spend about 80,000 hours working in your career: 40 hours a week, 50 weeks a year, for 40 years. So how to spend that time is one of the most important decisions you’ll ever make.
Choose wisely, and you will not only have a more rewarding and interesting life — you’ll also be able to help solve some of the world’s most pressing problems. But how should you choose?
To answer this question, we set up an independent nonprofit and have done over 10 years of research alongside Oxford academics. Our only aim is to help you have the greatest possible positive impact.
Along the way, we’ve discovered some surprising things, and over 10 million people have viewed our advice.
Source:https://80000hours.org/career-guide/introduction/
Narrated for 80,000 Hours by Perrin Walker of TYPE III AUDIO.
Share feedback on this narration.
client: ea_forum
project_id: curated
narrator: pw
qa: km
qa_time: 0h20m
---This post is a summary of some of my work as a field strategy consultant at Schmidt Futures' Act 2 program, where I spoke with over a hundred experts and did a deep dive into antimicrobial resistance to find impactful investment opportunities within the cause area. The full report can be accessed here.
Antimicrobials, the medicines we use to fight infections, have played a foundational role in improving the length and quality of human life since penicillin and other antimicrobials were first developed in the early and mid 20th century.
Antimicrobial resistance, or AMR, occurs when bacteria, viruses, fungi, and parasites evolve resistance to antimicrobials. As a result, antimicrobial medicine such as antibiotics and antifungals become ineffective and unable to fight infections in the body.
AMR is responsible for millions of deaths each year, more than HIV or malaria (ARC 2022). The AMR Visualisation Tool, produced by Oxford University and IHME, visualises IHME data which finds that 1.27 million deaths per year are attributable to bacterial resistance and 4.95 million deaths per year are associated with bacterial resistance, as shown below.
Source:
https://forum.effectivealtruism.org/posts/W93Pt7xch7eyrkZ7f/cause-area-report-antimicrobial-resistance
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
Share feedback on this narration.
client: ea_forum
project_id: curated
narrator: pw
qa: km
narrator_time: 1h15m
qa_time: 30m
---Charity Entrepreneurship is frequently contacted by individuals and donors who like our model. Several have expressed interest in seeing the model expanded, or seeing what a twist on the model would look like (e.g., different cause area, region, etc.) Although we are excited about maximizing CE’s impact, we are less convinced by the idea of growing the effective charity pool via franchising or other independent nonprofit incubators. This is because new incubators often do not address the actual bottlenecks faced by the nonprofit landscape, as we see them. There are lots of factors that prevent great new charities from being launched, and from eventually having a large impact. We have scaled CE to about 10 charities a year, and from our perspective, these are the three major bottlenecks to growing the new charity ecosystem further: Mid-stage funding, Founders and Multiplying effects.
Source:
https://forum.effectivealtruism.org/posts/ckokr9uhr2Cu3h5En/tips-for-people-considering-starting-new-incubators
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
Share feedback on this narration.
client: lesswrong
project_id: articles
feed_id: ai ai_safety
narrator: pw
qa: km
narrator_time: 4h30m
qa_time: 2h0m
---Philosopher David Chalmers asked: "Is there a canonical source for "the argument for AGI ruin" somewhere, preferably laid out as an explicit argument with premises and a conclusion?"
Unsurprisingly, the actual reason people expect AGI ruin isn't a crisp deductive argument; it's a probabilistic update based on many lines of evidence. The specific observations and heuristics that carried the most weight for someone will vary for each individual, and can be hard to accurately draw out. That said, Eliezer Yudkowsky's So Far: Unfriendly AI Edition might be a good place to start if we want a pseudo-deductive argument just for the sake of organizing discussion. People can then say which premises they want to drill down on. In The Basic Reasons I Expect AGI Ruin, I wrote: "When I say "general intelligence", I'm usually thinking about "whatever it is that lets human brains do astrophysics, category theory, etc. even though our brains evolved under literally zero selection pressure to solve astrophysics or category theory problems". It's possible that we should already be thinking of GPT-4 as "AGI" on some definitions, so to be clear about the threshold of generality I have in mind, I'll specifically talk about "STEM-level AGI", though I expect such systems to be good at non-STEM tasks too. STEM-level AGI is AGI that has "the basic mental machinery required to do par-human reasoning about all the hard sciences", though a specific STEM-level AGI could (e.g.) lack physics ability for the same reasons many smart humans can't solve physics problems, such as "lack of familiarity with the field".
Source:
https://www.lesswrong.com/posts/QzkTfj4HGpLEdNjXX/an-artificially-structured-argument-for-expecting-agi-ruin
Narrated for LessWrong by TYPE III AUDIO.
Share feedback on this narration.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 2h45m
qa_time:1h00m
---Thanks to Drake Thomas for feedback. I. Here’s a fun scatter plot. It has two thousand points, which I generated as follows: first, I drew two thousand x-values from a normal distribution with mean 0 and standard deviation 1. Then, I chose the y-value of each point by taking the x-value and then adding noise to it. The noise is also normally distributed, with mean 0 and standard deviation 1. Notice that there’s more spread along the y-axis than along the x-axis. That’s because each y-coordinate is a sum of two independently drawn numbers from the standard normal distribution. Because variances add, the y-values have variance 2 (standard deviation 1.41), not 1. Statisticians often talk about data forming an “elliptical cloud”.
Original text:
https://www.lesswrong.com/posts/nnDTgmzRrzDMiPF9B/how-much-do-you-believe-your-results
Narrated for LessWrong by TYPE III AUDIO.
Share feedback on this narration.
client: ea_forum
project_id: summaries
narrator: cs
We've just passed the half-year mark for this project! If you're reading this, please consider taking this 5 minute survey — all questions optional. If you listen to the podcast version, we have a separate survey for that here. Thanks to everyone that has responded to this already!
Original text:https://forum.effectivealtruism.org/posts/9QcmyGAjERHRFfrr7/summaries-of-top-forum-posts-1st-to-7th-may-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
Share feedback on this narration.
client: ea_forum
project_id: curated
feed_id: ai_safety
narrator: jc
How worried about AI risk will we feel in the future, when we can see advanced machine intelligence up close? We should worry accordingly now.
Original article:
https://joecarlsmith.com/2023/05/08/predictable-updating-about-ai-risk
Narrated by Joe Carlsmith and included on the Effective Altruism Forum by TYPE III AUDIO.
Share feedback on this narration.
client: ea_forum
project_id: summaries
narrator: cs
We've just passed the half-year mark for this project! If you're reading this, please consider taking this 5 minute survey — all questions optional. If you listen to the podcast version, we have a separate survey for that here. Thanks to everyone that has responded to this already!
Original text:https://forum.effectivealtruism.org/posts/wzn7hEj3BSz7us7ge/summaries-of-top-forum-posts-24th-30th-april-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: ea_forum
project_id: curated
narrator: pw
qa: km
qa_time: 0h35m
People often ask me for career advice related to AGI safety. This post summarizes the advice I most commonly give. I’ve split it into three sections: general mindset, alignment research and governance work. For each of the latter two, I start with high-level advice aimed primarily at students and those early in their careers, then dig into more details of the field. See also this post I wrote two years ago, containing a bunch of fairly general career advice. ## General mindset In order to have a big impact on the world you need to find a big lever. This document assumes that you think, as I do, that AGI safety is the biggest such lever. There are many ways to pull on that lever, though—from research and engineering to operations and field-building to politics and communications. I encourage you to choose between these based primarily on your personal fit—a combination of what you're really good at and what you really enjoy. In my opinion the difference between being a great versus a mediocre fit swamps other differences in the impactfulness of most pairs of AGI-safety-related jobs.
Original article:
https://forum.effectivealtruism.org/posts/xg7gxsYaMa6F3uH8h/agi-safety-career-advice
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/m2Y6HheC2Q2GLQ3oS/summaries-of-top-forum-posts-17th-23rd-april-2023
This podcast has just passed the 6-month mark! Please give us your feedback and suggestions so we can continue to improve — the survey should take no more than 10 minutes, and we really appreciate your input!
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: ea_forum
project_id: curated
narrator: pw
qa: km
qa_time: 0h45m
The excellent report from Rethink Priorities was my main source for this. Many of the substantial points I make are taken from it, though errors are my own. It’s worth reading! The authors are Gavriel Kleinwaks, Alastair Fraser-Urquhart, Jam Kraprayoon, and Josh Morrison.
Clean water
In the mid 19th century, London had a sewage problem. It relied on a patchwork of a few hundred sewers, of brick and wood, and hundreds of thousands of cesspits. The Thames — Londoners’ main source of drinking water — was near-opaque with waste. Here is Michael Faraday in an 1855 letter to The Times:
"Near the bridges the feculence rolled up in clouds so dense that they were visible at the surface even in water of this kind […] The smell was very bad, and common to the whole of the water. It was the same as that which now comes up from the gully holes in the streets. The whole river was for the time a real sewer […] If we neglect this subject, we cannot expect to do so with impunity; nor ought we to be surprised if, ere many years are over, a season give us sad proof of the folly of our carelessness."
Original article:
https://forum.effectivealtruism.org/posts/WLok4YuJ4kfFpDRTi/first-clean-water-now-clean-air
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
qa_time: 0h05m
This is a linkpost for a new 80,000 hours episode focused on how to engage in climate from an effective altruist perspective.
Rob and I are having a pretty wide-ranging conversation, here are the things we cover which I find most interesting for different audiences:
Original article:https://forum.effectivealtruism.org/posts/A3ZLLanDZZt9sgGQ9/new-80-000-hours-podcast-on-high-impact-climate-philanthropy
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
qa_time: 1h00m
---This is a post about mental health and disposition in relation to the alignment problem. It compiles a number of resources that address how to maintain wellbeing and direction when confronted with existential risk.
Many people in this community have posted their emotional strategies for facing Doom after Eliezer Yudkowsky’s “Death With Dignity” generated so much conversation on the subject. This post intends to be more touchy-feely, dealing more directly with emotional landscapes than questions of timelines or probabilities of success.
The resources section would benefit from community additions. Please suggest any resources that you would like to see added to this post.
Please note that this document is not intended to replace professional medical or psychological help in any way. Many preexisting mental health conditions can be exacerbated by these conversations. If you are concerned that you may be experiencing a mental health crisis, please consult a professional.
Original article:
https://www.lesswrong.com/posts/pLLeGA7aGaJpgCkof/mental-health-and-the-alignment-problem-a-compilation-of#
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/o3Gaoizs2So6SpgLH/summaries-of-top-forum-posts-27th-march-to-16th-april
This podcast has just passed the 6-month mark! Please give us your feedback and suggestions so we can continue to improve — the survey should take no more than 10 minutes, and we really appreciate your input!
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
qa_time: 0h15m
I recently spent some time reflecting on my career and my life, for a few reasons:
I wanted to have a better answer to these questions:
Original article:https://forum.effectivealtruism.org/posts/2DzLY6YP2z5zRDAGA/a-freshman-year-during-the-ai-midgame-my-approach-to-the
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
feed_id: ai_alignment, ai
narrator: pw
qa: km
qa_time: 0h50m
The primary talk of the AI world recently is about AI agents (whether or not it includes the question of whether we can’t help but notice we are all going to die.)
The trigger for this was AutoGPT, now number one on GitHub, which allows you to turn GPT-4 (or GPT-3.5 for us clowns without proper access) into a prototype version of a self-directed agent.
We also have a paper out this week where a simple virtual world was created, populated by LLMs that were wrapped in code designed to make them simple agents, and then several days of activity were simulated, during which the AI inhabitants interacted, formed and executed plans, and it all seemed like the beginnings of a living and dynamic world. Game version hopefully coming soon.
How should we think about this? How worried should we be?
Original article:
https://www.lesswrong.com/posts/566kBoPi76t8KAkoD/on-autogpt
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: articles
feed_id: ai_safety
narrator: pw
qa: mds
qa_time: 0h15m
This is a linkpost for https://www.forourposterity.com/want-to-win-the-agi-race-solve-alignment/
Society really cares about safety. Practically speaking, the binding constraint on deploying your AGI could well be your ability to align your AGI. Solving (scalable) alignment might be worth lots of $$$ and key to beating China.
Look, I really don't want Xi Jinping Thought to rule the world. If China gets AGI first, the ensuing rapid AI-powered scientific and technological progress could well give it a decisive advantage (cf potential for >30%/year economic growth with AGI). I think there's a very real specter of global authoritarianism here.
Or hey, maybe you just think AGI is cool. You want to go build amazing products and enable breakthrough science and solve the world’s problems.
So, race to AGI with reckless abandon then? At this point, people get into agonizing discussions about safety tradeoffs. And many people just mood affiliate their way to an answer: "accelerate, progress go brrrr," or "AI scary, slow it down."
I see this much more practically. And, practically, society cares about safety, a lot. Do you actually think that you’ll be able to and allowed to deploy an AI system that has, say, a 10% chance of destroying all of humanity?
Original article:
https://forum.effectivealtruism.org/posts/Ackzs8Wbk7isDzs2n/want-to-win-the-agi-race-solve-alignment
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
feed_id: ai_safety
narrator: pw
qa: km
qa_time: 0h10m
(Related text posted to Twitter; this version is edited and has a more advanced final section.)
Imagine yourself in a box, trying to predict the next word - assign as much probability mass to the next token as possible - for all the text on the Internet.
Koan: Is this a task whose difficulty caps out as human intelligence, or at the intelligence level of the smartest human who wrote any Internet text? What factors make that task easier, or harder? (If you don't have an answer, maybe take a minute to generate one, or alternatively, try to predict what I'll say next; if you do have an answer, take a moment to review it inside your mind, or maybe say the words out loud.)
Original article:https://www.lesswrong.com/posts/nH4c3Q9t9F3nJ7y8W/gpts-are-predictors-not-imitators
Narrated for LessWrong by TYPE III AUDIO.
client: t3a
feed_id: ai_safety_abstracts
narrator: ai
This episode covers 3 abstracts:
client: lesswrong
project_id: curated
narrator: pw
qa: mds
qa_time: 0h30m
---(This is a stylized version of a real conversation, where the first part happened as part of a public debate between John Wentworth and Eliezer Yudkowsky, and the second part happened between John and me over the following morning. The below is combined, stylized, and written in my own voice throughout. The specific concrete examples in John's part of the dialog were produced by me. It's over a year old. Sorry for the lag.)
(As to whether John agrees with this dialog, he said "there was not any point at which I thought my views were importantly misrepresented" when I asked him for comment.)
Original article: https://www.lesswrong.com/posts/fJBTRa7m7KnCDdzG5/a-stylized-dialogue-on-john-wentworth-s-claims-about-markets
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
feed_id: ai_safety
narrator: pw
qa: mds
qa_time: 1h00m
In late 2022, Nate Soares gave some feedback on my Cold Takes series on AI risk (shared as drafts at that point), stating that I hadn't discussed what he sees as one of the key difficulties of AI alignment.
I wanted to understand the difficulty he was pointing to, so the two of us had an extended Slack exchange, and I then wrote up a summary of the exchange that we iterated on until we were both reasonably happy with its characterization of the difficulty and our disagreement.1 My short summary is:
Original article:https://www.lesswrong.com/posts/iy2o4nQj9DnQD7Yhj/discussion-with-nate-soares-on-a-key-alignment-difficulty
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
feed_id: ai_safety
narrator: pw
qa: mds
qa_time: 0h45m
This post is an attempt to gesture at a class of AI notkilleveryoneism (alignment) problem that seems to me to go largely unrecognized. E.g., it isn’t discussed (or at least I don't recognize it) in the recent plans written up by OpenAI (1,2), by DeepMind’s alignment team, or by Anthropic, and I know of no other acknowledgment of this issue by major labs.
You could think of this as a fragment of my answer to “Where do plans like OpenAI’s ‘Our Approach to Alignment Research’ fail?”, as discussed in Rob and Eliezer’s challenge for AGI organizations and readers. Note that it would only be a fragment of the reply; there's a lot more to say about why AI alignment is a particularly tricky task to task an AI with. (Some of which Eliezer gestures at in a follow-up to his interview on Bankless.)
Original article:
https://www.lesswrong.com/posts/XWwvwytieLtEWaFJX/deep-deceptiveness
Narrated for LessWrong by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h30m
It seems unlikely that humans are near the ceiling of possible intelligences, rather than simply being the first such intelligence that happened to evolve. Computers far outperform humans in many narrow niches (e.g. arithmetic, chess, memory size), and there is reason to believe that similar large improvements over human performance are possible for general reasoning, technology design, and other tasks of interest. As occasional AI critic Jack Schwartz (1987) wrote:
"If artificial intelligences can be created at all, there is little reason to believe that initial successes could not lead swiftly to the construction of artificial superintelligences able to explore significant mathematical, scientific, or engi-neering alternatives at a rate far exceeding human ability, or to generate plans and take action on them with equally overwhelming speed. Since man’s near-monopoly of all higher forms of intelligence has been one of the most basic facts of human existence throughout the past history of this planet, such developments would clearly create a new economics, a new sociology, and a new history."
Why might AI “lead swiftly” to machine superintelligence? Below we consider some reasons.
Original article:https://drive.google.com/file/d/1QxMuScnYvyq-XmxYeqBRHKz7cZoOosHr/view
Authors:Luke Muehlhauser, Anna Salamon
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h30m
AI is undergoing a paradigm shift with the rise of models (e.g., BERT, DALL-E, GPT-3) that are trained on broad data at scale and are adaptable to a wide range of downstream tasks. We call these models foundation models to underscore their critically central yet incomplete character. This report provides a thorough account of the opportunities and risks of foundation models, ranging from their capabilities (e.g., language, vision, robotics, reasoning, human interaction) and technical principles(e.g., model architectures, training procedures, data, systems, security, evaluation, theory) to their applications (e.g., law, healthcare, education) and societal impact (e.g., inequity, misuse, economic and environmental impact, legal and ethical considerations). Though foundation models are based on standard deep learning and transfer learning, their scale results in new emergent capabilities,and their effectiveness across so many tasks incentivizes homogenization. Homogenization provides powerful leverage but demands caution, as the defects of the foundation model are inherited by all the adapted models downstream. Despite the impending widespread deployment of foundation models, we currently lack a clear understanding of how they work, when they fail, and what they are even capable of due to their emergent properties. To tackle these questions, we believe much of the critical research on foundation models will require deep interdisciplinary collaboration commensurate with their fundamentally sociotechnical nature.
Original article:https://arxiv.org/abs/2108.07258
Authors:Bommasani et al.
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h30m
Many real world learning tasks involve complex or hard-to-specify objectives, and using an easier-to-specify proxy can lead to poor performance or misaligned behavior. One solution is to have humans provide a training signal by demonstrating or judging performance, but this approach fails if the task is too complicated for a human to directly evaluate. We propose Iterated Amplification, an alternative training strategy which progressively builds up a training signal for difficult problems by combining solutions to easier subproblems. Iterated Amplification is closely related to Expert Iteration (Anthony et al., 2017; Silver et al., 2017), except that it uses no external reward function. We present results in algorithmic environments, showing that Iterated Amplification can efficiently learn complex behaviors.
Original article:https://arxiv.org/abs/1810.08575
Authors:Paul Christiano, Buck Shlegeris, Dario Amodei
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h30m
The field of AI has undergone a revolution over the last decade, driven by the success of deep learning techniques. This post aims to convey three ideas using a series of illustrative examples:
I’ll focus on four domains: vision, games, language-based tasks, and science. The first two have more limited real-world applications, but provide particularly graphic and intuitive examples of the pace of progress.
Original article:https://medium.com/@richardcngo/visualizing-the-deep-learning-revolution-722098eb9c5
Author:Richard Ngo
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 1h00m
Within the coming decades, artificial general intelligence (AGI) may surpass human capabilities at a wide range of important tasks. We outline a case for expecting that, without substantial effort to prevent it, AGIs could learn to pursue goals which are undesirable (i.e. misaligned) from a human perspective. We argue that if AGIs are trained in ways similar to today's most capable models, they could learn to act deceptively to receive higher reward, learn internally-represented goals which generalize beyond their training distributions, and pursue those goals using power-seeking strategies. We outline how the deployment of misaligned AGIs might irreversibly undermine human control over the world, and briefly review research directions aimed at preventing this outcome.
Original article:https://arxiv.org/abs/2209.00626
Authors:Richard Ngo, Lawrence Chan, Sören Mindermann
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h15m
According to the orthogonality thesis, intelligent agents may have an enormous range of possible final goals. Nevertheless, according to what we may term the “instrumental convergence” thesis, there are some instrumental goals likely to be pursued by almost any intelligent agent, because there are some objectives that are useful intermediaries to the achievement of almost any final goal. We can formulate this thesis as follows:
The instrumental convergence thesis:
"Several instrumental values can be identified which are convergent in the sense that their attainment would increase the chances of the agent’s goal being realized for a wide range of final goals and a wide range of situations, implying that these instrumental values are likely to be pursued by a broad spectrum of situated intelligent agents."
Original article:https://drive.google.com/file/d/1KewDov1taegTzrqJ4uurmJ2CJ0Y72EU3/view
Author:Nick Bostrom
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h30m
The two tasks of supervised learning: regression and classification. Linear regression, loss functions, and gradient descent.
How much money will we make by spending more dollars on digital advertising? Will this loan applicant pay back the loan or not? What’s going to happen to the stock market tomorrow?
Original article:https://medium.com/machine-learning-for-humans/supervised-learning-740383a2feab
Author:Vishal Maini
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h15m
One step towards building safe AI systems is to remove the need for humans to write goal functions, since using a simple proxy for a complex goal, or getting the complex goal a bit wrong, can lead to undesirable and even dangerous behavior. In collaboration with DeepMind’s safety team, we’ve developed an algorithm which can infer what humans want by being told which of two proposed behaviors is better.
Original article:https://openai.com/research/learning-from-human-preferences
Authors:Dario Amodei, Paul Christiano, Alex Ray
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h30m
Specification gaming is a behaviour that satisfies the literal specification of an objective without achieving the intended outcome. We have all had experiences with specification gaming, even if not by this name. Readers may have heard the myth of King Midas and the golden touch, in which the king asks that anything he touches be turned to gold - but soon finds that even food and drink turn to metal in his hands. In the real world, when rewarded for doing well on a homework assignment, a student might copy another student to get the right answers, rather than learning the material - and thus exploit a loophole in the task specification.
Original article:https://www.deepmind.com/blog/specification-gaming-the-flip-side-of-ai-ingenuity
Authors:Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, Shane Legg
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: agi_sf
project_id: core_readings
feed_id: agi_sf__alignment
narrator: pw
qa: mds
qa_time: 0h15m
One approach to the AI control problem goes like this:
This approach has the major advantage that we can begin empirical work today — we can actually build systems which observe user behavior, try to figure out what the user wants, and then help with that. There are many applications that people care about already, and we can set to work on making rich toy models.
It seems great to develop these capabilities in parallel with other AI progress, and to address whatever difficulties actually arise, as they arise. That is, in each domain where AI can act effectively, we’d like to ensure that AI can also act effectively in the service of goals inferred from users (and that this inference is good enough to support foreseeable applications).
This approach gives us a nice, concrete model of each difficulty we are trying to address. It also provides a relatively clear indicator of whether our ability to control AI lags behind our ability to build it. And by being technically interesting and economically meaningful now, it can help actually integrate AI control with AI practice.
Overall I think that this is a particularly promising angle on the AI safety problem.
Original article:https://www.alignmentforum.org/posts/h9DesGT3WT9u2k7Hr/the-easy-goal-inference-problem-is-still-hard
Authors:Paul Christiano
Narrated for the AGI Safety Fundamentals Course by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: jc
In my last essay, I looked at two stories (brute preference for systematic-ness, and money-pumps) about why ethical anti-realists should still be interested in ethics – two stories about why the “philosophy game” is worth playing, even if there are no objective normative truths, and you’re free to do whatever you want. I think some versions of these stories might well have a role to play; but I find that on their own, they don’t fully capture what feels alive to me about ethics. Here I try to say something that gets closer.
Original article:
https://joecarlsmith.com/2023/02/17/seeing-more-whole
Narrated by Joe Carlsmith and included on the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/idpbfmPjHFCvzj46L/ea-and-lw-forum-weekly-summary-13th-19th-march-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: cea
project_id: articles
narrator: pw
qa: mds
qa_time: 1h15m
Effective altruism is a project that aims to find the best ways to help others, and put them into practice.
It’s both a research field, which aims to identify the world’s most pressing problems and the best solutions to them, and a practical community that aims to use those findings to do good.
This project matters because, while many attempts to do good fail, some are enormously effective. For instance, some charities help 100 or even 1,000 times as many people as others, when given the same amount of resources.
This means that by thinking carefully about the best ways to help, we can do far more to tackle the world’s biggest problems.
Original article:
https://www.effectivealtruism.org/articles/introduction-to-effective-altruism
Narrated for effectivealtruism.org by TYPE III AUDIO.
client: t3a
feed_id: ai ai_safety
In several of my books and many of my talks, I take great care to spell out just how special recent times have been, for most Americans at least. For my entire life, and a bit more, there have been two essential features of the basic landscape:
American hegemony over much of the world, and relative physical safety for Americans.
An absence of truly radical technological change.
Unless you are very old, old enough to have taken in some of WWII, or were drafted into Korea or Vietnam, probably those features describe your entire life as well.
In other words, virtually all of us have been living in a bubble “outside of history".
Now, circa 2023, at least one of those assumptions is going to unravel, namely #2. AI represents a truly major, transformational technological advance. Biomedicine might too, but for this post I’ll stick to the AI topic, as I wish to consider existential risk.
Original article:https://marginalrevolution.com/marginalrevolution/2023/03/existential-risk-and-the-turn-in-human-history.html
Narrated by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 7h00m
qa_time: 2h00m
Longtermists have argued that humanity should significantly increase its efforts to prevent catastrophes like nuclear wars, pandemics, and AI disasters. But one prominent longtermist argument overshoots this conclusion: the argument also implies that humanity should reduce the risk of existential catastrophe even at extreme cost to the present generation. This overshoot means that democratic governments cannot use the longtermist argument to guide their catastrophe policy. In this paper, we show that the case for preventing catastrophe does not depend on longtermism. Standard cost-benefit analysis implies that governments should spend much more on reducing catastrophic risk. We argue that a government catastrophe policy guided by cost-benefit analysis should be the goal of longtermists in the political sphere. This policy would be democratically acceptable, and it would reduce existential risk by almost as much as a strong longtermist policy.
Original article:
https://forum.effectivealtruism.org/posts/DiGL5FuLgWActPBsf/how-much-should-governments-pay-to-prevent-catastrophes
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: km
qa_time: 0h45m
Oral rehydration therapy is now the standard treatment for dehydration. It’s saved millions of lives, and can be prepared at home in minutes. So why did it take so long to discover?
Written by Matt Reynolds for Asterisk Magazine.
Original article:https://asteriskmag.com/issues/2/salt-sugar-water-zinc-how-scientists-learned-to-treat-the-20th-century-s-biggest-killer-of-children
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
qa_time: 0h30m
In addition to technical challenges, plans to safely develop AI face lots of organizational challenges. If you're running an AI lab, you need a concrete plan for handling that.
In this post, I'll explore some of those issues, using one particular AI plan as an example. I first heard this described by Buck at EA Global London, and more recently with OpenAI's alignment plan. (I think Anthropic's plan has a fairly different ontology, although it still ultimately routes through a similar set of difficulties)
I'd call the cluster of plans similar to this "Carefully Bootstrapped Alignment."
Original article:
https://www.lesswrong.com/posts/thkAtqoQwN6DtaiGT/carefully-bootstrapped-alignment-is-organizationally-hard
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
qa_time: 0h30m
This is a linkpost for https://evals.alignment.org/blog/2023-03-18-update-on-recent-evals/
[Written for more of a general-public audience than alignment-forum audience. We're working on a more thorough technical report.]
We believe that capable enough AI systems could pose very large risks to the world. We don’t think today’s systems are capable enough to pose these sorts of risks, but we think that this situation could change quickly and it’s important to be monitoring the risks consistently. Because of this, ARC is partnering with leading AI labs such as Anthropic and OpenAI as a third-party evaluator to assess potentially dangerous capabilities of today’s state-of-the-art ML models. The dangerous capability we are focusing on is the ability to autonomously gain resources and evade human oversight.
We attempt to elicit models’ capabilities in a controlled environment, with researchers in-the-loop for anything that could be dangerous, to understand what might go wrong before models are deployed. We think that future highly capable models should involve similar “red team” evaluations for dangerous capabilities before the models are deployed or scaled up, and we hope more teams building cutting-edge ML systems will adopt this approach. The testing we’ve done so far is insufficient for many reasons, but we hope that the rigor of evaluations will scale up as AI systems become more capable.
Original article:
https://www.lesswrong.com/posts/4Gt42jX7RiaNaxCwP/more-information-about-the-dangerous-capability-evaluations
Narrated for LessWrong by TYPE III AUDIO.
client: t3a
project_id: ai_safety
narrator: pw
qa: km
Last week, The New York Times published the transcript of a conversation with Microsoft’s Bing (AKA Sydney) wherein over the course of a long chat the next-gen AI tried, very consistently, and without any prompting to do so, to break up the reporter’s marriage and to emotionally manipulate him in every way possible. I had been up late the night before researching reports coming out of similar phenomena as Sydney threatened and cajoled users across the globe, later arguing that same day in “I am Bing, and I am evil” that it was time to panic about AI safety. Like many, while I knew that current AIs were capable of these acts, what I didn’t expect was Microsoft to release one that was so obviously unhinged and yet, at the same time, so creepily convincing and intelligent.
Original article:
https://erikhoel.substack.com/p/how-to-navigate-the-ai-apocalypse
Narrated for Erik Hoel by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/fWGdsWbS6vtC9E7ii/ea-and-lw-forum-weekly-summary-6th-12th-march-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 1h20m
qa_time: 0h10m
~ A Parable of Forecasting Under Model Uncertainty ~
You, the monarch, need to know when the rainy season will begin, in order to properly time the planting of the crops. You have two advisors, Pronto and Eternidad, who you trust exactly equally.
You ask them both: "When will the next heavy rain occur?"
Pronto says, "Three weeks from today."
Eternidad says, "Ten years from today."
Original article:
https://www.lesswrong.com/posts/LzQtrHSYDafXynofq/the-parable-of-the-king-and-the-random-process#
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 1h15m
qa_time: 0h15m
Status: some mix of common wisdom (that bears repeating in our particular context), and another deeper point that I mostly failed to communicate.
Short version
Harmful people often lack explicit malicious intent. It’s worth deploying your social or community defenses against them anyway. I recommend focusing less on intent and more on patterns of harm.
(Credit to my explicit articulation of this idea goes in large part to Aella, and also in part to Oliver Habryka.)
Original article:
https://www.lesswrong.com/posts/zidQmfFhMgwFzcHhs/enemies-vs-malefactors
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/yCxsz9jk5iau2uvYH/ea-and-lw-forum-weekly-summary-27th-feb-5th-mar-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 3h30m
qa_time: 0h50m
In this article, I will present a mechanistic explanation of the Waluigi Effect and other bizarre "semiotic" phenomena which arise within large language models such as GPT-3/3.5/4 and their variants (ChatGPT, Sydney, etc). This article will be folklorish to some readers, and profoundly novel to others.
Original article:https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post
Narrated for LessWrong by TYPE III AUDIO.
Summary: Having thought a bunch about acausal trade — and proven some theorems relevant to its feasibility — I believe there do not exist powerful information hazards about it that stand up to clear and circumspect reasoning about the topic. I say this to be comforting rather than dismissive; if it sounds dismissive, I apologize.
With that said, I have four aims in writing this post:
Original article:
https://www.lesswrong.com/posts/3RSq3bfnzuL3sp46J/acausal-normalcy
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 1h30m
qa_time: 0h30m
This document looks at the predictions made by AI experts in The 2016 Expert Survey on Progress in AI, analyses the predictions on ‘Narrow tasks’, and gives a Brier score to the median of the experts’ predictions.
My analysis suggests that the experts did a fairly good job of forecasting (Brier score = 0.21), and would have been less accurate if they had predicted each development in AI to generally come, by a factor of 1.5, later (Brier score = 0.26) or sooner (Brier score = 0.29) than they actually predicted.
Original article:
https://forum.effectivealtruism.org/posts/tCkBsT6cAw6LEKAbm/scoring-forecasts-from-the-2016-expert-survey-on-progress-in
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 2h00m
qa_time: 1h00m
I don’t think the existing evidence justifies HLI's estimate of 50% household spillovers.
My main disagreements are:
Original article:
https://forum.effectivealtruism.org/posts/gr4epkwe5WoYJXF32/why-i-don-t-agree-with-hli-s-estimate-of-household
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/bEJ6SyrkSF45B2LWZ/ea-and-lw-forum-weekly-summary-20th-26th-feb-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: german_ea
narrator: ur
Effektiver Altruismus bezeichnet die Suche nach den besten Wegen, anderen zu helfen, und deren Umsetzung in die Praxis.
Es handelt sich um ein Projekt, das aus zwei komplementären Teilen besteht: zum Einen, einem Forschungsfeld, das sich darauf konzentriert, die wichtigsten globalen Probleme und die besten Lösungen für diese Probleme zu ermitteln. Zum anderen einer Gemeinschaft von Menschen, die in dem Bestreben geeint sind, auf der Grundlage dieser Erkenntnisse Gutes zu tun.
Dieses Projekt ist von großer Bedeutung, denn während viele Versuche, Gutes zu tun, scheitern, sind einige enorm effektiv. So helfen einige Hilfsorganisationen mit den gleichen Mitteln 100- oder sogar 1.000-mal so vielen Menschen wie andere.
Das bedeutet, dass wir die wichtigsten globalen Probleme weitaus besser angehen können, wenn wir sorgfältig über die besten Lösungswege nachdenken.
client: ea_forum
project_id: curated
narrator: not_t3a
Ethical philosophy often tries to systematize. That is, it seeks general principles that will explain, unify, and revise our more particular intuitions. And sometimes, this can lead to strange and uncomfortable places.
So why do it? If you believe in an objective ethical truth, you might talk about getting closer to that truth. But suppose that you don’t. Suppose you think that you’re “free to do whatever you want.” In that case, if “systematizing” starts getting tough and uncomfortable, why not just … stop? After all, you can always just do whatever’s most intuitive or common-sensical in a given case – and often, this is the choice the “ethics game” was trying so hard to validate, anyway. So why play?
I think it’s a reasonable question. And I’ve found it showing up in my life in various ways. So I wrote a set of two essays explaining part of my current take. This is the first essay. Here I describe the question in more detail, give some examples of where it shows up, and describe my dissatisfaction with two places anti-realists often look for answers.
Original article:https://joecarlsmith.com/2023/02/16/why-should-ethical-anti-realists-do-ethics
Narrated by Joe Carlsmith and included on the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/fAWotZTEnyycJnuxz/ea-and-lw-forum-weekly-summary-6th-19th-feb-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client:
project_id:
narrator: pw
qa: km
This is the story of how I came to see Wild Animal Welfare (WAW) as a less promising cause than I did initially. I summarise three articles I wrote on WAW: ‘Why it’s difficult to find cost-effective WAW interventions we could do now’, ‘Lobbying governments to improve WAW’, and ‘WAW in the far future’. I then draw some more general conclusions. The articles assume some familiarity with WAW ideas. See here or here for an intro to WAW ideas.
Original article:
https://forum.effectivealtruism.org/posts/saEQXBgzmDbob9GdH/why-i-no-longer-prioritize-wild-animal-welfare
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
Apply now to start a nonprofit in Biosecurity or Large-Scale Global Health
In this post we introduce our top five charity ideas for launch in 2023, in the areas of Biosecurity and Large-Scale Global Health. These are the result of five months’ work from our research team, and a six-stage iterative process that includes collaboration with partners and ideas from within and outside of the EA community.
We’re looking for people to launch these ideas through ourJuly - August 2023Incubation Program. The deadline for applications is March 12, 2023.
[APPLY NOW]
We provide cost-covered two-month training, stipends, ongoing mentorship, and grants up to $200,000 per project. You can learn more on our website. We also invite you to join our event on February 20, 6PM UK Time. Sam Hilton, our Director of Research, will introduce the ideas and answer your questions. Sign up here.
Original article:
https://forum.effectivealtruism.org/posts/xWRweQmmEKoLFwGyu/ce-announcing-our-2023-charity-ideas-apply-now-2
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
narrator_time: 5h30m
qa_time: 2h15m
There is a lot of disagreement and confusion about the feasibility and risks associated with automating alignment research. Some see it as the default path toward building aligned AI, while others expect limited benefit from near term systems, expecting the ability to significantly speed up progress to appear well after misalignment and deception. Furthermore, progress in this area may directly shorten timelines or enable the creation of dual purpose systems which significantly speed up capabilities research.
OpenAI recently released their alignment plan. It focuses heavily on outsourcing cognitive work to language models, transitioning us to a regime where humans mostly provide oversight to automated research assistants. While there have been a lot of objections to and concerns about this plan, there hasn’t been a strong alternative approach aiming to automate alignment research which also takes all of the many risks seriously.
The intention of this post is not to propose an end-all cure for the tricky problem of accelerating alignment using GPT models. Instead, the purpose is to explicitly put another point on the map of possible strategies, and to add nuance to the overall discussion.
Original article:
https://www.lesswrong.com/posts/bxt7uCiHam4QXrQAA/cyborgism
Narrated for LessWrong by TYPE III AUDIO.
client: joe_carlsmith
project_id:
narrator: pw
qa: km
narrator_time: 18h00m
qa_time: 5h30m
This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence. I proceed in two stages. First, I lay out a backdrop picture that informs such concern. On this picture, intelligent agency is an extremely powerful force, and creating agents much more intelligent than us is playing with fire -- especially given that if their objectives are problematic, such agents would plausibly have instrumental incentives to seek power over humans. Second, I formulate and evaluate a more specific six-premise argument that creating agents of this kind will lead to existential catastrophe by 2070. On this argument, by 2070: (1) it will become possible and financially feasible to build relevantly powerful and agentic AI systems; (2) there will be strong incentives to do so; (3) it will be much harder to build aligned (and relevantly powerful/agentic) AI systems than to build misaligned (and relevantly powerful/agentic) AI systems that are still superficially attractive to deploy; (4) some such misaligned systems will seek power over humans in high-impact ways; (5) this problem will scale to the full disempowerment of humanity; and (6) such disempowerment will constitute an existential catastrophe. I assign rough subjective credences to the premises in this argument, and I end up with an overall estimate of ~5% that an existential catastrophe of this kind will occur by 2070. (May 2022 update: since making this report public in April 2021, my estimate here has gone up, and is now at >10%.)
Original article:
https://arxiv.org/abs/2206.13353
Narrated for Joseph Carlsmith by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
This is a linkpost for https://escapingflatland.substack.com/p/childhoods
Let’s start with one of those insights that are as obvious as they are easy to forget: if you want to master something, you should study the highest achievements of your field. If you want to learn writing, read great writers, etc.
But this is not what parents usually do when they think about how to educate their kids. The default for a parent is rather to imitate their peers and outsource the big decisions to bureaucracies. But what would we learn if we studied the highest achievements?
Thinking about this question, I wrote down a list of twenty names—von Neumann, Tolstoy, Curie, Pascal, etc—selected on the highly scientific criteria “a random Swedish person can recall their name and think, Sounds like a genius to me”. That list is to me a good first approximation of what an exceptional result in the field of child-rearing looks like. I ordered a few piles of biographies, read, and took notes. Trying to be a little less biased in my sample, I asked myself if I could recall anyone exceptional that did not fit the patterns I saw in the biographies, which I could, and so I ordered a few more biographies.
This kept going for an unhealthy amount of time.
I sampled writers (Virginia Woolf, Lev Tolstoy), mathematicians (John von Neumann, Blaise Pascal, Alan Turing), philosophers (Bertrand Russell, René Descartes), and composers (Mozart, Bach), trying to get a diverse sample.
In this essay, I am going to detail a few of the patterns that have struck me after having skimmed 42 biographies. I will sort the claims so that I start with more universal patterns and end with patterns that are less common.
Original article:
https://www.lesswrong.com/posts/CYN7swrefEss4e3Qe/childhoods-of-exceptional-people
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
Hi everyone,
I've been reading up on H5N1 this weekend, and I'm pretty concerned. Right now my hunch is that there is a non-zero chance that it will cost more than 10,000 people their lives.
To be clear, I think it is unlikely that H5N1 will become a pandemic anywhere close to the size of covid.
Nevertheless, I think our community should be actively following the news and start thinking about ways to be helpful if the probability increases. I am creating this thread as a place where people can discuss and share information about H5N1. We have a lot of pandemic experts in this community, do chime in!
Original article:
https://forum.effectivealtruism.org/posts/QMMFyAX3ajf9vF5sb/h5n1-thread-for-information-sharing-planning-and-action
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
(Epistemic status: attempting to clear up a misunderstanding about points I have attempted to make in the past. This post is not intended as an argument for those points.)
I have long said that the lion's share of the AI alignment problem seems to me to be about pointing powerful cognition at anything at all, rather than figuring out what to point it at.
It’s recently come to my attention that some people have misunderstood this point, so I’ll attempt to clarify here.
Original article:
https://www.lesswrong.com/posts/NJYmovr9ZZAyyTBwM/what-i-mean-by-alignment-is-in-large-part-about-making
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 1h20m
qa_time: 0h15m
This is a linkpost for https://epochai.org/blog/literature-review-of-transformative-artificial-intelligence-timelines
We summarize and compare several models and forecasts predicting when transformative AI will be developed.
Highlights
Original article:
https://forum.effectivealtruism.org/posts/4Ckc2zNrAKQwnAyA2/literature-review-of-transformative-artificial-intelligence
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 5h00m
qa_time: 2h00m
A Chemical Hunger (a), a series by the authors of the blog Slime Mold Time Mold (SMTM), argues that the obesity epidemic is entirely caused (a) by environmental contaminants.
In my last post, I investigated SMTM’s main suspect (lithium).[1] This post collects other observations I have made about SMTM’s work, not narrowly related to lithium, but rather focused on the broader thesis of their blog post series.
I think that the environmental contamination hypothesis of the obesity epidemic is a priori plausible. After all, we know that chemicals can affect humans, and our exposure to chemicals has plausibly changed a lot over time. However, I found that several of what seem to be SMTM’s strongest arguments in favor of the contamination theory turned out to be dubious, and that nearly all of the interesting things I thought I’d learned from their blog posts turned out to actually be wrong. I’ll explain that in this post.
Original article:
https://www.lesswrong.com/posts/NRrbJJWnaSorrqvtZ/on-not-getting-contaminated-by-the-wrong-obesity-ideas
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/Qzfew7EBPgdCzsxED/ea-and-lw-forum-weekly-summary-30th-jan-5th-feb-2023
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: 80000_hours
project_id: articles
narrator: pw
qa: mds
narrator_time: 3h00m
qa_time: 0h45m
Lots of people say they want to “make a difference,” “do good,” “have a social impact,” or “make the world a better place” — but they rarely say what they mean by those terms.
By clarifying your definition, you can better target your efforts, and make a difference more effectively.
But how should you define social impact?
Thousands of years of philosophy have gone into that question. We’re going to try to sum up that thinking; introduce a practical, rough-and-ready definition of social impact; and explain why we think it’s a good definition to focus on.
This is a bit ambitious for one article, so to the philosophers in the audience, please forgive the enormous simplifications. We hope the usefulness of the definition will make up for it.
Original article:
https://80000hours.org/articles/what-is-social-impact-definition/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: mds
narrator_time: 3h00m
qa_time: 1h15m
In a nutshell: People in operations roles act as multipliers, maximising the productivity of others in the organisation by building systems that keep the organisation functioning effectively at a high level. As a result, people who excel in these positions require significant creativity, self-direction, social skills, and conscientiousness. If you’re a good fit, operations could be the highest-impact role for you.
This career review is based largely on our 2017 survey of talent needs and input from 12 people who have worked in these roles (often in leadership positions) — you can see the full list of contributors in the online version of this article. However, the views presented here do not necessarily reflect those of everyone listed.
Original article:
https://80000hours.org/articles/operations-management/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: mds
narrator_time: 2h30m
qa_time: 0h30m
Many of the highest-impact people in history have been communicators and advocates of one kind or another.
Take Rosa Parks, who in 1955 refused to give up her seat to a white man on a bus, sparking a protest which led to a Supreme Court ruling that segregated buses were unconstitutional. Parks was a seamstress in her day job, but in her spare time she was involved with the civil rights movement. After she was arrested, she and the NAACP used widely distributed fliers to launch a total boycott of buses in a city with 40,000 African Americans, while simultaneously pushing forward with legal action. This led to major progress for civil rights.
Communication can be aimed at a broad audience (like in Parks’s case) or a narrow influential group. This means there are also many examples of important communicators you’ve never heard of, like Viktor Zhdanov.
In the 20th century, smallpox killed around 400 million people — far more than died in all the century’s wars and political famines.
Although credit for the elimination of smallpox often goes to D.A. Henderson (who directly oversaw the programme), it was Viktor Zhdanov who lobbied the World Health Organization to start the elimination campaign in the first place — while facing significant opposition from the members of the World Health Assembly (the proposal passed by just two votes). Without his involvement, smallpox’s elimination probably would not have happened until much later, costing millions of lives, and possibly not at all.
So why has communicating important ideas sometimes been so effective?
Original article:
https://80000hours.org/articles/communication/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
TL;DR
Anomalous tokens: a mysterious failure mode for GPT (which reliably insulted Matthew)
Prompt generation: a new interpretability method for language models (which reliably finds prompts that result in a target completion). This is good for:
In this post, we'll introduce the prototype of a new model-agnostic interpretability method for language models which reliably generates adversarial prompts that result in a target completion. We'll also demonstrate a previously undocumented failure mode for GPT-2 and GPT-3 language models, which results in bizarre completions (in some cases explicitly contrary to the purpose of the model), and present the results of our investigation into this phenomenon. Further detail can be found in a follow-up post.
Original article:
https://www.lesswrong.com/posts/aPeJE8bSo6rAFoLqg/solidgoldmagikarp-plus-prompt-generation
Narrated for LessWrong by TYPE III AUDIO.
client: german_ea
narrator: ur
ABSTRACT. Mit sehr fortschrittlicher Technologie könnte eine sehr große Population von Personen, die glückliche Leben leben, in der erreichbaren Region des Universums aufrechterhalten werden. Daraus ergeben sich entsprechende Opportunitätskosten für jedes Jahr, um das sich die Entwicklung solcher Technologien und die folgliche Kolonialisierung des Universums verzögert: Ein potenzielles Gut, nämlich das lebenswerter Leben, wird nicht realisiert. Unter plausiblen Annahmen sind diese Kosten extrem groß. Jedoch lautet die Lehre für den Standardutilitaristen nicht, dass wir die Geschwindigkeit des technologischen Fortschritts maximieren sollten, sondern dass wir seine Sicherheit maximieren sollten, d. h. die Wahrscheinlichkeit, dass die Kolonisierung des Weltalls tatsächlich stattfinden wird. Dieses Ziel hat eine so hohe Utilität, dass Standardutilitaristen all ihre Energie darauf verwenden sollten. Utilitaristen der „personenbezogenen“ Sorte sollten eine modifizierte Version dieser Schlussfolgerung akzeptieren. Manch andere ethische Sichtweisen, welche utilitaristische Erwägungen mit anderen Kriterien kombinieren, werden zu einem ähnlichen Fazit kommen.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
narrator_time: 0h40m
qa_time: 0h15m
Writing down something I’ve found myself repeating in different conversations:
If you're looking for ways to help with the whole “the world looks pretty doomed” business, here's my advice: look around for places where we're all being total idiots.
Look for places where everyone's fretting about a problem that some part of you thinks it could obviously just solve.
Look around for places where something seems incompetently run, or hopelessly inept, and where some part of you thinks you can do better.
Then do it better.
Original article:
https://www.lesswrong.com/posts/Zp6wG5eQFLGWwcG6j/focus-on-the-places-where-you-feel-shocked-everyone-s
Narrated for the LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 1h00m
qa_time: 0h30m
This post outlines the capability approach to thinking about human welfare. I think that this approach, while very popular in international development, is neglected in EA. While the capability approach has problems, I think that it provides a better approach to thinking about improving human welfare than approaches based on measuring happiness or subjective wellbeing (SWB) or approaches based on preference satisfaction. Finally, even if you disagree that the capability approach is best, I think this post will be useful to you because it may clarify why many people and organizations in the international development or global health space take the positions that they do. I will be drawing heavily on the work of Amartya Sen, but I will often not be citing specific texts because I’m an academic and getting to write without careful citations is thrilling.
Original article:
https://forum.effectivealtruism.org/posts/zy6jGPeFKHaoxKEfT/the-capability-approach
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/hzc26vGa4RLns7TvK/ea-and-lw-forum-weekly-summary-23rd-29th-jan-23
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
narrator_time: 3h40m
qa_time: 1h45m
This post is meant to be a linkable resource. Its core is a short list of guidelines (you can link directly to the list) that are intended to be fairly straightforward and uncontroversial, for the purpose of nurturing and strengthening a culture of clear thinking, clear communication, and collaborative truth-seeking.
"Alas," said Dumbledore, "we all know that what should be, and what is, are two different things. Thank you for keeping this in mind."
There is also (for those who want to read more than the simple list) substantial expansion/clarification of each specific guideline, along with justification for the overall philosophy behind the set.
Original article:
https://www.lesswrong.com/posts/XPv4sYrKnPzeJASuk/basics-of-rationalist-discourse-1
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
narrator_time: 0h45m
qa_time: 0h15m
I think that EA [editor note: "Effective Altruism"] burnout usually results from prolonged dedication to satisfying the values you think you should have, while neglecting the values you actually have.
Setting aside for the moment what “values” are and what it means to “actually” have one, suppose that I actually value these things (among others):
True Values:
One day I learn about “global catastrophic risk”: Perhaps we’ll all die in a nuclear war, or an AI apocalypse, or a bioengineered global pandemic, and perhaps one of these things will happen quite soon.
I recognize that GCR is a direct threat to The Wellbeing Of Others and to Personal Longevity, and as I do, I get scared. I get scared in a way I have never been scared before, because I’ve never before taken seriously the possibility that everyone might die, leaving nobody to continue the species or even to remember that we ever existed—and because this new perspective on the future of humanity has caused my own personal mortality to hit me harder than the lingering perspective of my Christian upbringing ever allowed. For the first time in my life, I’m really aware that I, and everyone I will ever care about, may die.
Original article:
https://www.lesswrong.com/posts/pDzdb4smpzT3Lwbym/my-model-of-ea-burnout
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
narrator_time: 2h15m
qa_time: 1h15m
Casus Belli: As I was scanning over my (rather long) list of essays-to-write, I realized that roughly a fifth of them were of the form "here's a useful standalone concept I'd like to reify," à la cup-stacking skills, fabricated options, split and commit, and sazen. Some notable entries on that list (which I name here mostly in the hope of someday coming back and turning them into links) include: red vs. white, walking with three, setting the zero point[1], seeding vs. weeding, hidden hinges, reality distortion fields, and something-about-layers-though-that-one-obviously-needs-a-better-word.
While it's still worthwhile to motivate/justify each individual new conceptual handle (and the planned essays will do so), I found myself imagining a general objection of the form "this is just making up terms for things," or perhaps "this is too many new terms, for too many new things." I realized that there was a chunk of argument, repeated across all of the planned essays, that I could factor out, and that (to the best of my knowledge) there was no single essay aimed directly at the question "why new words/phrases/conceptual handles at all?"
So ... voilà.
Original article:
https://www.lesswrong.com/posts/PCrTQDbciG4oLgmQ5/sapir-whorf-for-rationalists
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:https://forum.effectivealtruism.org/posts/6Ezg8HgHib9bpWCFr/ea-and-lw-forum-weekly-summary-16th-22nd-jan-23
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: german_ea
service: human_edits_ai_narrates
narrator: nnm
qa: ph
Ihre Namen werden wir nie kennen.
Das erste Opfer hatte nicht erfasst werden können, denn damals gab es noch keine Schriftsprache, in der man es aufzeichnen konnte. Die Opfer waren jemandes Töchter oder Söhne, ein Freund oder eine Freundin. Sie wurden von den Menschen in ihrem Umfeld geliebt. Und sie hatten Schmerzen, ihre Haut war von Ausschlägen bedeckt, sie waren verwirrt, verängstigt, wussten nicht, warum ihnen das geschah oder was sie dagegen tun konnten – sie waren die Opfer eines zornigen, unmenschlichen Gottes. Es gab nichts, was man tun konnte – die Menschheit war noch nicht stark genug, noch nicht sensibilisiert genug, nicht sachkundig genug, um ein Ungeheuer zu bekämpfen, das man nicht sehen konnte.
client: german_ea
service: human_narration
narrator: ur
qa: not_t3a
This is an audio narration of the German translation of The Fable of the Dragon-Tyrant by Nick Bostrom. The translation was done by Franz Fuchs, edited by Stephan Dalügge and narrated by Uta Reichardt. You can find the original paper at nickbostrom.com. Links and related reading suggestions are in the episode description.
Es war einmal vor langer, langer Zeit, da wurde unser Planet von einem riesigen Drachen tyrannisiert. Der Drache überragte selbst die höchste Kathedrale und war mit einem dicken Panzer aus schwarzen Schuppen bedeckt. Seine roten Augen glühten vor Hass und aus seinem furchtbaren Maul floss beständig ein übel riechender, gelblich-grüner Schleim. Er verlangte der Menschheit einen Furcht einflößenden Tribut ab: Um seinen gigantischen Appetit zu stillen, mussten jeden Tag beim Einbruch der Dunkelheit zehntausend Männer und Frauen zum Fuß des Berges gebracht werden, wo der tyrannische Drache lebte. Manchmal verschlang der Drache die Unglücklichen sofort; manchmal wiederum kerkerte er sie im Berg ein. Dort mussten sie Monate oder Jahre dahinsiechen, bis sie schließlich verspeist wurden.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 2h15m
qa_time: 1h0m
Original article:
https://forum.effectivealtruism.org/posts/Qk3hd6PrFManj8K6o/rethink-priorities-welfare-range-estimates
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: km
narrator_time: 0h45m
qa_time: 0h15m
For many years, I've actively lived in avoidance of idolizing behavior and in pursuit of a nuanced view of even those I respect most deeply. I think this has helped me in numerous ways and has been of particular help in weathering the past few months within the EA community. Below, I discuss how I think about the act of idolizing behavior, some of my personal experiences, and how this mentality can be of use to others.
Note: I want more people to post on the EA Forum and have their ideas taken seriously regardless of whether they conform to Forum stylistic norms. I'm perfectly capable of writing a version of this post in the style typical to the Forum, but this post is written the way I actually like to write. If this style doesn’t work for you, you might want to read the first section “Anarchists have no idols” and then skip ahead to the section “Living without idols, Pt. 1” toward the end. You’ll lose some of the insights contained in my anecdotes, but still get most of the core ideas I want to convey here.
Original article:
https://forum.effectivealtruism.org/posts/jgspXC8GKA7RtxMRE/on-living-without-idols
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 1h15m
qa_time: 0h40m
One of the most discussed topics online recently has been friendships and loneliness. Ever since the infamous chart showing more people are not having sex than ever before first made the rounds, there’s been increased interest in the social state of things. Polling has demonstrated a marked decline in all spheres of social life, including close friends, intimate relationships, trust, labor participation, and community involvement. The trend looks to have worsened since the pandemic, although it will take some years before this is clearly established.
The decline comes alongside a documented rise in mental illness, diseases of despair, and poor health more generally. In August 2022, the CDC announced that U.S. life expectancy has fallen further and is now where it was in 1996. Contrast this to Western Europe, where it has largely rebounded to pre-pandemic numbers. Still, even before the pandemic, the years 2015-2017 saw the longest sustained decline in U.S. life expectancy since 1915-18. While my intended angle here is not health-related, general sociability is closely linked to health. The ongoing shift has been called the “friendship recession” or the “social recession.”
My intention here is not to present a list of miserable points, but to group them together in a meaningful context whose consequences are far-reaching. While most of what I will outline here focuses on the United States, many of these same trends are present elsewhere because its catalyst is primarily the internet itself. With no signs of abating, a new kind of sociability has only started to affect what people ask of the world through the prism of themselves.
Original article:
https://www.lesswrong.com/posts/Xo7qmDakxiizG7B9c/the-social-recession-by-the-numbers
Narrated for LessWrong by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: mds
narrator_time: 3h00m
qa_time: 1h00m
For the right person, becoming a journalist could be very impactful. Good journalists help keep people informed, positively shape public discourse on important topics, and can provide a platform for people and ideas that the public might not otherwise hear about.
But the most influential positions in the field are highly competitive, and journalists face a lot of mixed incentives that may detract from their ability to have a positive impact.
Original article:
https://80000hours.org/career-reviews/journalism/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 1h45m
qa_time: 0h30m
I think Zvi's Immoral Mazes sequence is really important, but comes with more worldview-assumptions than are necessary to make the points actionable. I conceptualize Zvi as arguing for multiple hypotheses. In this post I want to articulate one sub-hypothesis, which I call "Recursive Middle Manager Hell". I'm deliberately not covering some other components of his model[1].
tl;dr:
Something weird and kinda horrifying happens when you add layers of middle-management. This has ramifications on when/how to scale organizations, and where you might want to work, and maybe general models of what's going on in the world.
You could summarize the effect as "the org gets more deceptive, less connected to its original goals, more focused on office politics, less able to communicate clearly within itself, and selected for more for sociopathy in upper management."
You might read that list of things and say "sure, seems a bit true", but one of the main points here is "Actually, this happens in a deeper and more insidious way than you're probably realizing, with much higher costs than you're acknowledging. If you're scaling your organization, this should be one of your primary worries."
Original article:
https://www.lesswrong.com/posts/pHfPvb4JMhGDr4B7n/recursive-middle-manager-hell
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:
https://forum.effectivealtruism.org/posts/DNWpFLrtrJXe4mted/ea-and-lw-forum-summaries-9th-jan-to-15th-jan-23
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 2h20m
qa_time: 0h45m
I previously wrote an entry for the Open Philanthropy Cause Exploration Prize on why preventing violence against women and girls is a global priority. For an introduction to the area, I have written a brief summary below. In this post, I will extend that work, diving deeper into the literature and the landscape of organisations in the field, as well as creating a cost-effectiveness model for some of the most promising preventative interventions. Based on this, I will offer some concrete recommendations that different stakeholders should take - from individuals looking to donate, to funders, to charity evaluators and incubators.
Original article:
https://forum.effectivealtruism.org/posts/uH9akQzJkzpBD5Duw/what-you-can-do-to-help-stop-violence-against-women-and
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 4h45m
qa_time: 1h30m
In this post, we point out that short AI timelines would cause real interest rates to be high, and would do so under expectations of either unaligned or aligned AI. However, 30- to 50-year real interest rates are low. We argue that this suggests one of two possibilities:
In the rest of this post we flesh out this argument.
Original article:
https://forum.effectivealtruism.org/posts/8c7LycgtkypkgYjZx/agi-and-the-emh-markets-are-not-expecting-aligned-or
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:
https://forum.effectivealtruism.org/posts/JZuCg7TtfzzaX9bBY/ea-and-lw-forum-summaries-holiday-edition-19th-dec-8th-jan
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
narrator_time: 2h00m
qa_time: 0h30m
Should I quit my job? Which of my offers should I take? Which long-term options should I explore?
These decisions will affect how you spend years of your time, so the stakes are high. But they’re also an area where you shouldn’t expect your intuition to be a reliable guide. This means it’s worth taking a more systematic approach.
What might a good career decision process look like?
A common approach is to make a pro and con list, but it’s possible to do a lot better. Pro and con lists make it easy to put too much weight on an unimportant factor. More importantly, they don’t encourage you to make use of the most powerful decision-making methods, which can greatly improve the quality of your decisions.
Original article:
https://80000hours.org/career-decision/article/
Narrated for the 80,000 Hours by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 2h15m
qa_time: 0h35m
A few collaborators and I recently released a new paper: Discovering Latent Knowledge in Language Models Without Supervision. For a quick summary of our paper, you can check out this Twitter thread.
In this post I will describe how I think the results and methods in our paper fit into a broader scalable alignment agenda. Unlike the paper, this post is explicitly aimed at an alignment audience and is mainly conceptual rather than empirical.
Tl;dr: unsupervised methods are more scalable than supervised methods, deep learning has special structure that we can exploit for alignment, and we may be able to recover superhuman beliefs from deep learning representations in a totally unsupervised way.
Original article:
https://www.lesswrong.com/posts/L4anhrxjv8j2yRKKp/how-discovering-latent-knowledge-in-language-models-without
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 1h05m
qa_time: 0h10m
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.
In terms of content, this has a lot of overlap with Reward is not the optimization target. I'm basically rewriting a part of that post in language I personally find clearer, emphasising what I think is the core insight.
When thinking about deception and RLHF training, a simplified threat model is something like this:
Before continuing, I would encourage you to really engage with the above. Does it make sense to you? Is it making any hidden assumptions? Is it missing any steps? Can you rewrite it to be more mechanistically correct?
I believe that when people use the above threat model, they are either using it as shorthand for something else or they misunderstand how reinforcement learning works. Most alignment researchers will be in the former category. However, I was in the latter.
Original article:
https://www.lesswrong.com/posts/TWorNr22hhYegE4RT/models-don-t-get-reward
Narrated for LessWrong by TYPE III AUDIO.
client: lesswrong
project_id: curated
narrator: pw
qa: km
narrator_time: 1h05m
qa_time: 0h10m
Here’s a story you may recognize. There's a bright up-and-coming young person - let's call her Alice. Alice has a cool idea. It seems like maybe an important idea, a big idea, an idea which might matter. A new and valuable idea. It’s the first time Alice has come up with a high-potential idea herself, something which she’s never heard in a class or read in a book or what have you.
So Alice goes all-in pursuing this idea. She spends months fleshing it out. Maybe she writes a paper, or starts a blog, or gets a research grant, or starts a company, or whatever, in order to pursue the high-potential idea, bring it to the world.
And sometimes it just works!
… but more often, the high-potential idea doesn’t actually work out. Maybe it turns out to be basically-the-same as something which has already been tried. Maybe it runs into some major barrier, some not-easily-patchable flaw in the idea. Maybe the problem it solves just wasn’t that important in the first place.
Original article:
https://www.lesswrong.com/posts/mfPHTWsFhzmcXw8ta/the-feeling-of-idea-scarcity
Narrated for LessWrong by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: km
narrator_time: 2h45m
qa_time: 0h50m
At Anima International, we recently decided to suspend our campaign against live fish sales in Poland indefinitely. After a few years of running the campaign, we are now concerned about the effects of our efforts, specifically the possibility of a net negative result for the lives of animals. We believe that by writing about it openly we can help foster a culture of intellectual honesty, information sharing and accountability. Ideally, our case can serve as a good example on reflecting on potential unintended consequences of advocacy interventions.
Original article:
https://forum.effectivealtruism.org/posts/snnfmepzrwpAsAoDT/why-anima-international-suspended-the-campaign-to-end-live
This is a linkpost for https://animainternational.org/blog/why-anima-international-suspended-the-campaign-to-end-live-fish-sales-in-poland
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: mds
narrator_time: 1h10m
qa_time: 0h15m
We believe that some of the career paths open to you likely have over 100 times more positive impact than other paths you might take.
Why? In our key ideas series, we’ve shown that you can have more impact by:
We’ve also shown that there are big differences for each factor:
On top of that, you can further increase your impact by having a good career strategy, such as by striking the right balance between investing in yourself and having an impact right away.
Original article:
https://80000hours.org/articles/careers-differ-in-impact/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: ea_forum
project_id: curated
qa: mds
narrator_time: 1h40m
qa_time: 0h0m
This is a linkpost for https://simonm.substack.com/p/strongminds-should-not-be-a-top-rated
GWWC lists StrongMinds as a “top-rated” charity. Their reason for doing so is because Founders Pledge has determined they are cost-effective in their report into mental health.
I could say here, “and that report was written in 2019 - either they should update the report or remove the top rating” and we could all go home. In fact, most of what I’m about to say does consist of “the data really isn’t that clear yet”.
I think the strongest statement I can make (which I doubt StrongMinds would disagree with) is:
“StrongMinds have made limited effort to be quantitative in their self-evaluation, haven’t continued monitoring impact after intervention, haven’t done the research they once claimed they would. They have not been vetted sufficiently to be considered a top charity, and only one independent group has done the work to look into them.”
Original article:
https://forum.effectivealtruism.org/posts/ffmbLCzJctLac3rDu/strongminds-should-not-be-a-top-rated-charity-yet
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: km
narrator_time: 5h00m
qa_time: 1h50m
If you fear that someone will build a machine that will seize control of the world and annihilate humanity, then one kind of response is to try to build further machines that will seize control of the world even earlier without destroying it, forestalling the ruinous machine’s conquest. An alternative or complementary kind of response is to try to avert such machines being built at all, at least while the degree of their apocalyptic tendencies is ambiguous.
The latter approach seems to me like the kind of basic and obvious thing worthy of at least consideration, and also in its favor, fits nicely in the genre ‘stuff that it isn’t that hard to imagine happening in the real world’. Yet my impression is that for people worried about extinction risk from artificial intelligence, strategies under the heading ‘actively slow down AI progress’ have historically been dismissed and ignored (though ‘don’t actively speed up AI progress’ is popular).
Original article:
https://forum.effectivealtruism.org/posts/vwK3v3Mekf6Jjpeep/let-s-think-about-slowing-down-ai-1
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: not_t3a
qa: not_t3a
In previous pieces, I argued that there's a real and large risk of AI systems' aiming to defeat all of humanity combined - and succeeding.
I first argued that this sort of catastrophe would be likely without specific countermeasures to prevent it. I then argued that countermeasures could be challenging, due to some key difficulties of AI safety research.But while I think misalignment risk is serious and presents major challenges, I don’t agree with sentiments along the lines of “We haven’t figured out how to align an AI, so if transformative AI comes soon, we’re doomed.” Here I’m going to talk about some of my **high-level hopes for how we might end up avoiding this risk.
Original article:**https://forum.effectivealtruism.org/posts/rJRw78oihoT5paFGd/high-level-hopes-for-ai-alignment
Narrated by Holden Karnofsky for the Cold Takes blog.
client: lesswrong
project_id: curated
narrator: pw
qa: mds
narrator_time: 4h30m
qa_time: 2h0m
This post is inspired by What 2026 looks like and an AI vignette workshop guided by Tamay Besiroglu. I think of this post as “what would I expect the world to look like if these timelines (median compute for transformative AI ~2036) were true” or “what short-to-medium timelines feel like” since I find it hard to translate a statement like “median TAI year is 20XX” into a coherent imaginable world.
I expect some readers to think that the post sounds wild and crazy but that doesn’t mean its content couldn’t be true. If you had told someone in 1990 or 2000 that there would be more smartphones and computers than humans in 2020, that probably would have sounded wild to them. The same could be true for AIs, i.e. that in 2050 there are more human-level AIs than humans. The fact that this sounds as ridiculous as ubiquitous smartphones sounded to the 1990/2000 person, might just mean that we are bad at predicting exponential growth and disruptive technology.
Original article:
https://www.lesswrong.com/posts/qRtD4WqKRYEtT5pi3/the-next-decades-might-be-wild
Narrated for LessWrong by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
narrator_time: 2h40m
qa_time: 0h50m
We’ve argued that preventing an AI-related catastrophe may be the world’s most pressing problem, and that while progress in AI over the next few decades could have enormous benefits, it could also pose severe, possibly existential risks. As a result, we think that working on some technical AI research — research related to AI safety — may be a particularly high-impact career path.
But there are many ways of approaching this path that involve researching or otherwise advancing AI capabilities — meaning making AI systems better at some specific skills — rather than only doing things that are purely in the domain of safety. In short, this is because capabilities work and some forms of safety work are intertwined, and many available ways of learning enough about AI to contribute to safety are via capabilities-enhancing roles.
So if you want to help prevent an AI-related catastrophe, should you be open to roles that also advance AI capabilities, or steer clear of them?
Original article:
https://80000hours.org/articles/ai-capabilities/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:
https://forum.effectivealtruism.org/posts/8bcPkqdLYG78YbnTh/ea-and-lw-forums-weekly-summary-5th-dec-11th-dec-22
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
narrator_time: 1h0m
qa_time: 0h10m
If you want to have an impact, the aim is to find a job that has the potential to make a big contribution to a pressing problem, and that’s a good fit for you. But how can you find a job like that?
In the strategy section of our key ideas series, we discuss the value of exploration and career capital, as well as many other ideas, like why to be more ambitious. Here we sum them up into a simple career strategy.
Original article:
https://80000hours.org/articles/key-career-stages/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
narrator_time: 1h0m
qa_time: 0h30m
We took 10 years of research and what we’ve learned from advising more than 1,000 people on how to build high-impact careers, compressed that into an eight-week course to create your career plan, and then compressed that into this summary of the main points.
(It’s especially aimed at people who want a career that’s both satisfying and has a significant positive impact, but much of the advice applies to all career decisions.)
Original article:
https://80000hours.org/career-planning/summary/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
narrator_time: 1h30m
qa_time: 20m
Why your career is your biggest opportunity to make a difference to the worldWhen people think of living ethically, they most often think of things like recycling, fair trade, and volunteering.
But that’s missing something huge: your choice of career.
We believe that what you do with your career is probably the most important ethical decision of your life.
The first reason is the huge amount of time at stake. You have about 80,000 hours in your career: 40 hours per week, 50 weeks per year, for 40 years. That’s more time than you’ll spend eating, socialising, and watching Netflix put together.
And it means (unless you happen to be the heir to a large estate) that time is the biggest resource you have to help others.
So if you can increase the overall impact of your career by just a tiny amount, it will likely do more good than changes you could make to other parts of your life.
Or, to look at it another way: it’s worth thinking a lot about how to make even just small improvements to your career. For instance, if you could increase the impact of your career by 1%, it would be worth spending up to 800 hours working out how to do that.
And that brings us to the second reason why your choice of career is so important: some careers give you the opportunity to do vastly more good for the world than others — to a much greater extent than people realise.
In fact, we’ll argue that some career paths open to you likely have 10, 100, or even 1,000 times more impact than others. And this makes it even more important to think hard about your career.
Why do careers differ so much in impact?
Original article:
https://80000hours.org/make-a-difference-with-your-career/
Narrated for 80,000 Hours by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: km
narrator_time: 3h15m
qa_time: 0h30m
Question 1: What are these lists based on?
Our aim is to find the problems where an additional person can have the greatest positive social impact, given how effort is already allocated in society.
The primary way we do that is by trying to compare global issues based on their scale, neglectedness, and tractability. To learn about this framework, see our introductory article on prioritising world problems.
To assess the problems based on this framework, we mainly draw upon research and advice from subject-matter experts and advisors in the effective altruism research community — including the Global Priorities Institute, Rethink Priorities, and Open Philanthropy — though we also make some of our own judgement calls in borderline cases.
To see the reasons why we listed each individual problem, click through to see the full profiles.
Assessments of the scale and tractability of different global issues depend on your values and worldview. You can see some of the most important aspects of our worldview in the ‘foundations’ section of our key ideas series, especially our article on how we define social impact.
All this has led to a few themes in the issues we tend to prioritise most highly:
Original article:
https://80000hours.org/problem-profiles/#problems-faq
Narrated for 80,000 Hours by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: mds
narrator_time: 1h50m
qa_time: 45m
Key Takeaways
Original article:
https://forum.effectivealtruism.org/posts/tnSg6o7crcHFLc395/the-welfare-range-table
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: curated
narrator: pw
qa: km
narrator_time: 1h30m
qa_time: 25m
Hiya folks! I'm Patrick McKenzie, better known on the Internets as patio11. (Proof.) Long-time-listener, first-time-caller; I don't think I would consider myself an EA but I've been reading y'all, and adjacent intellectual spaces, for some time now.
Epistemic status: Arbitrarily high confidence with regards to facts of the VaccinateCA experience (though speaking only for myself), moderately high confidence with respect to inferences made about vaccine policy and mechanisms for impact last year, one geek's opinion with respect to implicit advice to you all going forward.
A Thing That Happened Last Year
As some of the California-based EAs may remember, the rollout of the covid-19 vaccines in California and across the U.S. was... not optimal. I accidentally ended up founding a charity, VaccinateCA, which ran the national shadow vaccine location information infrastructure for 6 months.
The core product at the start of the sprint, which some of you may be familiar with, was a site which listed places to get the vaccine in California, sourced by a volunteer-driven operation to conduct an ongoing census of medical providers by calling them. Importantly, that was not our primary vector for impact, though it was very important to our trajectory.
Original article:
https://forum.effectivealtruism.org/posts/NkPghabDd54nkG3kX/some-observations-from-an-ea-adjacent-charitable-effort
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
client: ea_forum
project_id: summaries
narrator: cs
Original article:
https://forum.effectivealtruism.org/posts/LdEPDqyZvucQkxhWH/ea-and-lw-forums-weekly-summary-28th-nov-4th-dec-22
This is part of a weekly series summarizing the top posts on the EA Forum — you can see the full collection here. The first post includes some details on purpose and methodology. Feedback, thoughts, and corrections are welcomed.
Narrated by Coleman Jackson Snell. Summaries written by Zoe Williams (Rethink Priorities).
Published by TYPE III AUDIO on behalf of the Effective Altruism Forum.
client: ea_forum
project_id: curated
narrator: pw
qa: km
narrator_time: 2h
qa_time: 35m
Excerpt:
This is a linkpost for https://docs.google.com/document/d/1p50vw84-ry2taYmyOIl4B91j7wkCurlB/edit?rtpof=true&sd=true
Key Takeaways* Several influential EAs have suggested using neuron counts as rough proxies for animals’ relative moral weights. We challenge this suggestion. * We take the following ideas to be the strongest reasons in favor of a neuron count proxy: + neuron counts are correlated with intelligence and intelligence is correlated with moral weight, + additional neurons result in “more consciousness” or “more valenced consciousness,” and + increasing numbers of neurons are required to reach thresholds of minimal information capacity required for morally relevant cognitive abilities. * However: + in regards to intelligence, we can question boththe extent to which more neurons are correlated with intelligence and whether more intelligence in fact predicts greater moral weight; + many ways of arguing that more neurons results in more valenced consciousness seem incompatible with our current understanding of how the brain is likely to work; and + there is no straightforward empirical evidence or compelling conceptual arguments indicating that relative differences in neuron counts within or between species reliably predicts welfare relevant functional capacities. * Overall, we suggest that neuron counts should not be used as a sole proxy for moral weight, but cannot be dismissed entirely. Rather, neuron counts should be combined with other metrics in an overall weighted score that includes information about whether different species have welfare-relevant capacities.
Original article:
https://forum.effectivealtruism.org/posts/Mfq7KxQRvkeLnJvoB/why-neuron-counts-shouldn-t-be-used-as-proxies-for-moral
Narrated for the Effective Altruism Forum by TYPE III AUDIO.
Excerpt:
At EA Global: San Francisco 2022, the following organisations held a joint session to discuss their different approaches to measuring ‘good’:
A representative from each organisation gave a five-minute lightning talk summarising their approach before the audience broke out into table discussions.
Original article:
https://forum.effectivealtruism.org/posts/8whqn2GrJfvTjhov6/measuring-good-better-1
Edited for the Effective Altruism Forum by TYPE III AUDIO.
client: 80000_hours
project_id: articles
narrator: pw
qa: mds
narrator_time: 3h 13m
qa_time: 2h
Excerpt:
Could climate change lead to the end of civilisation?
Across the world, over half of young people worry that, as a result of climate change, humanity is doomed. They feel angry, powerless, and — above all — afraid about what the future may hold.
Climate change matters so much, to so many, not just because of the suffering and injustice it’s already causing, but also because it’s one of the few issues that has obvious potential to affect our world over many future generations. We think safeguarding future generations is a key moral priority, and should be a crucial consideration in prioritising problems on which to work.
If climate change could lead to the end of civilisation, then that would mean future generations might never get to exist – or they could live in a permanently worse world. If so, then preventing climate change, and adapting to its effects, might be more important than working on almost any other issue.
So – what does the science say?
The Intergovernmental Panel on Climate Change (IPCC) Sixth Assessment Report is, to our knowledge, the most authoritative and comprehensive source on climate change. The report is clear: climate change will be hugely destructive. We’ll see floods, famines, fires, and droughts — and the world’s poorest people will be affected the most.
But even when we try to account for unknown unknowns, nothing in the IPCC’s report suggests that civilisation will be destroyed.
This isn’t to say society shouldn’t do far more to tackle climate change.
Original article:
https://80000hours.org/problem-profiles/climate-change/
Narrated for 80,000 Hours by TYPE III AUDIO.
Just the bottom lines from our key ideas series.
https://80000hours.org/key-ideas/summary/
TLDR: Get good at something that lets you effectively contribute to big and neglected global problems.
What ultimately makes for an impactful career? You can have more positive impact over the course of your career by aiming to:
narrator_time: 1h
editing_time: 0
narrator: pw
editor: pw
qa: MDS
client: 80000_hours
project_id: 80000_hours
https://forum.effectivealtruism.org/posts/9Y6Y6qoAigRC7A8eX/my-take-on-what-we-owe-the-future
Cross-posted from Foxy Scout
OverviewWhat We Owe The Future (WWOTF) by Will MacAskill has recently been released with much fanfare. While I strongly agree that future people matter morally and we should act based on this, I think the book isn’t clear enough about MacAskill’s views on longtermist priorities, and to the extent it is it presents a mistaken view of the most promising longtermist interventions.
I argue that MacAskill:
I highlight and expand on these disagreements in part to contribute to the debate on these topics, but also make a practical recommendation.
While I like some aspects of the book, I think The Precipice is a substantially better introduction for potential longtermist direct workers, e.g. as a book given away to talented university students. For instance, I’m worried people will feel bait-and-switched if they get into EA via WWOTF then do an 80,000 Hours call or hang out around their EA university group and realize most people think AI risk is the biggest longtermist priority, many thinking this by a large margin.[1] more
What I disagree with[2]Underestimating risk of misaligned AI takeover
Overall probability of takeover
In endnote 2.22 (p. 274), MacAskill writes [emphasis mine]:
I put that possibility [of misaligned AI takeover] at around 3 percent this century… I think most of the risk we face comes from scenarios where there is a hot or cold war between great powers.
narrator_time: 4h
editing_time: 0
narrator: pw
editor: pw
qa: KM
client: ea_forum
project_id: EA Forum Curated
https://forum.effectivealtruism.org/posts/PyZCqLrDTJrQofEf7/how-bad-could-a-war-get
Acknowledgements: Thanks to Joe Benton for research advice and Ben Harack and Max Daniel for feedback on earlier drafts.
Author contributions: Stephen and Rani both did research for this post; Stephen wrote it and Rani gave comments and edits.
Previously in this series: "Modelling great power conflict as an existential risk factor" and "How likely is World War III?"
Introduction & ContextIn “How Likely is World War III?”, Stephen suggested the chance of an extinction-level war occurring sometime this century is just under 1%. This was a simple, rough estimate, made in the following steps:
Not everybody was convinced. Arden Koehler of 80,000 Hours, for example, slammed it as “[overstating] the risk because it doesn’t consider that wars would be unlikely to continue once 90% or more of the population has been killed.” While our friendship may never recover, I (Stephen) have to admit that some skepticism is justified. An extinction-level war would be 30-to-100 times larger than World War II, the most severe war humanity has experienced so far.[1] Is it reasonable to just assume number go up? Would the same escalatory dynamics that shape smaller wars apply at this scale?
Forecasting the likelihood of enormous wars is difficult. Stephen’s extrapolatory approach creates estimates that are sensitive to the data included and the kind of distribution fit, particularly in the tails. But such efforts are important despite their defects. Estimates of the likelihood of major conflict are an important consideration for cause prioritization. And out-of-sample conflicts may account for most of the x-risk accounted for by global conflict. So in this post we interrogate two of the assumptions made in “How Likely is World War III?”:
Our findings are:
https://forum.effectivealtruism.org/posts/jk7A3NMdbxp65kcJJ/500-million-but-not-a-single-one-more
We will never know their names.
The first victim could not have been recorded, for there was no written language to record it. They were someone’s daughter, or son, and someone’s friend, and they were loved by those around them. And they were in pain, covered in rashes, confused, scared, not knowing why this was happening to them or what they could do about it — victims of a mad, inhuman god. There was nothing to be done — humanity was not strong enough, not aware enough, not knowledgeable enough, to fight back against a monster that could not be seen.
It was in Ancient Egypt, where it attacked slave and pharaoh alike. In Rome, it effortlessly decimated armies. It killed in Syria. It killed in Moscow. In India, five million dead. It killed a thousand Europeans every day in the 18th century. It killed more than fifty million Native Americans. From the Peloponnesian War to the Civil War, it slew more soldiers and civilians than any weapon, any soldier, any army. (Not that this stopped the most foolish and empty souls from attempting to harness the demon as a weapon against their enemies.)
Cultures grew and faltered, and it remained. Empires rose and fell, and it thrived. Ideologies waxed and waned, but it did not care. Kill. Maim. Spread. An ancient, mad god, hidden from view, that could not be fought, could not be confronted, could not even be comprehended. Not the only one of its kind, but the most devastating.
For a long time, there was no hope — only the bitter, hollow endurance of survivors.
In China, in the 10th century, humanity began to fight back.
It was observed that survivors of the mad god’s curse would never be touched again: They had taken a portion of that power into themselves, and were so protected from it. Not only that, but this power could be shared by consuming a remnant of the wounds. There was a price, for you could not take the god’s power without first defeating it — but a smaller battle, on humanity’s terms.
By the 16th century, the technique spread to India, then across Asia, the Ottoman Empire and, in the 18th century, Europe. In 1796, a more powerful technique was discovered by Edward Jenner.
An idea began to take hold: Perhaps the ancient god could be killed.
narrator_time: 45m
editing_time: 0
narrator: pw
editor: pw
qa: MDS
client: ea_forum
project_id: EA Forum Curated
https://www.lesswrong.com/posts/SqjQFhn5KTarfW8v7/lessons-learned-from-talking-to-greater-than-100-academics
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.
I’d like to thank MH, Jaime Sevilla and Tamay Besiroglu for their feedback.
During my Master's and Ph.D. (still ongoing), I have spoken with many academics about AI safety. These conversations include chats with individual PhDs, poster presentations and talks about AI safety.
I think I have learned a lot from these conversations and expect many other people concerned about AI safety to find themselves in similar situations. Therefore, I want to detail some of my lessons and make my thoughts explicit so that others can scrutinize them.
TL;DR: People in academia seem more and more open to arguments about risks from advanced intelligence over time and I would genuinely recommend having lots of these chats. Furthermore, I underestimated how much work related to some aspects AI safety already exists in academia and that we sometimes reinvent the wheel. Messaging matters, e.g. technical discussions got more interest than alarmism and explaining the problem rather than trying to actively convince someone received better feedback.
narrator_time: 2h20m
narrator: pw
editor: pw
qa: km
client: 80000_hours
project_id: 80000_hours
https://80000hours.org/career-reviews/china-related-ai-safety-and-governance-paths/#strong-networking-abilities-especially-for-policy-roles
Expertise in China and its relations with the world might be critical in tackling some of the world’s most pressing problems. In particular, China’s relationship with the US is arguably the most important bilateral relationship in the world, with these two countries collectively accounting for over 40% of global GDP.1 These considerations led us to publish a guide to improving China–Western coordination on global catastrophic risks and other key problems in 2018. Since then, we have seen an increase in the number of people exploring this area.
China is one of the most important countries developing and shaping advanced artificial intelligence (AI). The Chinese government’s spending on AI research and development is estimated to be on the same order of magnitude as that of the US government,2 and China’s AI research is prominent on the world stage and growing.
Because of the importance of AI from the perspective of improving the long-run trajectory of the world, we think relations between China and the US on AI could be among the most important aspects of their relationship. Insofar as the EU and/or UK influence advanced AI development through labs based in their countries or through their influence on global regulation, the state of understanding and coordination between European and Chinese actors on AI safety and governance could also be significant.
That, in short, is why we think working on AI safety and governance in China and/or building mutual understanding between Chinese and Western actors in these areas is likely to be one of the most promising China-related career paths. Below we provide more arguments and detailed information on this option.
If you are interested in pursuing a career path described in this profile, contact 80,000 Hours’ one-on-one team and we may be able to put you in touch with a specialist advisor.
narrator_time: 4h
narrator: pw
editor: pw
qa: mds
client: 80000_hours
project_id: 80000_hours
https://www.lesswrong.com/posts/6LzKRP88mhL9NKNrS/how-my-team-at-lightcone-sometimes-gets-stuff-done
Disclaimer: I originally wrote this as a private doc for the Lightcone team. I then showed it to John and he said he would pay me to post it here. That sounded awfully compelling. However, I wanted to note that I’m an early founder who hasn't built anything truly great yet. I’m writing this doc because as Lightcone is growing, I have to take a stance on these questions. I need to design our org to handle more people. Still, I haven’t seen the results long-term, and who knows if this is good advice. Don’t overinterpret this.
Suppose you went up on stage in front of a company you founded, that now had grown to 100, or 1000, 10 000+ people. You were going to give a talk about your company values. You can say things like “We care about moving fast, taking responsibility, and being creative” -- but I expect these words would mostly fall flat. At the end of the day, the path the water takes down the hill is determined by the shape of the territory, not the sound the water makes as it swooshes by. To manage that many people, it seems to me you need clear, concrete instructions. What are those? What are things you could write down on a piece of paper and pass along your chain of command, such that if at the end people go ahead and just implement them, without asking what you meant, they would still preserve some chunk of what makes your org work?
narrator_time: 1h15m
editing_time: 0h15m
narrator: pw
editor: pw
qa: mds
client: LessWrong
project_id: LessWrong Curated
Read the original writing here.
In this note I’ll summarize the bio-anchors report, describe my initial reactions to it, and take a closer look at two disagreements that I have with background assumptions used by (readers of) the report. (Thanks to Steph Lin for comments on a draft of this review.)
Summary of the report This report attempts to forecast the year when the amount of compute required to train a transformative AI (TAI) model will first become available, as the year when a forecast for the amount of compute required to train TAI in a given year will intersect a forecast for the amount of compute that will be available for a training run of a single project in a given year.
The report estimates the former by first estimating the amount of compute needed to train TAI in 2020 assuming that 2020 algorithms anchored to the sizes of various biological processes can scale to TAI, then correcting for how this requirement may fall over time due to incremental algorithmic progress. It estimates the latter by multiplying an estimate for compute available per dollar in a given year with an estimate for dollars available per training run in a given year.
Most of the report focuses on generating the 2020 training compute requirements distribution.
narrator_time: 3h20m
editing_time:
narrator: pw
editor: pw
qa: mds
client: EA Forum
project_id: EA Forum Red Team Prize
https://forum.effectivealtruism.org/posts/nGrmemHzQvBpnXkNX/what-matters-to-shrimps-factors-affecting-shrimp-welfare-in
This is a linkpost for https://www.shrimpwelfareproject.org/shrimp-welfare-report
Shrimp Welfare Project (SWP) produced this report to guide our decision making on funding for further research into shrimp welfare and on which interventions to allocate our resources. We are cross-posting this on the forum because we think it may be useful to share the complexity of understanding the needs of beneficiaries who cannot communicate with us. We also hope it will be useful for other organisations working on shrimp welfare, and it’s also hopefully an interesting read!
The report was written by Lucas Lewit-Mendes, with detailed feedback provided by Sasha Saugh and Aaron Boddy. We are thankful for and build on the work and feedback of other NGOs, including Charity Entrepreneurship, Rethink Priorities, Aquatic Life Institute, Fish Welfare Initiative, Compassion in World Farming and Crustacean Compassion. All errors and shortcomings are our own.
Executive SummaryWhile many environmental conditions and farming practices could plausibly affect the welfare of shrimps, little research has been done to assess which factors most affect shrimp welfare.
This report aims to assess the importance of various factors for the welfare of farmed shrimps, with a particular focus on Litopenaeus vannamei (also known as Penaeus vannamei, or whiteleg shrimp), due to the scale and intensity of farming (~171-405 billion globally per annum) (Mood and Brooke, 2019). Where evidence is scarce, we extend our research to other shrimps, other decapods, or even other aquatic animals. Further research into the most significant factors and practices affecting farmed shrimp welfare is needed.
Conclusions from our review are summarised below:
Eyestalk Ablation: Shrimps demonstrate aversive behavioural responses to eyestalk ablation, and applying anaesthesia before ablation has therapeutic effects. We believe this is strongly indicative that eyestalk ablation is a welfare concern.
Disease: Infectious diseases cause significant mortality events. This is likely to both cause suffering prior to death and increase the total number of shrimps who are farmed and experience suffering.
Stunning and Slaughter: Current slaughter practices (asphyxiation or immersion in ice slurry) are likely to be inhumane. While evidence on the optimal slaughter method for shrimps is limited, electrical stunning appears to be the most promising method to effectively stun and kill shrimps.
Stocking Density: There is strong experimental evidence to suggest that reductions in stocking density indirectly improve welfare by improving water quality, reducing disease, and increasing survival. There is also some tentative evidence that stocking density directly impacts shrimp behaviour and measurable stress biomarkers (e.g. serotonin).
Environmental Enrichment (EE): Environmental enrichments
https://michaelnotebook.com/eanotes/
*Long and rough notes on Effective Altruism (EA). Written to help me get to the bottom of several questions: what do I like and think is important about EA? Why do I find the mindset so foreign? Why am I not an EA? And to start me thinking about: what do alternatives to EA look like? The notes are not aimed at effective altruists, though they may perhaps be of interest to EA-adjacent people. Thoughtful, informed comments and corrections welcome (especially detailed, specific corrections!) - see the comment area at the bottom.
narrator_time: 2h45m
editing_time: 1h30m
narrator: pw
editor: pw
qa: mds
client: EA Forum
project_id: Redteam Prize
---*
https://www.lesswrong.com/posts/rP66bz34crvDudzcJ/decision-theory-does-not-imply-that-we-get-to-have-nice
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.
(Note: I wrote this with editing help from Rob and Eliezer. Eliezer's responsible for a few of the paragraphs.)
A common confusion I see in the tiny fragment of the world that knows about logical decision theory (FDT/UDT/etc.), is that people think LDT agents are genial and friendly for each other.[1]
One recent example is Will Eden’s tweet about how maybe a molecular paperclip/squiggle maximizer would leave humanity a few stars/galaxies/whatever on game-theoretic grounds. (And that's just one example; I hear this suggestion bandied around pretty often.)
I'm pretty confident that this view is wrong (alas), and based on a misunderstanding of LDT. I shall now attempt to clear up that confusion.
To begin, a parable: the entity Omicron (Omega's little sister) fills box A with $1M and box B with $1k, and puts them both in front of an LDT agent saying "You may choose to take either one or both, and know that I have already chosen whether to fill the first box". The LDT agent takes both.
"What?" cries the CDT agent. "I thought LDT agents one-box!"
LDT agents don't cooperate because they like cooperating. They don't one-box because the name of the action starts with an 'o'. They maximize utility, using counterfactuals that assert that the world they are already in (and the observations they have already seen) can (in the right circumstances) depend (in a relevant way) on what they are later going to do.
A paperclipper cooperates with other LDT agents on a one-shot prisoner's dilemma because they get more paperclips that way. Not because it has a primitive property of cooperativeness-with-similar-beings. It needs to get the more paperclips.
narrator_time: 1h40m
editing_time: 0h50m
narrator: pw
editor: pw
qa: mds
client: LessWrong
project_id: LessWrong Curated
https://forum.effectivealtruism.org/posts/cXBznkfoPJAjacFoT/are-you-really-in-a-race-the-cautionary-tales-of-szilard-and#Summary
OR The Tragedy of the Einstein Letter and the Gaither Report; Cautionary Lessons from the Manhattan Project and the ‘Missile Gap’; Beware Assuming You’re in an AI Race; The illusory Atomic Gap, the illusory Missile Gap and the AGI Gap
SummaryIn both the 1940s and 1950s, well-meaning and good people – the brightest of their generation – were convinced they were in an existential race with an expansionary, totalitarian regime. Because of this belief, they advocated for and participated in a ‘sprint’ race: the Manhattan Project to develop a US atomic bomb (1939-1945); and the ‘missile gap’ project to build up a US ICBM capability (1957-1962). These were both based on a mistake, however - the Nazis decided against a Manhattan Project in 1942, and the Soviets decided against an ICBM build-up in 1958. The main consequence of both was to unilaterally speed up dangerous developments and increase existential risk. Key participants, such as Albert Einstein and Daniel Ellsberg, described their involvement as the greatest mistake of their life.
Our current situation with AGI shares certain striking similarities and certain lessons suggest themselves: make sure you’re actually in a race (information on whether you are is very valuable), be careful when secrecy is emphasised, and don’t give up your power as an expert too easily.
I briefly cover the two case studies, discuss the atmosphere at RAND, then draw the comparison with AGI and explain my three takeaways. This short piece is mainly based on Richard Rhodes’ The Making of the Atomic Bomb and Daniel Ellsberg’s The Doomsday Machine. It was inspired by a Slack discussion with Di Cooke.
narrator_time: 2h30m
editing_time: 1h
narrator: pw
editor: pw
qa: km
client: EA Forum
project_id: redteam prize
https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like#2022
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.
This was written for the Vignettes Workshop.[1] The goal is to write out a detailed future history (“trajectory”) that is as realistic (to me) as I can currently manage, i.e. I’m not aware of any alternative trajectory that is similarly detailed and clearly more plausible to me. The methodology is roughly: Write a future history of 2022. Condition on it, and write a future history of 2023. Repeat for 2024, 2025, etc. (I'm posting 2022-2026 now so I can get feedback that will help me write 2027+. I intend to keep writing until the story reaches singularity/extinction/utopia/etc.)
What’s the point of doing this? Well, there are a couple of reasons:
https://forum.effectivealtruism.org/posts/dQvDxDMyueLyydHw4/population-ethics-without-axiology-a-framework#fnref-zciTjACpXYMxgfpCF-31
This post introduces a framework for thinking about population ethics: “population ethics without axiology.” In its last section, I sketch the implications of adopting my framework for evaluating the thesis of longtermism.
Before explaining what’s different about my proposal, I’ll describe what I understand to be the standard approach it seeks to replace, which I call “axiology-focused.” (Skip to the section SUMMARY: “Population ethics without axiology” for a summary of my proposal.)
The axiology-focused approach goes as follows. First, there’s the search for axiology, a theory of (intrinsic) value. (E.g., the axiology may state that good experiences are what’s valuable.) Then, there’s further discussion on whether ethics contains other independent parts or whether everything derives from that axiology. For instance, a consequentialist may frame their disagreement with deontology as follows. “Consequentialism is the view that making the world a better place is all that matters, while deontologists think that other things (e.g., rights, duties) matter more.” Similarly, someone could frame population-ethical disagreements as follows. “Some philosophers think that all that matters is more value in the world and less disvalue (“totalism”). Others hold that further considerations also matter – for instance, it seems odd to compare someone’s existence to never having been born, so we can discuss what it means to benefit a person in such contexts.”
In both examples, the discussion takes for granted that there’s something that’s valuable in itself. The still-open questions come afterward, after “here’s what’s valuable.”
narrator_time: TBC
editing_time: TBC
narrator: pw
editor: pw
qa: km
client: ea_forum
project_id: TBC
https://forum.effectivealtruism.org/posts/AjxqsDmhGiW9g8ju6/effective-altruism-in-the-garden-of-ends
Huge thanks to Alex Zhu, Anders Sandberg, Andrés Gómez Emilsson, Andrew Roberts, Anne-Lorraine Selke (who I've subbed in entire sentences from), Crichton Atkinson, Ellie Hain, George Walker, Jesper Östman, Joe Edelman, Liza Simonova, Kathryn Devaney, Milan Griffes, Morgan Sutherland, Nathan Young, Rafael Ruiz, Tasshin Fogelman, Valerie Zhang, and Xiq for reviewing or helping me develop my ideas here. Further thanks to Allison Duettmann, Anders Sandberg, Howie Lempel, Julia Wise, and Tildy Stokes, for inspiring me through their lived examples.
I did not believe that a Cause which stood for a beautiful ideal […] should demand the denial of life & joy. – Emma Goldman, Living My Life
This essay is a reconciliation of moral commitment and the good life. Here is its essence in two paragraphs:
Totalized by an ought, I sought its source outside myself. I found nothing. The ought came from me, an internal whip toward a thing which, confusingly, I already wanted – to see others flourish. I dropped the whip. My want now rested, commensurate, amidst others of its kind – terminal wants for ends-in-themselves: loving, dancing, and the other spiritual requirements of my particular life. To say that these were lesser seemed to say, “It is more vital and urgent to eat well than to drink or sleep well.” No – I will eat, sleep, and drink well to feel alive; so too will I love and dance as well as help.
Once, the material requirements of life were in competition: If we spent time building shelter it might jeopardize daylight that could have been spent hunting. We built communities to take the material requirements of life out of competition. For many of us, the task remains to do the same for our spirits. Particularly so for those working outside of organized religion on huge, consuming causes. I suggest such a community might practice something like “fractal altruism,” taking the good life at the scale of its individuals out of competition with impact at the scale of the world.
If you’re a Blinkist or Sparknotes person, you can stop here.
narrator_time: TBC
editing_time: TBC
narrator: pw
editor: pw
qa: km
client: ea_forum
project_id: TBC
https://forum.effectivealtruism.org/posts/KqCybin8rtfP3qztq/agi-and-lock-in
This is a linkpost for https://docs.google.com/document/d/1mkLFhxixWdT5peJHq4rfFzq4QbHyfZtANH1nou68q88/edit#
The long-term future of intelligent life is currently unpredictable and undetermined. In the linked document, we argue that the invention of artificial general intelligence (AGI) could change this by making extreme types of lock-in technologically feasible. In particular, we argue that AGI would make it technologically feasible to (i) perfectly preserve nuanced specifications of a wide variety of values or goals far into the future, and (ii) develop AGI-based institutions that would (with high probability) competently pursue any such values for at least millions, and plausibly trillions, of years.
The rest of this post contains the summary (6 pages), with links to relevant sections of the main document (40 pages) for readers who want more details.
0.0 The claim
Life on Earth could survive for millions of years. Life in space could plausibly survive for trillions of years. What will happen to intelligent life during this time? Some possible claims are:
A. Humanity will almost certainly go extinct in the next million years.
B. Under Darwinian pressures, intelligent life will spread throughout the stars and rapidly evolve toward maximal reproductive fitness.
C. Through moral reflection, intelligent life will reliably be driven to pursue some specific “higher” (non-reproductive) goal, such as maximizing the happiness of all creatures.
narrator_time: TBC
editing_time: TBC
narrator: pw
editor: pw
qa: km
client: ea_forum
project_id: TBC
https://forum.effectivealtruism.org/posts/zoWypGfXLmYsDFivk/counterarguments-to-the-basic-ai-risk-case
This is cross-posted from the AI Impacts blog
This is going to be a list of holes I see in the basic argument for existential risk from superhuman AI systems1.
To start, here’s an outline of what I take to be the basic case2:
I. If superhuman AI systems are built, any given system is likely to be ‘goal-directed’
Reasons to expect this:
II. If goal-directed superhuman AI systems are built, their desired outcomes will probably be about as bad as an empty universe by human lights---
narrator_time: TBC
editing_time: TBC
narrator: pw
editor: pw
qa: km
client: ea_forum
project_id: TBC
https://www.lesswrong.com/posts/7cHgjJR2H5e4w4rxT/quintin-s-alignment-papers-roundup-week-1
Introduction
I've decided to start a weekly roundup of papers that seem relevant to alignment, focusing on papers or approaches that might be new to safety researchers. Unlike the Alignment Newsletter, I'll be spending relatively little effort on summarizing the papers. I'll just link them, copy their abstracts, and potentially describe some of my thoughts on how the paper relates to alignment. Hopefully, this will let me keep to a weekly schedule.
The purpose of this series isn't so much to share insights directly with the reader, but instead to make them aware of already existing research that may be relevant to the reader's own research.
narrator_time: TBC
narrator: pw
qa: km
client: lesswrong
https://www.lesswrong.com/posts/rCJQAkPTEypGjSJ8X/how-might-we-align-transformative-ai-if-it-s-developed-very
This post is part of my AI strategy nearcasting series: trying to answer key strategic questions about transformative AI, under the assumption that key events will happen very soon, and/or in a world that is otherwise very similar to today's.
This post gives my understanding of what the set of available strategies for aligning transformative AI would be if it were developed very soon, and why they might or might not work. It is heavily based on conversations with Paul Christiano, Ajeya Cotra and Carl Shulman, and its background assumptions correspond to the arguments Ajeya makes in this piece (abbreviated as “Takeover Analysis”).
I premise this piece on a nearcast in which a major AI company (“Magma,” following Ajeya’s terminology) has good reason to think that it can develop transformative AI very soon (within a year), using what Ajeya calls “human feedback on diverse tasks” (HFDT) - and has some time (more than 6 months, but less than 2 years) to set up special measures to reduce the risks of misaligned AI before there’s much chance of someone else deploying transformative AI.
narrator_time: TBC
narrator: pw
qa: km
client: lesswrong
https://www.lesswrong.com/posts/gNodQGNoPDjztasbh/lies-damn-lies-and-fabricated-options
This is an essay about one of those "once you see it, you will see it everywhere" phenomena. It is a psychological and interpersonal dynamic roughly as common, and almost as destructive, as motte-and-bailey, and at least in my own personal experience it's been quite valuable to have it reified, so that I can quickly recognize the commonality between what I had previously thought of as completely unrelated situations.
The original quote referenced in the title is "There are three kinds of lies: lies, damned lies, and statistics."
Background 1: Gyroscopes
Gyroscopes are weird.
Except they're not. They're quite normal and mundane and straightforward. The weirdness of gyroscopes is a map-territory confusion—gyroscopes seem weird because my map is poorly made, and predicts that they will do something other than their normal, mundane, straightforward thing.
In large part, this is because I don't have the consequences of physical law engraved deeply enough into my soul that they make intuitive sense.
I can imagine a world that looks exactly like the world around me, in every way, except that in this imagined world, gyroscopes don't have any of their strange black-magic properties. It feels coherent to me. It feels like a world that could possibly exist.
"Everything's the same, except gyroscopes do nothing special." Sure, why not.
But in fact, this world is deeply, deeply incoherent. It is Not Possible with capital letters. And a physicist with sufficiently sharp intuitions would know this—would be able to see the implications of a world where gyroscopes "don't do anything weird," and tell me all of the ways in which reality falls apart.
narrator_time: TBC
narrator: pw
qa: km
client: lesswrong
https://www.lesswrong.com/posts/fFY2HeC9i2Tx8FEnK/my-resentful-story-of-becoming-a-medical-miracle
This is a linkpost for https://acesounderglass.com/2022/10/13/my-resentful-story-of-becoming-a-medical-miracle/
You know those health books with “miracle cure” in the subtitle? The ones that always start with a preface about a particular patient who was completely hopeless until they tried the supplement/meditation technique/healing crystal that the book is based on? These people always start broken and miserable, unable to work or enjoy life, perhaps even suicidal from the sheer hopelessness of getting their body to stop betraying them. They’ve spent decades trying everything and nothing has worked until their friend makes them see the book’s author, who prescribes the same thing they always prescribe, and the patient immediately stands up and starts dancing because their problem is entirely fixed (more conservative books will say it took two sessions). You know how those are completely unbelievable, because anything that worked that well would go mainstream, so basically the book is starting you off with a shit test to make sure you don’t challenge its bullshit later?
Well 5 months ago I became one of those miraculous stories, except worse, because my doctor didn’t even do it on purpose. This finalized some already fermenting changes in how I view medical interventions and research. Namely: sometimes knowledge doesn’t work and then you have to optimize for luck.
I assure you I’m at least as unhappy about this as you are.
narrator_time: TBC
narrator: pw
qa: km
client: lesswrong
https://www.lesswrong.com/posts/REA49tL5jsh69X3aM/introduction-to-abstract-entropy
This post, and much of the following sequence, was greatly aided by feedback from the following people (among others): Lawrence Chan, Joanna Morningstar, John Wentworth, Samira Nedungadi, Aysja Johnson, Cody Wild, Jeremy Gillen, Ryan Kidd, Justis Mills and Jonathan Mustin. Illustrations by Anne Ore.
Introduction & motivation
In the course of researching optimization, I decided that I had to really understand what entropy is.[1] But there are a lot of other reasons why the concept is worth studying:
narrator_time: TBC
narrator: pw
qa: km
client: lesswrong
https://www.lesswrong.com/posts/8vesjeKybhRggaEpT/consider-your-appetite-for-disagreements
Poker
There was a time about five years ago where I was trying to get good at poker. If you want to get good at poker, one thing you have to do is review hands. Preferably with other people.
For example, suppose you have ace king offsuit on the button. Someone in the highjack opens to 3 big blinds preflop. You call. Everyone else folds. The flop is dealt. It's a rainbow Q75. You don't have any flush draws. You missed. Your opponent bets. You fold. They take the pot and you move to the next hand.
Once you finish your session, it'd be good to come back and review this hand. Again, preferably with another person. To do this, you would review each decision point in the hand. Here, there were two decision points.
The first was when you faced a 3BB open from HJ preflop with AKo. In the hand, you decided to call. However, this of course wasn't your only option. You had two others: you could have folded, and you could have raised. Actually, you could have raised to various sizes. You could have raised small to 8BB, medium to 10BB, or big to 12BB. Or hell, you could have just shoved 200BB! But that's not really a realistic option, nor is folding. So in practice your decision was between calling and raising to various realistic sizes.
narrator_time: TBC
narrator: pw
qa: km
client: lesswrong
https://www.lesswrong.com/posts/K4urTDkBbtNuLivJx/why-i-think-strong-general-ai-is-coming-soon
I think there is little time left before someone builds AGI (median ~2030). Once upon a time, I didn't think this.
This post attempts to walk through some of the observations and insights that collapsed my estimates.
The core ideas are as follows:
Some notes up front
narrator_time: TBC
narrator: pw
qa: km
client: lesswrong
https://forum.effectivealtruism.org/posts/coryFCkmcMKdJb7Pz/does-economic-growth-meaningfully-improve-well-being-an
https://forum.effectivealtruism.org/posts/rXYW9GPsmwZYu3doX/what-happens-on-the-average-day
narrator_time: TBC
editing_time: TBC
narrator: pw
editor: pw
qa: km
client: 80000_hours
project_id: TBC
https://www.lesswrong.com/posts/nbq2bWLcYmSGup9aF/a-transparency-and-interpretability-tech-tree
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.
Thanks to Chris Olah, Neel Nanda, Kate Woolverton, Richard Ngo, Buck Shlegeris, Daniel Kokotajlo, Kyle McDonell, Laria Reynolds, Eliezer Yudkowksy, Mark Xu, and James Lucassen for useful comments, conversations, and feedback that informed this post.
The more I have thought about AI safety over the years, the more I have gotten to the point where the only worlds I can imagine myself actually feeling good about humanity’s chances are ones in which we have powerful transparency and interpretability tools that lend us insight into what our models are doing as we are training them.[1] Fundamentally, that’s because if we don’t have the feedback loop of being able to directly observe how the internal structure of our models changes based on how we train them, we have to essentially get that structure right on the first try—and I’m very skeptical of humanity’s ability to get almost anything right on the first try, if only just because there are bound to be unknown unknowns that are very difficult to predict in advance.
Certainly, there are other things that I think are likely to be necessary for humanity to succeed as well—e.g. convincing leading actors to actually use such transparency techniques, having a clear training goal that we can use our transparency tools to enforce, etc.—but I currently feel that transparency is the least replaceable necessary condition and yet the one least likely to be solved by default.
Nevertheless, I do think that it is a tractable problem to get to the point where transparency and interpretability is reliably able to give us the sort of insight into our models that I think is necessary for humanity to be in a good spot. I think many people who encounter transparency and interpretability, however, have a hard time envisioning what it might look like to actually get from where we are right now to where we need to be. Having such a vision is important both for enabling us to better figure out how to make that vision into reality and also for helping us tell how far along we are at any point—and thus enabling us to identify at what point we’ve reached a level of transparency and interpretability that we can trust it to reliably solve different sorts of alignment problems.
The goal of this post, therefore, is to attempt to lay out such a vision by providing a “tech tree” of transparency and interpretability problems, with each successive problem tackling harder and harder parts of what I see as the core difficulties. This will only be my tech tree, in terms of the relative difficulties, dependencies, and orderings that I expect as we make transparency and interpretability progress—I could, and probably will, be wrong in various ways, and I’d encourage others to try to build their own tech trees to represent their pictures of progress as well.
narrator_time: 4h30m
narrator: pw
qa: km
client: lesswrong
By Benjamin Hilton.
Abstract:
AI might bring huge benefits—if we avoid the risks.
Source URL:
https://80000hours.org/problem-profiles/artificial-intelligence/
By Benjamin Hilton.
Abstract:
AI might bring huge benefits—if we avoid the risks.
Source URL:
https://80000hours.org/problem-profiles/artificial-intelligence/