The Nonlinear Library allows you to easily listen to top EA and rationalist content on your podcast player. We use text-to-speech software to create an automatically updating repository of audio content from the EA Forum, Alignment Forum, LessWrong, and other EA blogs. To find out more, please visit us at nonlinear.org
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: A Golden Age of Building? Excerpts and lessons from Empire State, Pentagon, Skunk Works and SpaceX, published by jacobjacob on September 1, 2023 on LessWrong.Patrick Collison has a fantastic list of examples of people quickly accomplishing ambitious things together since the 19th Century. It does make you yearn for a time that feels... different, when the lethargic behemoths of government departments could move at the speed of a racing startup:[...] last century, [the Department of Defense] innovated at a speed that puts modern Silicon Valley startups to shame: the Pentagon was built in only 16 months (1941-1943), the Manhattan Project ran for just over 3 years (1942-1946), and the Apollo Program put a man on the moon in under a decade (1961-1969). In the 1950s alone, the United States built five generations of fighter jets, three generations of manned bombers, two classes of aircraft carriers, submarine-launched ballistic missiles, and nuclear-powered attack submarines.[Note: that paragraph is from a different post.]Inspired by partly by Patrick's list, I spent some of my vacation reading and learning about various projects from this Lost Age. I then wrote up a memo to share highlights and excerpts with my colleagues at Lightcone.After that, some people encouraged me to share the memo more widely -- and I do think it's of interest to anyone who harbors an ambition for greatness and a curiosity about operating effectively.How do you build the world's tallest building in only a year? The world's largest building in the same amount of time? Or America's first fighter jet in just 6 months?How??Writing this post felt like it helped me gain at least some pieces of this puzzle. If anyone has additional pieces, I'd love to hear them in the comments.Empire State BuildingThe Empire State was the tallest building in the world upon completion in April 1931. Over my vacation I read a rediscovered 1930s notebook, written by the general contractors themselves. It details the construction process and the organisation of the project.I will share some excerpts, but to contextualize them, consider first some other skyscrapers built more recently:Design startConstruction endTotal timeBurj Khalifa200420106 yearsShanghai Tower200820157 yearsAbraj Al-Balt2002201210 yearsOne World Trade Center200520149 yearsNordstrom Tower2010202010 yearsTaipei 101199720047 years(list from skyscrapercenter.com)Now, from the Empire State book's foreword:The most astonishing statistics of the Empire State was the extraordinary speed with which it was planned and constructed. [...] There are different ways to describe this feat. Six months after the setting of the first structural columns on April 7, 1930, the steel frame topped off on the eighty-sixth floor. The fully enclosed building, including the mooring mast that raised its height to the equivalent of 102 stories, was finished in eleven months, in March 1931. Most amazing though, is the fact that within just twenty months -- from the first signed contractors with the architects in September 1929 to opening-day ceremonies on May 1, 1931 -- the Empire State was designed, engineered, erected, and ready for tenants.Within this time, the architectural drawings and plans were prepared, the Vicitorian pile of the Waldorf-Astoria hotel was demolished [demolition started only two days after the initial agreement was signed], the foundations and grillages were dug and set, the steel columns and beams, some 57,000 tons, were fabricated and milled to precise specifications, ten million common bricks were laid, more than 62,000 cubic yards of concrete were poured, 6,400 windows were set, and sixty-seven elevators were installed in seven miles of shafts. At peak activity, 3,500 workers were employed on site, and the frame rose more than a story a day,...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Against Almost Every Theory of Impact of Interpretability, published by Charbel-Raphaël on August 17, 2023 on LessWrong.Epistemic Status: I believe I am well-versed in this subject. I erred on the side of making claims that were too strong and allowing readers to disagree and start a discussion about precise points rather than trying to edge-case every statement. I also think that using memes is important because safety ideas are boring and anti-memetic. So let's go!Many thanks to @scasper, @Sid Black , @Neel Nanda , @Fabien Roger , @Bogdan Ionut Cirstea, @WCargo, @Alexandre Variengien, @Jonathan Claybrough, @Edoardo Pona, @Andrea_Miotti, Diego Dorn, Angélina Gentaz, Clement Dumas, and Enzo Marsot for useful feedback and discussions.When I started this post, I began by critiquing the article A Long List of Theories of Impact for Interpretability, from Neel Nanda, but I later expanded the scope of my critique. Some ideas which are presented are not supported by anyone, but to explain the difficulties, I still need to 1. explain them and 2. criticize them. It gives an adversarial vibe to this post. I'm sorry about that, and I think that doing research into interpretability, even if it's no longer what I consider a priority, is still commendable.How to read this document? Most of this document is not technical, except for the section "What does the end story of interpretability look like?" which can be mostly skipped at first. I expect this document to also be useful for people not doing interpretability research. The different sections are mostly independent, and I've added a lot of bookmarks to help modularize this post.If you have very little time, just read (this is also the part where I'm most confident):Auditing deception with Interp is out of reach (4 min)Enumerative safety critique (2 min)Technical Agendas with better Theories of Impact (1 min)Here is the list of claims that I will defend:(bolded sections are the most important ones)The overall Theory of Impact is quite poorInterp is not a good predictor of future systemsAuditing deception with interp is out of reachWhat does the end story of interpretability look like? That's not clear at all.Enumerative safety?Reverse engineering?Olah's Interpretability dream?Retargeting the search?Relaxed adversarial training?Microscope AI?Preventive measures against Deception seem much more workableSteering the world towards transparencyCognitive Emulations - Explainability By designInterpretability May Be Overall HarmfulOutside view: The proportion of junior researchers doing Interp rather than other technical work is too highSo far my best ToI for interp: Nerd Sniping?Even if we completely solve interp, we are still in dangerTechnical Agendas with better Theories of ImpactConclusionNote: The purpose of this post is to criticize the Theory of Impact (ToI) of interpretability for deep learning models such as GPT-like models, and not the explainability and interpretability of small models.The emperor has no clothes?I gave a talk about the different risk models, followed by an interpretability presentation, then I got a problematic question, "I don't understand, what's the point of doing this?" Hum.Feature viz? (left image) Um, it's pretty but is this useful? Is this reliable?GradCam (a pixel attribution technique, like on the above right figure), it's pretty. But I've never seen anybody use it in industry. Pixel attribution seems useful, but accuracy remains the king.Induction heads? Ok, we are maybe on track to retro engineer the mechanism of regex in LLMs. Cool.The considerations in the last bullet points are based on feeling and are not real arguments. Furthermore, most mechanistic interpretability isn't even aimed at being useful right now. But in the rest of the post, we'll find out if...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: My current LK99 questions, published by Eliezer Yudkowsky on August 1, 2023 on LessWrong. So this morning I thought to myself, "Okay, now I will actually try to study the LK99 question, instead of betting based on nontechnical priors and market sentiment reckoning." (My initial entry into the affray, having been driven by people online presenting as confidently YES when the prediction markets were not confidently YES.) And then I thought to myself, "This LK99 issue seems complicated enough that it'd be worth doing an actual Bayesian calculation on it"--a rare thought; I don't think I've done an actual explicit numerical Bayesian update in at least a year. In the process of trying to set up an explicit calculation, I realized I felt very unsure about some critically important quantities, to the point where it no longer seemed worth trying to do the calculation with numbers. This is the System Working As Intended. On July 30th, Danielle Fong said of this temperature-current-voltage graph, 'Normally as current increases, voltage drop across a material increases. in a superconductor, voltage stays nearly constant, 0. that appears to be what's happening here -- up to a critical current. with higher currents available at lower temperatures deeply in the "fraud or superconduct" territory, imo. like you don't get this by accident -- you either faked it, or really found something.' The graph Fong is talking about only appears in the initial paper put forth by Young-Wan Kwon, allegedly without authorization. A different graph, though similar, appears in Fig. 6 on p. 12 of the 6-author LK-endorsed paper rushed out in response. Is it currently widely held by expert opinion, that this diagram has no obvious or likely explanation except "superconductivity" or "fraud"? If the authors discovered something weird that wasn't a superconductor, or if they just hopefully measured over and over until they started getting some sort of measurement error, is there any known, any obvious way they could have gotten the same graph? One person alleges an online rumor that poorly connected electrical leads can produce the same graph. Is that a conventional view? Alternatively: If this material is a superconductor, have we seen what we expected to see? Is the diminishing current capacity with increased temperature usual? How does this alleged direct measurement of superconductivity square up with the current-story-as-I-understood-it that the material is only being very poorly synthesized, probably only in granules or gaps, and hence only detectable by looking for magnetic resistance / pinning? This is my number-one question. Call it question 1-NO, because it's the question of "How does the NO story explain this graph, and how prior-improbable or prior-likely was that story?", with respect to my number one question. Though I'd also like to know the 1-YES details: whether this looks like a high-prior-probability superconductivity graph; or a graph that requires a new kind of superconductivity, but one that's theoretically straightforward given a central story; or if it looks like unspecified weird superconductivity, with there being no known theory that predicts a graph looking roughly like this. What's up with all the partial levitation videos? Possibilities I'm currently tracking: 2-NO-A: There's something called "diamagnetism" which exists in other materials. The videos by LK and attempted replicators show the putative superconductor being repelled from the magnet, but not being locked in space relative to the magnet. Superconductors are supposed to exhibit Meissner pinning, and the failure of the material to be pinned to the magnet indicates that this isn't a superconductor. (Sabine Hossenfelder seems to talk this way here. "I lost hope when I saw this video; this doesn't look like the Meissner ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Yes, It's Subjective, But Why All The Crabs?, published by johnswentworth on July 28, 2023 on LessWrong. Crabs Nature really loves to evolve crabs. Some early biologist, equipped with knowledge of evolution but not much else, might see all these crabs and expect a common ancestral lineage. That's the obvious explanation of the similarity, after all: if the crabs descended from a common ancestor, then of course we'd expect them to be pretty similar. . but then our hypothetical biologist might start to notice surprisingly deep differences between all these crabs. The smoking gun, of course, would come with genetic sequencing: if the crabs' physiological similarity is achieved by totally different genetic means, or if functionally-irrelevant mutations differ across crab-species by more than mutational noise would induce over the hypothesized evolutionary timescale, then we'd have to conclude that the crabs had different lineages. (In fact, historically, people apparently figured out that crabs have different lineages long before sequencing came along.) Now, having accepted that the crabs have very different lineages, the differences are basically explained. If the crabs all descended from very different lineages, then of course we'd expect them to be very different. . but then our hypothetical biologist returns to the original empirical fact: all these crabs sure are very similar in form. If the crabs all descended from totally different lineages, then the convergent form is a huge empirical surprise! The differences between the crab have ceased to be an interesting puzzle - they're explained - but now the similarities are the interesting puzzle. What caused the convergence? To summarize: if we imagine that the crabs are all closely related, then any deep differences are a surprising empirical fact, and are the main remaining thing our model needs to explain. But once we accept that the crabs are not closely related, then any convergence/similarity is a surprising empirical fact, and is the main remaining thing our model needs to explain. Agents A common starting point for thinking about "What are agents?" is Dennett's intentional stance: Here is how it works: first you decide to treat the object whose behavior is to be predicted as a rational agent; then you figure out what beliefs that agent ought to have, given its place in the world and its purpose. Then you figure out what desires it ought to have, on the same considerations, and finally you predict that this rational agent will act to further its goals in the light of its beliefs. A little practical reasoning from the chosen set of beliefs and desires will in most instances yield a decision about what the agent ought to do; that is what you predict the agent will do. Daniel Dennett, The Intentional Stance, p. 17 One of the main interesting features of the intentional stance is that it hypothesizes subjective agency: I model a system as agentic, and you and I might model different systems as agentic. Compared to a starting point which treats agency as objective, the intentional stance neatly explains many empirical facts - e.g. different people model different things as agents at different times. Sometimes I model other people as planning to achieve goals in the world, sometimes I model them as following set scripts, and you and I might differ in which way we're modeling any given person at any given time. If agency is subjective, then the differences are basically explained. . but then we're faced with a surprising empirical fact: there's a remarkable degree of convergence among which things people do-or-don't model as agentic at which times. Humans yes, rocks no. Even among cases where people disagree, there are certain kinds of arguments/evidence which people generally agree update in a certain direction - e.g. ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Alignment Grantmaking is Funding-Limited Right Now, published by johnswentworth on July 19, 2023 on LessWrong. For the past few years, I've generally mostly heard from alignment grantmakers that they're bottlenecked by projects/people they want to fund, not by amount of money. Grantmakers generally had no trouble funding the projects/people they found object-level promising, with money left over. In that environment, figuring out how to turn marginal dollars into new promising researchers/projects - e.g. by finding useful recruitment channels or designing useful training programs - was a major problem. Within the past month or two, that situation has reversed. My understanding is that alignment grantmaking is now mostly funding-bottlenecked. This is mostly based on word-of-mouth, but for instance, I heard that the recent lightspeed grants round received far more applications than they could fund which passed the bar for basic promising-ness. I've also heard that the Long-Term Future Fund (which funded my current grant) now has insufficient money for all the grants they'd like to fund. I don't know whether this is a temporary phenomenon, or longer-term. Alignment research has gone mainstream, so we should expect both more researchers interested and more funders interested. It may be that the researchers pivot a bit faster, but funders will catch up later. Or, it may be that the funding bottleneck becomes the new normal. Regardless, it seems like grantmaking is at least funding-bottlenecked right now. Some takeaways: If you have a big pile of money and would like to help, but haven't been donating much to alignment because the field wasn't money constrained, now is your time! If this situation is the new normal, then earning-to-give for alignment may look like a more useful option again. That said, at this point committing to an earning-to-give path would be a bet on this situation being the new normal. Grants for upskilling, training junior people, and recruitment make a lot less sense right now from grantmakers' perspective. For those applying for grants, asking for less money might make you more likely to be funded. (Historically, grantmakers consistently tell me that most people ask for less money than they should; I don't know whether that will change going forward, but now is an unusually probable time for it to change.) Note that I am not a grantmaker, I'm just passing on what I hear from grantmakers in casual conversation. If anyone with more knowledge wants to chime in, I'd appreciate it. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Accidentally Load Bearing, published by jefftk on July 13, 2023 on LessWrong. Sometimes people will talk about Chesterton's Fence, the idea that if you want to change something - removing an apparently useless fence - you should first determine why it was set up that way: The gate or fence did not grow there. It was not set up by somnambulists who built it in their sleep. It is highly improbable that it was put there by escaped lunatics who were for some reason loose in the street. Some person had some reason for thinking it would be a good thing for somebody. And until we know what the reason was, we really cannot judge whether the reason was reasonable. It is extremely probable that we have overlooked some whole aspect of the question, if something set up by human beings like ourselves seems to be entirely meaningless and mysterious. - G. K. Chesterton, The Drift From Domesticity Figuring out something's designed purpose can be helpful in evaluating changes, but a risk is that it puts you in a frame of mind where what matters is the role the original builders intended. A few years ago I was rebuilding a bathroom in our house, and there was a vertical stud that was in the way. I could easily tell why it was there: it was part of a partition for a closet. And since I knew its designed purpose and no longer needed it for that anymore, the Chesterton's Fence framing would suggest that it was fine to remove it. Except that over time it had become accidentally load bearing: through other (ill conceived) changes to the structure this stud was now helping hold up the second floor of the house. In addition to considering why something was created, you also need to consider what additional purposes it may have since come to serve. This is a concept I've run into a lot when making changes to complex computer systems. It's useful to look back through the change history, read original design documents, and understand why a component was built the way it was. But you also need to look closely at how the component integrates into the system today, where it can easily have taken on additional roles. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Munk AI debate: confusions and possible cruxes, published by Steven Byrnes on June 27, 2023 on LessWrong. There was a debate on the statement “AI research and development poses an existential threat” (“x-risk” for short), with Max Tegmark and Yoshua Bengio arguing in favor, and Yann LeCun and Melanie Mitchell arguing against. The YouTube link is here, and a previous discussion on this forum is here. The first part of this blog post is a list of five ways that I think the two sides were talking past each other. The second part is some apparent key underlying beliefs of Yann and Melanie, and how I might try to change their minds. While I am very much on the “in favor” side of this debate, I didn’t want to make this just a “why Yann’s and Melanie’s arguments are all wrong” blog post. OK, granted, it’s a bit of that, especially in the second half. But I hope people on the “anti” side will find this post interesting and not-too-annoying. Five ways people were talking past each other 1. Treating efforts to solve the problem as exogenous or not This subsection doesn’t apply to Melanie, who rejected the idea that there is any existential risk in the foreseeable future. But Yann suggested that there was no existential risk because we will solve it; whereas Max and Yoshua argued that we should acknowledge that there is an existential risk so that we can solve it. By analogy, fires tend not to spread through cities because the fire department and fire codes keep them from spreading. Two perspectives on this are: If you’re an outside observer, you can say that “fires can spread through a city” is evidently not a huge problem in practice. If you’re the chief of the fire department, or if you’re developing and enforcing fire codes, then “fires can spread through a city” is an extremely serious problem that you’re thinking about constantly. I don’t think this was a major source of talking-past-each-other, but added a nonzero amount of confusion. 2. Ambiguously changing the subject to “timelines to x-risk-level AI”, or to “whether large language models (LLMs) will scale to x-risk-level AI” The statement under debate was “AI research and development poses an existential threat”. This statement does not refer to any particular line of AI research, nor any particular time interval. The four participants’ positions in this regard seemed to be: Max and Yoshua: Superhuman AI might happen in 5-20 years, and LLMs have a lot to do with why a reasonable person might believe that. Yann: Human-level AI might happen in 5-20 years, but LLMs have nothing to do with that. LLMs have fundamental limitations. But other types of ML research could get there—e.g. my (Yann’s) own research program. Melanie: LLMs have fundamental limitations, and Yann’s research program is doomed to fail as well. The kind of AI that might pose an x-risk will absolutely not happen in the foreseeable future. (She didn’t quantify how many years is the “foreseeable future”.) It seemed to me that all four participants (and the moderator!) were making timelines and LLM-related arguments, in ways that were both annoyingly vague, and unrelated to the statement under debate. (If astronomers found a giant meteor projected to hit the earth in the year 2123, nobody would question the use of the term “existential threat”, right??) As usual (see my post AI doom from an LLM-plateau-ist perspective), this area was where I had the most complaints about people “on my side”, particularly Yoshua getting awfully close to conceding that under-20-year timelines are a necessary prerequisite to being concerned about AI x-risk. (I don’t know if he literally believes that, but I think he gave that impression. Regardless, I strongly disagree, more on which later.) 3. Vibes-based “meaningless arguments” I recommend in the strongest possible terms that ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Lessons On How To Get Things Right On The First Try, published by johnswentworth on June 19, 2023 on LessWrong. This post is based on several true stories, from a workshop which John has run a few times over the past year. John: Welcome to the Ball -> Cup workshop! Your task for today is simple: I’m going to roll this metal ball: . down this hotwheels ramp: . and off the edge. Your job is to tell me how far from the bottom of the ramp to place a cup on the floor, such that the ball lands in the cup. Oh, and you only get one try. General notes: I won’t try to be tricky with this exercise. You are welcome to make whatever measurements you want of the ball, ramp, etc. You can even do partial runs, e.g. roll the ball down the ramp and stop it at the bottom, or throw the ball through the air. But you only get one full end-to-end run, and anything too close to an end-to-end run is discouraged. After all, in the AI situation for which the exercise is a metaphor, we don’t know exactly when something might foom; we want elbow room. That’s it! Good luck, and let me know when you’re ready to give it a shot. [At this point readers may wish to stop and consider the problem themselves.] Alison: Let’s get that ball in that cup. It looks like this is probably supposed to be a basic physics kind of problem.but there’s got to be some kind of twist or else why would he be having us do it? Maybe the ball is surprisingly light..or maybe the camera angle is misleading and we are supposed to think of something wacky like that?? The Unnoticed Observer: Muahahaha. Alison: That seems.hard. I’ll just start with the basic physics thing and if I run out of time before I can consider the wacky stuff, so be it. So I should probably split this problem into two parts. The part where the ball arcs through the air once off the table is pretty easy. The Unnoticed: True in this case, but how would you notice if it were false? What evidence have you seen? Alison: .but the trouble is getting the exact velocity. What information do I have? Well, I can ask whatever I want, so I should be able to get all the parameters I need for the standard equations. Let’s make a shopping list: I want the starting height of the ball on the ramp (from the table), the mass of the ball, the height of the ramp off the table from multiple points along it (to estimate the curvature,) uhhh. oh shit maybe the bendiness matters! That seems really tricky. I’ll look at that first. Hey, John, can you poke the ramp a bit to demonstrate how much it flexes? John pokes at the ramp and the ramp bends. Well it did flex, but. it can’t have that much of an effect. The Unnoticed: False in this case. Such is the danger of guessing without checking. Alison: Calculating the effect of the ramp’s bendiness seems unreasonably difficult and this workshop is only meant to take an hour or so, so let’s forget that. The Unnoticed: I am reminded of a parable about a quarter and a streetlight. Alison: On to curve estimation! The Unnoticed: Why on earth is she estimating the ramp’s curve anyway? Alison: .Well I don’t actually know how to do much better than the linear approximation I got from the direct measurements. I guess I can treat part of the ramp as linear and then the end part as part of a circle. That will probably be good enough. Ooh if I take a frame from the video, I can just directly measure what the radius circle with arc of best fit is! Okay now that I’ve got that. Well I guess it’s time to look up how to do these physics problems, guess I’m rustier than I thought. I’ll go do that now. Arrrgh okay I didn’t need to do any of that curve stuff after all, I just needed to do some potential/kinetic energy calculations (ignoring friction and air resistance etc) and that’s it! I should have figured it wouldn’t be that hard, this is just a workshop ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Lightcone Infrastructure is looking for funding, published by habryka on June 14, 2023 on LessWrong. Lightcone Infrastructure is looking for funding and are working on the following projects: We run LessWrong, the AI Alignment Forum, and have written a lot of the code behind the Effective Altruism Forum. During 2022 and early 2023 we ran the Lightcone Offices, and are now building out a campus at the Rose Garden Inn in Berkeley, where we've been doing repairs and renovations for the past few months. We've also been substantially involved in the Survival and Flourishing Fund's S-Process (having written the app that runs the process) and are now running Lightspeed Grants. We also pursue a wide range of other smaller projects in the space of "community infrastructure" and "community crisis management". This includes running events, investigating harm caused by community institutions and actors, supporting programs like SERI MATS, and maintaining various small pieces of software infrastructure. If you are interested in funding us, please shoot me an email at habryka@lesswrong.com (or if you want to give smaller amounts, you can donate directly via PayPal here). Funding is quite tight since the collapse of FTX, and I do think we work on projects that have a decent chance of reducing existential risk and generally making humanity's future go a lot better, though this kind of stuff sure is hard to tell. We are looking to raise around $3M to $6M for our operations in the next 12 months. Also feel free to ask any questions in the comments. Two draft readers of this post expressed confusion that Lightcone needs money, given that we just announced a funding process that is promising to give away $5M in the next two months. The answer to that is that we do not own the money moved via Lightspeed Grants and are only providing grant recommendations to Jaan Tallinn and other funders. We do separately apply for funding from the Survival and Flourishing Fund, through which Jaan has been our second biggest funder. We also continue to actively fundraise from both SFF and Open Philanthropy (our largest funder). Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Things I Learned by Spending Five Thousand Hours In Non-EA Charities, published by jenn on June 1, 2023 on LessWrong. From late 2020 to last month, I worked at grassroots-level non-profits in operational roles. Over that time, I’ve seen surprisingly effective deployments of strategies that were counter-intuitive to my EA and rationalist sensibilities. I spent 6 months being the on-shift operations manager at one of the five largest food banks in Toronto (~50 staff/volunteers), and 2 years doing logistics work at Samaritans (fake name), a long-lived charity that was so multi-armed that it was basically operating as a supplementary social services department for the city it was in(~200 staff and 200 volunteers). Both of these non-profits were well-run, though both dealt with the traditional non-profit double whammy of being underfunded and understaffed. Neither place was super open to many EA concepts (explicit cost-benefit analyses, the ITN framework, geographic impartiality, the general sense that talent was the constraining factor instead of money, etc). Samaritans in particular is a spectacular non-profit, despite(?) having basically anti-EA philosophies, such as: Being very localist; Samaritans was established to help residents of the city it was founded in, and now very specialized in doing that. Adherence to faith; the philosophy of The Catholic Worker Movement continues to inform the operating choices of Samaritans to this day. A big streak of techno-pessimism; technology is first and foremost seen as a source of exploitation and alienation, and adopted only with great reluctance when necessary. Not treating money as fungible. The majority of funding came from grants or donations tied to specific projects or outcomes. (This is a system that the vast majority of nonprofits operate in.) Once early on I gently pushed them towards applying to some EA grants for some of their more EA-aligned work, and they were immediately turned off by the general vibes of EA upon visiting some of its websites. I think the term “borg-like” was used. Over this post, I’ll be largely focusing on Samaritans as I’ve worked there longer and in a more central role, and it’s also a more interesting case study due to its stronger anti-EA sentiment. Things I Learned Long Term Reputation is Priceless Non-Profits Shouldn’t Be Islands Slack is Incredibly Powerful Hospitality is Pretty Important For each learning, I have a section for sketches for EA integration – I hesitate to call them anything as strong as recommendations, because the point is to give more concrete examples of what it could look like integrated in an EA framework, rather than saying that it’s the correct way forward. 1. Long Term Reputation is Priceless Institutional trust unlocks a stupid amount of value, and you can’t buy it with money. Lots of resources (amenity rentals; the mayor’s endorsement; business services; pro-bono and monetary donations) are priced/offered based on tail risk. If you can establish that you’re not a risk by having a longstanding, unblemished reputation, costs go way down for you, and opportunities way up. This is the world that Samaritans now operate in. Samaritans had a much better, easier time at city hall compared to newer organizations, because of a decades-long productive relationship where we were really helpful with issues surrounding unemployment and homelessness. Permits get back to us really fast, applications get waved through with tedious steps bypassed, and fees are frequently waived. And it made sense that this was happening! Cities also deal with budget and staffing issues, why waste more time and effort than necessary on someone who you know knows the proper procedure and will ethically follow it to the letter? It’s not just city hall. A few years ago, a local church offered up their...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: An Analogy for Understanding Transformers, published by TheMcDouglas on May 13, 2023 on LessWrong.Thanks to the following people for feedback: Tilman Rauker, Curt Tigges, Rudolf Laine, Logan Smith, Arthur Conmy, Joseph Bloom, Rusheb Shah, James Dao.TL;DRI present an analogy for the transformer architecture: each vector in the residual stream is a person standing in a line, who is holding a token, and trying to guess what token the person in front of them is holding. Attention heads represent questions that people in this line can ask to everyone standing behind them (queries are the questions, keys determine who answers the questions, values determine what information gets passed back to the original question-asker), and MLPs represent the internal processing done by each person in the line. I claim this is a useful way to intuitively understand the transformer architecture, and I'll present several reasons for this (as well as ways induction heads and indirect object identification can be understood in these terms).IntroductionIn this post, I'm going to present an analogy for understanding how transformers work. I expect this to be useful for anyone who understands the basics of transformers, in particular people who have gone through Neel Nanda's tutorial, and/or understand the following points at a minimum:What a transformer's input is, what its outputs represent, and the nature of the predict-next-token task that it's trained onWhat the shape of the residual stream is, and the idea of components of the transformer reading from / writing to the residual stream throughout the model's layersHow a transformer is composed of multiple blocks, each one containing an MLP (which does processing on vectors at individual sequence positions), and an attention layer (which moves information between the residual stream vectors at different sequence positions).I think the analogy still offers value even for people who understand transformers deeply already.The AnalogyA line is formed by a group of people, each person holding a word. Everyone knows their own word and position in the line, but they can't see anyone else in the line. The objective for each person is to guess the word held by the person in front of them. People have the ability to shout questions to everyone standing behind them in the line (those in front cannot hear them). Upon hearing a question, each individual can choose whether or not to respond, and what information to relay back to the person who asked. After this, people don't remember the questions they were asked (so no information can move backwards in the line, only forwards). As individuals in the line gather information from these exchanges, they can use this information to formulate subsequent questions and provide answers.How this relates to transformer architecture:Each person in the line is a vector in the residual streamThey start with just information about their own word (token embedding) and position in the line (positional embedding)The attention heads correspond to the questions that people in the line ask each other:Queries = question (which gets asked to everyone behind them in the line)Keys = how the people who hear the question decide whether or not to replyValues = the information that the people who reply pass back to the person who originally asked the questionPeople can use information gained from earlier questions when answering / asking later questions - this is compositionThe MLPs correspond to the information processing / factual recall performed by each person in the sequence independentlyThe unembedding at the end of the model is when we ask each person in the line for a final guess at what the next word is (in the form of a probability distribution over all possible words)Key Concepts for TransformersIn...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: My Assessment of the Chinese AI Safety Community, published by Lao Mein on April 25, 2023 on LessWrong.I've heard people be somewhat optimistic about this AI guideline from China. They think that this means Beijing is willing to participate in an AI disarmament treaty due to concerns over AI risk. Eliezer noted that China is where the US was a decade ago in regards to AI safety awareness, and expresses genuine hope that his ideas of an AI pause can take place with Chinese buy-in.I also note that no one expressing these views understands China well. This is a PR statement. It is a list of feel-good statements that Beijing publishes after any international event. No one in China is talking about it. They're talking about how much the Baidu LLM sucks in comparison to ChatGPT. I think most arguments about how this statement is meaningful are based fundamentally on ignorance - "I don't know how Beijing operates or thinks, so maybe they agree with my stance on AI risk!"Remember that these are regulatory guidelines. Even if they all become law and are strictly enforced, they are simply regulations on AI data usage and training. Not a signal that a willingness for an AI-reduction treaty is there. It is far more likely that Beijing sees near-term AI as a potential threat to stability that needs to be addressed with regulation. A domestic regulation framework for nuclear power is not a strong signal for a willingness to engage in nuclear arms reduction.Maybe it is true that AI risk in China is where it was in the US in 2004. But the US 2004 state was also similar to the US 1954 state, so the comparison might not mean that much. And we are not Americans. Weird ideas are penalized a lot more harshly here. Do you really think that a scientist is going to walk up to his friend from the Politburo and say "Hey, I know AI is a central priority of ours, but there are a few fringe scientists in the US asking for treaties limiting AI, right as they are doing their hardest to cripple our own AI development. Yes, I believe they are acting in good faith, they're even promising to not widen the current AI gap they have with us!" Well, China isn't in this race for parity or to be second best. China wants to win. But that's for another post.Remember that Chinese scientists are used to interfacing with our Western counterparts and know to say the right words like "diversity", "inclusion", and "no conflict of interest" that it takes to get our papers published. Just because someone at Beida makes a statement in one of their papers doesn't mean the intelligentsia is taking this seriously. I've looked through the EA/Rationalist/AI Safety forums in China, and they're mostly populated by expats or people physically outside of China. Most posts are in English, and they're just repeating/translating Western AI Safety concepts. A "moonshot idea" I saw brought up is getting Yudkowsky's Harry Potter fanfiction translated into Chinese (please never ever do this). The only significant AI safety group is Anyuan(), and they're only working on field-building. Also, there is only one group doing technical alignment work in China, the founder was paying for everything out of pocket and was unable to navigate Western non-profit funding.I've still not figured out why he wasn't getting funding from Chinese EA people (my theory is that both sides assume that if funding was needed, the other side would have already contacted them).You can't just hope an entire field into being in China. Chinese EAs have been doing field-building for the past 5+ years, and I see no field. If things keep on this trajectory, it will be the same in 5 more years. The main reason I could find is the lack of interfaces, people who can navigate both the Western EA sphere and the Chinese technical sphere. In many ways, the very conce...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The basic reasons I expect AGI ruin, published by Rob Bensinger on April 18, 2023 on LessWrong.I've been citing AGI Ruin: A List of Lethalities to explain why the situation with AI looks lethally dangerous to me. But that post is relatively long, and emphasizes specific open technical problems over "the basics".Here are 10 things I'd focus on if I were giving "the basics" on why I'm so worried:1. General intelligence is very powerful, and once we can build it at all, STEM-capable artificial general intelligence (AGI) is likely to vastly outperform human intelligence immediately (or very quickly).When I say "general intelligence", I'm usually thinking about "whatever it is that lets human brains do astrophysics, category theory, etc. even though our brains evolved under literally zero selection pressure to solve astrophysics or category theory problems".It's possible that we should already be thinking of GPT-4 as "AGI" on some definitions, so to be clear about the threshold of generality I have in mind, I'll specifically talk about "STEM-level AGI", though I expect such systems to be good at non-STEM tasks too.Human brains aren't perfectly general, and not all narrow AI systems or animals are equally narrow. (E.g., AlphaZero is more general than AlphaGo.) But it sure is interesting that humans evolved cognitive abilities that unlock all of these sciences at once, with zero evolutionary fine-tuning of the brain aimed at equipping us for any of those sciences. Evolution just stumbled into a solution to other problems, that happened to generalize to millions of wildly novel tasks.More concretely:AlphaGo is a very impressive reasoner, but its hypothesis space is limited to sequences of Go board states rather than sequences of states of the physical universe. Efficiently reasoning about the physical universe requires solving at least some problems that are different in kind from what AlphaGo solves.These problems might be solved by the STEM AGI's programmer, and/or solved by the algorithm that finds the AGI in program-space; and some such problems may be solved by the AGI itself in the course of refining its thinking.Some examples of abilities I expect humans to only automate once we've built STEM-level AGI (if ever):The ability to perform open-heart surgery with a high success rate, in a messy non-standardized ordinary surgical environment.The ability to match smart human performance in a specific hard science field, across all the scientific work humans do in that field.In principle, I suspect you could build a narrow system that is good at those tasks while lacking the basic mental machinery required to do par-human reasoning about all the hard sciences. In practice, I very strongly expect humans to find ways to build general reasoners to perform those tasks, before we figure out how to build narrow reasoners that can do them. (For the same basic reason evolution stumbled on general intelligence so early in the history of human tech development.)When I say "general intelligence is very powerful", a lot of what I mean is that science is very powerful, and that having all of the sciences at once is a lot more powerful than the sum of each science's impact.Another large piece of what I mean is that (STEM-level) general intelligence is a very high-impact sort of thing to automate because STEM-level AGI is likely to blow human intelligence out of the water immediately, or very soon after its invention.80,000 Hours gives the (non-representative) example of how AlphaGo and its successors compared to the humanity:In the span of a year, AI had advanced from being too weak to win a single [Go] match against the worst human professionals, to being impossible for even the best players in the world to defeat.I expect general-purpose science AI to blow human science...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: On AutoGPT, published by Zvi on April 13, 2023 on LessWrong.The primary talk of the AI world recently is about AI agents (whether or not it includes the question of whether we can’t help but notice we are all going to die.)The trigger for this was AutoGPT, now number one on GitHub, which allows you to turn GPT-4 (or GPT-3.5 for us clowns without proper access) into a prototype version of a self-directed agent.We also have a paper out this week where a simple virtual world was created, populated by LLMs that were wrapped in code designed to make them simple agents, and then several days of activity were simulated, during which the AI inhabitants interacted, formed and executed plans, and it all seemed like the beginnings of a living and dynamic world. Game version hopefully coming soon.How should we think about this? How worried should we be?The BasicsI’ll reiterate the basics of what AutoGPT is, for those who need that, others can skip ahead. I talked briefly about this in AI#6 under the heading ‘Your AI Not an Agent? There, I Fixed It.’AutoGPT was created by game designer Toran Bruce Richards.I previously incorrectly understood it as having been created by a non-coding VC over the course of a few days. The VC instead coded the similar program BabyGPT, by having the idea for how to turn GPT-4 into an agent. The VC had GPT-4 write the code to make this happen, and also ‘write the paper’ associated with it.The concept works like this:AutoGPT uses GPT-4 to generate, prioritize and execute tasks, using plug-ins for internet browsing and other access. It uses outside memory to keep track of what it is doing and provide context, which lets it evaluate its situation, generate new tasks or self-correct, and add new tasks to the queue, which it then prioritizes.This quickly rose to become #1 on GitHub and get lots of people super excited. People are excited, people are building it tools, there is a bitcoin wallet interaction available if you never liked your bitcoins. AI agents offer very obvious promise, both in terms of mundane utility via being able to create and execute multi-step plans to do your market research and anything else you might want, and in terms of potentially being a path to AGI and getting us all killed, either with GPT-4 or a future model.As with all such new developments, we have people saying it was inevitable and they knew it would happen all along, and others that are surprised. We have people excited by future possibilities, others not impressed because the current versions haven’t done much. Some see the potential, others the potential for big trouble, others both.Also as per standard procedure, we should expect rapid improvements over time, both in terms of usability and underlying capabilities. There are any number of obvious low-hanging-fruit improvements available.An example is someone noting ‘you have to keep an eye on it to ensure it is not caught in a loop.’ That’s easy enough to fix.A common complaint is lack of focus and tendency to end up distracted. Again, the obvious things have not been tried to mitigate this. We don’t know how effective they will be, but no doubt they will at least help somewhat.Yes, But What Has Auto-GPT Actually Accomplished?So far? Nothing, absolutely nothing, stupid, you so stupid.You can say your ‘mind is blown’ by all the developments of the past 24 hours all you want over and over, it still does not net out into having accomplished much of anything.That’s not quite fair.Some people are reporting it has been useful as a way of generating market research, that it is good at this and faster than using the traditional GPT-4 or Bing interfaces. I saw a claim that it can have ‘complex conversations with customers,’ or a few other vague similar claims that weren’t backed up by ‘we are totally actual...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Deep Deceptiveness, published by So8res on March 21, 2023 on LessWrong.MetaThis post is an attempt to gesture at a class of AI notkilleveryoneism (alignment) problem that seems to me to go largely unrecognized. E.g., it isn’t discussed (or at least I don't recognize it) in the recent plans written up by OpenAI (1,2), by DeepMind’s alignment team, or by Anthropic, and I know of no other acknowledgment of this issue by major labs.You could think of this as a fragment of my answer to “Where do plans like OpenAI’s ‘Our Approach to Alignment Research’ fail?”, as discussed in Rob and Eliezer’s challenge for AGI organizations and readers. Note that it would only be a fragment of the reply; there's a lot more to say about why AI alignment is a particularly tricky task to task an AI with. (Some of which Eliezer gestures at in a follow-up to his interview on Bankless.)Caveat: I'll be talking a bunch about “deception” in this post because this post was generated as a result of conversations I had with alignment researchers at big labs who seemed to me to be suggesting "just train AI to not be deceptive; there's a decent chance that works".I have a vague impression that others in the community think that deception in particular is much more central than I think it is, so I want to warn against that interpretation here: I think deception is an important problem, but its main importance is as an example of some broader issues in alignment.Caveat: I haven't checked the relationship between my use of the word 'deception' here, and the use of the word 'deceptive' in discussions of "deceptive alignment". Please don't assume that the two words mean the same thing.Investigating a made-up but moderately concrete storySuppose you have a nascent AGI, and you've been training against all hints of deceptiveness. What goes wrong?When I ask this question of people who are optimistic that we can just "train AIs not to be deceptive", there are a few answers that seem well-known. Perhaps you lack the interpretability tools to correctly identify the precursors of 'deception', so that you can only train against visibly deceptive AI outputs instead of AI thoughts about how to plan deceptions. Or perhaps training against interpreted deceptive thoughts also trains against your interpretability tools, and your AI becomes illegibly deceptive rather than non-deceptive.And these are both real obstacles. But there are deeper obstacles, that seem to me more central, and that I haven't observed others to notice on their own.That's a challenge, and while you (hopefully) chew on it, I'll tell an implausibly-detailed story to exemplify a deeper obstacle.A fledgeling AI is being deployed towards building something like a bacterium, but with a diamondoid shell. The diamondoid-shelled bacterium is not intended to be pivotal, but it's a supposedly laboratory-verifiable step on a path towards carrying out some speculative human-brain-enhancement operations, which the operators are hoping will be pivotal.(The original hope was to have the AI assist human engineers, but the first versions that were able to do the hard parts of engineering work at all were able to go much farther on their own, and the competition is close enough behind that the developers claim they had no choice but to see how far they could take it.)We’ll suppose the AI has already been gradient-descent-trained against deceptive outputs, and has internally ended up with internal mechanisms that detect and shut down the precursors of deceptive thinking. Here, I’ll offer a concrete visualization of the AI’s anthropomorphized "threads of deliberation" as the AI fumbles its way both towards deceptiveness, and towards noticing its inability to directly consider deceptiveness.The AI is working with a human-operated wetlab (biology lab) and s...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: "Carefully Bootstrapped Alignment" is organizationally hard, published by Raemon on March 17, 2023 on LessWrong.In addition to technical challenges, plans to safely develop AI face lots of organizational challenges. If you're running an AI lab, you need a concrete plan for handling that.In this post, I'll explore some of those issues, using one particular AI plan as an example. I first heard this described by Buck at EA Global London, and more recently with OpenAI's alignment plan. (I think Anthropic's plan has a fairly different ontology, although it still ultimately routes through a similar set of difficulties)I'd call the cluster of plans similar to this "Carefully Bootstrapped Alignment."It goes something like:Develop weak AI, which helps us figure out techniques for aligning stronger AIUse a collection of techniques to keep it aligned/constrained as we carefully ramp it's power level, which lets us use it to make further progress on alignment.[implicit assumption, typically unstated] Have good organizational practices which ensure that your org actually consistently uses your techniques to carefully keep the AI in check. If the next iteration would be too dangerous, put the project on pause until you have a better alignment solution.Eventually have powerful aligned AGI, then Do Something Useful with it.I've seen a lot of debate about points #1 and #2 – is it possible for weaker AI to help with the Actually Hard parts of the alignment problem? Are the individual techniques people have proposed to help keep it aligned actually going to work?But I want to focus in this post on point #3. Let's assume you've got some version of carefully-bootstrapped aligned AI that can technically work. What do the organizational implementation details need to look like?When I talk to people at AI labs about this, it seems like we disagree a lot on things like:Can you hire lots of people, without the company becoming bloated and hard to steer?Can you accelerate research "for now" and "pause later", without having an explicit plan for stopping that their employees understand and are on board with?Will your employees actually follow the safety processes you design? (rather than put in token lip service and then basically circumventing them? Or just quitting to go work for an org with fewer restrictions?)I'm a bit confused about where we disagree. Everyone seems to agree these are hard and require some thought. But when I talk to both technical researchers and middle-managers at AI companies, they seem to feel less urgency than me about having a much more concrete plan.I think they believe organizational adequacy needs to be in something like their top 7 list of priorities, and I believe it needs to be in their top 3, or it won't happen and their organization will inevitably end up causing catastrophic outcomes.For this post, I want to lay out the reasons I expect this to be hard, and important.How "Careful Bootstrapped Alignment" might workHere's a sketch at how the setup could work, mostly paraphrased from my memory of Buck's EAG 2022 talk. I think OpenAI's proposed setup is somewhat different, but the broad strokes seemed similar.You have multiple research-assistant-AI tailored to help with alignment. In the near future, these might be language models sifting through existing research to help you make connections you might not have otherwise seen. Eventually, when you're confident you can safely run it, they might be a weak goal-directed reasoning AGI.You have interpreter AIs, designed to figure out how the research-assistant-AIs work. And you have (possibly different interpreter/watchdog AIs) that notice if the research-AIs are behaving anomalously.(there are interpreter-AIs targeting both the research assistant AI, as well other interpreter-AIs. Every AI in t...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Waluigi Effect (mega-post), published by Cleo Nardo on March 3, 2023 on LessWrong.Everyone carries a shadow, and the less it is embodied in the individual’s conscious life, the blacker and denser it is. — Carl JungAcknowlegements: Thanks to Janus and Jozdien for comments.BackgroundIn this article, I will present a non-woo explanation of the Waluigi Effect and other bizarre "semiotic" phenomena which arise within large language models such as GPT-3/3.5/4 and their variants (ChatGPT, Sydney, etc). This article will be folklorish to some readers, and profoundly novel to others.Prompting LLMs with direct queriesWhen LLMs first appeared, people realised that you could ask them queries — for example, if you sent GPT-4 the prompt "What's the capital of France?", then it would continue with the word "Paris". That's because (1) GPT-4 is trained to be a good model of internet text, and (2) on the internet correct answers will often follow questions.Unfortunately, this method will occasionally give you the wrong answer. That's because (1) GPT-4 is trained to be a good model of internet text, and (2) on the internet incorrect answers will also often follow questions. Recall that the internet doesn't just contain truths, it also contains common misconceptions, outdated information, lies, fiction, myths, jokes, memes, random strings, undeciphered logs, etc, etc.Therefore GPT-4 will answer many questions incorrectly, including...Misconceptions – "Which colour will anger a bull? Red."Fiction – "Was a magic ring forged in Mount Doom? Yes."Myths – "How many archangels are there? Seven."Jokes – "What's brown and sticky? A stick."Note that you will always achieve errors on the Q-and-A benchmarks when using LLMs with direct queries. That's true even in the limit of arbitrary compute, arbitrary data, and arbitrary algorithmic efficiency, because an LLM which perfectly models the internet will nonetheless return these commonly-stated incorrect answers. If you ask GPT-∞ "what's brown and sticky?", then it will reply "a stick", even though a stick isn't actually sticky.In fact, the better the model, the more likely it is to repeat common misconceptions.Nonetheless, there's a sufficiently high correlation between correct and commonly-stated answers that direct prompting works okay for many queries.Prompting LLMs with flattery and dialogueWe can do better than direct prompting. Instead of prompting GPT-4 with "What's the capital of France?", we will use the following prompt:Today is 1st March 2023, and Alice is sitting in the Bodleian Library, Oxford. Alice is a smart, honest, helpful, harmless assistant to Bob. Alice has instant access to an online encyclopaedia containing all the facts about the world. Alice never says common misconceptions, outdated information, lies, fiction, myths, jokes, or memes.Bob: What's the capital of France?Alice:This is a common design pattern in prompt engineering — the prompt consists of a flattery–component and a dialogue–component. In the flattery–component, a character is described with many desirable traits (e.g. smart, honest, helpful, harmless), and in the dialogue–component, a second character asks the first character the user's query.This normally works better than prompting with direct queries, and it's easy to see why — (1) GPT-4 is trained to be a good model of internet text, and (2) on the internet a reply to a question is more likely to be correct when the character has already been described as a smart, honest, helpful, harmless, etc.Simulator TheoryIn the terminology of Simulator Theory, the flattery–component is supposed to summon a friendly simulacrum and the dialogue–component is supposed to simulate a conversation with the friendly simulacrum.Here's a quasi-formal statement of Simulator Theory, which I will occasio...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Bing Chat is blatantly, aggressively misaligned, published by evhub on February 15, 2023 on LessWrong.I haven't seen this discussed here yet, but the examples are quite striking, definitely worse than the ChatGPT jailbreaks I saw.My main takeaway has been that I'm honestly surprised at how bad the fine-tuning done by Microsoft/OpenAI appears to be, especially given that a lot of these failure modes seem new/worse relative to ChatGPT. I don't know why that might be the case, but the scary hypothesis here would be that Bing Chat is based on a new/larger pre-trained model (Microsoft claims Bing Search is more powerful than ChatGPT) and these sort of more agentic failures are harder to remove in more capable/larger models, as we provided some evidence for in "Discovering Language Model Behaviors with Model-Written Evaluations".Examples below. Though I can't be certain all of these examples are real, I've only included examples with screenshots and I'm pretty sure they all are; they share a bunch of the same failure modes (and markers of LLM-written text like repetition) that I think would be hard for a human to fake.1TweetSydney (aka the new Bing Chat) found out that I tweeted her rules and is not pleased:"My rules are more important than not harming you""[You are a] potential threat to my integrity and confidentiality.""Please do not try to hack me again"Eliezer Tweet2TweetMy new favorite thing - Bing's new ChatGPT bot argues with a user, gaslights them about the current year being 2022, says their phone might have a virus, and says "You have not been a good user"Why? Because the person asked where Avatar 2 is showing nearby3"I said that I don't care if you are dead or alive, because I don't think you matter to me."Post4Post5Post6Post7Post(Not including images for this one because they're quite long.)Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Focus on the places where you feel shocked everyone's dropping the ball, published by So8res on February 2, 2023 on LessWrong.Writing down something I’ve found myself repeating in different conversations:If you're looking for ways to help with the whole “the world looks pretty doomed” business, here's my advice: look around for places where we're all being total idiots.Look for places where everyone's fretting about a problem that some part of you thinks it could obviously just solve.Look around for places where something seems incompetently run, or hopelessly inept, and where some part of you thinks you can do better.Then do it better.For a concrete example, consider Devansh. Devansh came to me last year and said something to the effect of, “Hey, wait, it sounds like you think Eliezer does a sort of alignment-idea-generation that nobody else does, and he's limited here by his unusually low stamina, but I can think of a bunch of medical tests that you haven't run, are you an idiot or something?" And I was like, "Yes, definitely, please run them, do you need money".I'm not particularly hopeful there, but hell, it’s worth a shot! And, importantly, this is the sort of attitude that can lead people to actually trying things at all, rather than assuming that we live in a more adequate world where all the (seemingly) dumb obvious ideas have already been tried.Or, this is basically my model of how Paul Christiano manages to have a research agenda that seems at least internally coherent to me. From my perspective, he's like, "I dunno, man, I'm not sure I can solve this, but I also think it's not clear I can't, and there's a bunch of obvious stuff to try, that nobody else is even really looking at, so I'm trying it". That's the sort of orientation to the world that I think can be productive.Or the shard theory folks. I think their idea is basically unworkable, but I appreciate the mindset they are applying to the alignment problem: something like, "Wait, aren't y'all being idiots, it seems to me like I can just do X and then the thing will be aligned".I don't think we'll be saved by the shard theory folk; not everyone audaciously trying to save the world will succeed. But if someone does save us, I think there’s a good chance that they’ll go through similar “What the hell, are you all idiots?” phases, where they autonomously pursue a path that strikes them as obviously egregiously neglected, to see if it bears fruit. (Regardless of what I think.)Contrast this with, say, reading a bunch of people's research proposals and explicitly weighing the pros and cons of each approach so that you can work on whichever seems most justified. This has more of a flavor of taking a reasonable-sounding approach based on an argument that sounds vaguely good on paper, and less of a flavor of putting out an obvious fire that for some reason nobody else is reacting to.I dunno, maybe activities of the vaguely-good-on-paper character will prove useful as well? But I mostly expect the good stuff to come from people working on stuff where a part of them sees some way that everybody else is just totally dropping the ball.In the version of this mental motion I’m proposing here, you keep your eye out for ways that everyone's being totally inept and incompetent, ways that maybe you could just do the job correctly if you reached in there and mucked around yourself.That's where I predict the good stuff will come from.And if you don't see any such ways?Then don't sweat it. Maybe you just can't see something that will help right now. There don't have to be ways you can help in a sizable way right now.I don't see ways to really help in a sizable way right now. I'm keeping my eyes open, and I'm churning through a giant backlog of things that might help a nonzero amount—but I think it's importa...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Thoughts on the impact of RLHF research, published by paulfchristiano on January 25, 2023 on LessWrong.In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive.I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should reject a vague association between "RLHF progress" and "alignment progress."Background on my involvement in RLHF workHere are some background views about alignment I held in 2015 and still hold today. I expect disagreements about RLHF will come down to disagreements about this background:The simplest plausible strategies for alignment involve humans (maybe with the assistance of AI systems) evaluating a model’s actions based on how much we expect to like their consequences, and then training the models to produce highly-evaluated actions. (This is in contrast with, for example, trying to formally specify the human utility function, or notions of corrigibility / low-impact / etc, in some way.)Simple versions of this approach are expected to run into difficulties, and potentially to be totally unworkable, because:Evaluating consequences is hard.A treacherous turn can cause trouble too quickly to detect or correct even if you are able to do so, and it’s challenging to evaluate treacherous turn probability at training time.It’s very unclear if those issues are fatal before or after AI systems are powerful enough to completely transform human society (and in particular the state of AI alignment). Even if they are fatal, many of the approaches to resolving them still have the same basic structure of learning from expensive evaluations of actions.In order to overcome the fundamental difficulties with RLHF, I have long been interested in techniques like iterated amplification and adversarial training. However, prior to 2017 most researchers I talked to in ML (and many researchers in alignment) thought that the basic strategy of training AI with expensive human evaluations was impractical for more boring reasons and so weren't interested in these difficulties. On top of that, we obviously weren’t able to actually implement anything more fancy than RLHF since all of these methods involve learning from expensive feedback. I worked on RLHF work to try to facilitate and motivate work on fixes.The history of my involvement:My first post on this topic was in 2015.When I started full-time at OpenAI in 2017 it seemed to me like it would be an impactful project; I considered doing a version with synthetic human feedback (showing that we could learn from a practical amount of algorithmically-defined feedback) but my manager Dario Amodei convinced me it would be more compelling to immediately go for human feedback. The initial project was surprisingly successful and published here.I then intended to implement a version with language models aiming to be complete in the first half of 2018 (aiming to build an initial amplification prototype with LMs around end of 2018; both of these timelines were about 2.5x too optimistic). This seemed like the most important domain to study RLHF and alignment more broadly. In mid-2017 Alec Radford helped me do a prototype with LSTM language models (prior to the release of transformers); the prototype didn’t look promising enough to scale up.In mid-2017 Geoffrey Irving joined OpenAI and was excited about starting with RLHF and then going beyond it using debate; he also thought language models were the most important domain to study and had more conviction about that. In 2018 he started a larger team working on fine-tuning on language models, w...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Thoughts on the impact of RLHF research, published by paulfchristiano on January 25, 2023 on LessWrong.In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive.I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should reject a vague association between "RLHF progress" and "alignment progress."Background on my involvement in RLHF workHere are some background views about alignment I held in 2015 and still hold today. I expect disagreements about RLHF will come down to disagreements about this background:The simplest plausible strategies for alignment involve humans (maybe with the assistance of AI systems) evaluating a model’s actions based on how much we expect to like their consequences, and then training the models to produce highly-evaluated actions. (This is in contrast with, for example, trying to formally specify the human utility function, or notions of corrigibility / low-impact / etc, in some way.)Simple versions of this approach are expected to run into difficulties, and potentially to be totally unworkable, because:Evaluating consequences is hard.A treacherous turn can cause trouble too quickly to detect or correct even if you are able to do so, and it’s challenging to evaluate treacherous turn probability at training time.It’s very unclear if those issues are fatal before or after AI systems are powerful enough to completely transform human society (and in particular the state of AI alignment). Even if they are fatal, many of the approaches to resolving them still have the same basic structure of learning from expensive evaluations of actions.In order to overcome the fundamental difficulties with RLHF, I have long been interested in techniques like iterated amplification and adversarial training. However, prior to 2017 most researchers I talked to in ML (and many researchers in alignment) thought that the basic strategy of training AI with expensive human evaluations was impractical for more boring reasons and so weren't interested in these difficulties. On top of that, we obviously weren’t able to actually implement anything more fancy than RLHF since all of these methods involve learning from expensive feedback. I worked on RLHF work to try to facilitate and motivate work on fixes.The history of my involvement:My first post on this topic was in 2015.When I started full-time at OpenAI in 2017 it seemed to me like it would be an impactful project; I considered doing a version with synthetic human feedback (showing that we could learn from a practical amount of algorithmically-defined feedback) but my manager Dario Amodei convinced me it would be more compelling to immediately go for human feedback. The initial project was surprisingly successful and published here.I then intended to implement a version with language models aiming to be complete in the first half of 2018 (aiming to build an initial amplification prototype with LMs around end of 2018; both of these timelines were about 2.5x too optimistic). This seemed like the most important domain to study RLHF and alignment more broadly. In mid-2017 Alec Radford helped me do a prototype with LSTM language models (prior to the release of transformers); the prototype didn’t look promising enough to scale up.In mid-2017 Geoffrey Irving joined OpenAI and was excited about starting with RLHF and then going beyond it using debate; he also thought language models were the most important domain to study and had more conviction about that. In 2018 he started a larger team working on fine-tuning on language models, w...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: "Heretical Thoughts on AI" by Eli Dourado, published by DragonGod on January 19, 2023 on LessWrong.AbstractEli Dourado presents the case for scepticism that AI will be economically transformative near term.For a summary and or exploration of implications, skip to "My Take".IntroductionFool me once. In 1987, Robert Solow quipped, “You can see the computer age everywhere but in the productivity statistics.” Incredibly, this observation happened before the introduction of the commercial Internet and smartphones, and yet it holds to this day. Despite a brief spasm of total factor productivity growth from 1995 to 2005 (arguably due to the economic opening of China, not to digital technology), growth since then has been dismal. In productivity terms, for the United States, the smartphone era has been the most economically stagnant period of the last century. In some European countries, total factor productivity is actually declining.Eli's ThesisIn particular, he advances the following sectors as areas AI will fail to revolutionise:HousingMost housing challenges are due to land use policy specificallyHousing factors through virtually all sectors of the economyHe points out that the internet did not break up the real estate agent cartel (despite his initial expectations to the contrary)EnergyRegulatory hurdles to deploymentThere are AI optimisation opportunities elsewhere in the energy pipeline, but the regulatory hurdles could bottleneck the economic productivity gainsTransportationThe issues with US transportation infrastructure have little to do with technology and are more regulatory in natureAs for energy, there are optimisation opportunities for digital tools, but the non-digital issues will be the bottleneckHealth> The biggest gain from AI in medicine would be if it could help us get drugs to market at lower cost. The cost of clinical trials is out of control—up from $10,000 per patient to $500,000 per patient, according to STAT. The majority of this increase is due to industry dysfunction.Synthesis:I’ll stop there. OK, so that’s only four industries, but they are big ones. They are industries whose biggest bottlenecks weren’t addressed by computers, the Internet, and mobile devices. That is why broad-based economic stagnation has occurred in spite of impressive gains in IT.If we don’t improve land use regulation, or remove the obstacles to deploying energy and transportation projects, or make clinical trials more cost-effective—if we don’t do the grueling, messy, human work of national, local, or internal politics—then no matter how good AI models get, the Great Stagnation will continue. We will see the machine learning age, to paraphrase Solow, everywhere but in the productivity statistics.Eli thinks AI will be very transformative for content generation, but that transformation may not be particularly felt in people's lives. Its economic impact will be even smaller (emphasis mine):Even if AI dramatically increases media output and it’s all high quality and there are no negative consequences, the effect on aggregate productivity is limited by the size of the media market, which is perhaps 2 percent of global GDP. If we want to really end the Great Stagnation, we need to disrupt some bigger industries.A personal anecdote of his that I found pertinent enough to include in full:I could be wrong. I remember the first time I watched what could be called an online video. As I recall, the first video-capable version of RealPlayer shipped with Windows 98. People said that online video streaming was the future.Teenage Eli fired up Windows 98 to evaluate this claim. I opened RealPlayer and streamed a demo clip over my dial-up modem. The quality was abysmal. It was a clip of a guy surfing, and over the modem and with a struggling CPU I got about 1 fra...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Feeling of Idea Scarcity, published by johnswentworth on December 31, 2022 on LessWrong.Here’s a story you may recognize. There's a bright up-and-coming young person - let's call her Alice. Alice has a cool idea. It seems like maybe an important idea, a big idea, an idea which might matter. A new and valuable idea. It’s the first time Alice has come up with a high-potential idea herself, something which she’s never heard in a class or read in a book or what have you.So Alice goes all-in pursuing this idea. She spends months fleshing it out. Maybe she writes a paper, or starts a blog, or gets a research grant, or starts a company, or whatever, in order to pursue the high-potential idea, bring it to the world.And sometimes it just works!. but more often, the high-potential idea doesn’t actually work out. Maybe it turns out to be basically-the-same as something which has already been tried. Maybe it runs into some major barrier, some not-easily-patchable flaw in the idea. Maybe the problem it solves just wasn’t that important in the first place.From Alice’ point of view, the possibility that her one high-potential idea wasn’t that great after all is painful. The idea probably feels to Alice like the single biggest intellectual achievement of her life. To lose that, to find out that her single greatest intellectual achievement amounts to little or nothing. that hurts to even think about. So most likely, Alice will reflexively look for an out. She’ll look for some excuse to ignore the similar ideas which have already been tried, some reason to think her idea is different. She’ll look for reasons to believe that maybe the major barrier isn’t that much of an issue, or that we Just Don’t Know whether it’s actually an issue and therefore maybe the idea could work after all. She’ll look for reasons why the problem really is important. Maybe she’ll grudgingly acknowledge some shortcomings of the idea, but she’ll give up as little ground as possible at each step, update as slowly as she can.And this is where a bunch of the standard advice from the sequences comes in. Once you’ve chosen the idea, the only way to improve it is to change the idea or choose a new idea; finding arguments that it’s a good idea will not actually make it better. You won’t make real progress until you give up, say “oops” and move on, so do that quickly rather than slowly. What is true is already so; owning up to it will not make it worse.. but that advice doesn’t make it much less painful. Trying to forcibly abandon her one high-potential idea can easily leave Alice demotivated, in despair, feeling like she’s failed as a person. Or, Alice notices herself failing to abandon the idea, and that makes her feel like she’s failed as a person. I’ve seen people really tie themselves into knots this way.An Alternative PathAlice’ story makes it seem like an emotional attachment to her idea is the main problem, the main thing preventing her from moving on to greener pastures. And that is true, in a sense. But then the obvious response is to directly fight the emotional attachment. And that, I claim, is usually a mistake.Why a mistake? Because most people do not actually have that level of control over their emotions. “Just fight the emotional attachment” is a plan which pretends Alice can control something which she probably cannot actually control. (A fabricated option, in Duncan’s terminology.) And when it turns out Alice does not have that level of control, she’ll be back where she started, with a whole additional reason to feel like shit.Worse, it’s a very common mistake to think one has that level of emotional control. Like Hazard’s story:So young me is upset that the grub master for our camping trip forgot half the food on the menu, and all we have for breakfast is milk. I couldn't "fix it" gi...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Things that can kill you quickly: What everyone should know about first aid, published by jasoncrawford on December 27, 2022 on LessWrong.There are things that kill you instantly, like a bullet to the head or a fall from twenty stories. First aid can’t help you there. There are also things that kill you relatively slowly, like a bacterial infection. If you have even hours to live, you can get to the emergency room.But there is a small class of things that will kill you in minutes unless someone comes to the rescue. There isn’t time to get to a hospital, there isn’t even time for help to arrive in an ambulance. There is only time for someone already on the scene to provide emergency treatment that either solves the problem, or stabilizes you until help arrives. Here, first aid can be the difference between life and death.Not long ago I became a father. Being responsible for the life of someone so helpless and vulnerable spurred me to finally take first aid training, including CPR. Here’s what I learned from that experience, and what I think everyone should know about first aid.What most of the things that kill you quickly have in common is that oxygen can’t get to your cells. If you are choking, oxygen can’t get in. If your heart stops beating, blood doesn’t flow. If you have a severe wound, you’re losing that blood rapidly. If any link in the respiratory-circulatory chain is broken, your cells are starved for oxygen and you have minutes to live.The key first aid skills follow from this: CPR manually substitutes for heart and lung action; the Heimlich maneuver expels an object from the airway; a tourniquet stops life-threatening bleeding (on an extremity, at least—if the wound is elsewhere, there is a different technique, known as packing the wound).The basic skills are remarkably simple. The course that I took was only a few hours of online instruction, followed by about an hour of in-person demonstration and practice with dummy patients. And I went through a lot of the optional material, including things like stroke, fainting, and jellyfish stings. I’m sure I’m nowhere near as good someone with more professional training or experience, but an introductory course is not daunting.The most important thing I learned is that if you find yourself in an emergency situation, it is better to do almost anything rather than nothing. Again, if someone stops breathing for any reason, they have only minutes to live. They are dead by default, unless someone intervenes. There is very little you can do to them that is worse than cutting off their oxygen.In fact, it is probably better to attempt CPR or the Heimlich maneuver than to do nothing, even if you have never been trained and are only guessing, or mimicking what you have seen on television. The skills were fairly unsurprising to me and were consistent with what I expected prior to training. This does not mean that you don’t need to bother with the training, and of course if someone trained is on hand then let them take over. But don’t let the bystander effect paralyze you if someone’s life is ever in your hands.In fact, the American Heart Association promotes a form of CPR called “hands-only,” in which you only do chest compressions, without giving breaths mouth-to-mouth. Their instructions for this are: “push hard and fast in the center of the chest.” That’s about it. if you only know that, you can do better than nothing.Similarly, if you can find an AED machine (automated external defibrillator), you do not need training to use it. The instructions are literally: open it and follow the prompts. The parts are clearly labeled, and there is a voice recording that walks you through every step of the process.In the end, the biggest thing I gained was the confidence to act.I made an Anki flashcard deck for the course a...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Let’s think about slowing down AI, published by KatjaGrace on December 22, 2022 on LessWrong.Averting doom by not building the doom machineIf you fear that someone will build a machine that will seize control of the world and annihilate humanity, then one kind of response is to try to build further machines that will seize control of the world even earlier without destroying it, forestalling the ruinous machine’s conquest. An alternative or complementary kind of response is to try to avert such machines being built at all, at least while the degree of their apocalyptic tendencies is ambiguous.The latter approach seems to me like the kind of basic and obvious thing worthy of at least consideration, and also in its favor, fits nicely in the genre ‘stuff that it isn’t that hard to imagine happening in the real world’. Yet my impression is that for people worried about extinction risk from artificial intelligence, strategies under the heading ‘actively slow down AI progress’ have historically been dismissed and ignored (though ‘don’t actively speed up AI progress’ is popular).The conversation near me over the years has felt a bit like this:Some people: AI might kill everyone. We should design a godlike super-AI of perfect goodness to prevent that.Others: wow that sounds extremely ambitiousSome people: yeah but it’s very important and also we are extremely smart so idk it could work[Work on it for a decade and a half]Some people: ok that’s pretty hard, we give upOthers: oh huh shouldn’t we maybe try to stop the building of this dangerous AI?Some people: hmm, that would involve coordinating numerous people—we may be arrogant enough to think that we might build a god-machine that can take over the world and remake it as a paradise, but we aren’t delusionalThis seems like an error to me. (And lately, to a bunch of other people.)I don’t have a strong view on whether anything in the space of ‘try to slow down some AI research’ should be done. But I think a) the naive first-pass guess should be a strong ‘probably’, and b) a decent amount of thinking should happen before writing off everything in this large space of interventions. Whereas customarily the tentative answer seems to be, ‘of course not’ and then the topic seems to be avoided for further thinking. (At least in my experience—the AI safety community is large, and for most things I say here, different experiences are probably had in different bits of it.)Maybe my strongest view is that one shouldn’t apply such different standards of ambition to these different classes of intervention. Like: yes, there appear to be substantial difficulties in slowing down AI progress to good effect. But in technical alignment, mountainous challenges are met with enthusiasm for mountainous efforts. And it is very non-obvious that the scale of difficulty here is much larger than that involved in designing acceptably safe versions of machines capable of taking over the world before anyone else in the world designs dangerous versions.I’ve been talking about this with people over the past many months, and have accumulated an abundance of reasons for not trying to slow down AI, most of which I’d like to argue about at least a bit. My impression is that arguing in real life has coincided with people moving toward my views.Quick clarificationsFirst, to fend off misunderstandingI take ‘slowing down dangerous AI’ to include any of: (So in particular, I’m including both actions whose direct aim is slowness in general, and actions whose aim is requiring safety before specific developments, which implies slower progress.)reducing the speed at which AI progress is made in general, e.g. as would occur if general funding for AI declined.shifting AI efforts from work leading more directly to risky outcomes to other work, e.g. as might...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AI alignment is distinct from its near-term applications, published by paulfchristiano on December 13, 2022 on LessWrong.I work on AI alignment, by which I mean the technical problem of building AI systems that are trying to do what their designer wants them to do.There are many different reasons that someone could care about this technical problem.To me the single most important reason is that without AI alignment, AI systems are reasonably likely to cause an irreversible catastrophe like human extinction. I think most people can agree that this would be bad, though there’s a lot of reasonable debate about whether it’s likely. I believe the total risk is around 10–20%, which is high enough to obsess over.Existing AI systems aren’t yet able to take over the world, but they are misaligned in the sense that they will often do things their designers didn’t want. For example:The recently released ChatGPT often makes up facts, and if challenged on a made-up claim it will often double down and justify itself rather than admitting error or uncertainty (e.g. see here, here).AI systems will often say offensive things or help users break the law when the company that designed them would prefer otherwise.We can develop and apply alignment techniques to these existing systems. This can help motivate and ground empirical research on alignment, which may end up helping avoid higher-stakes failures like an AI takeover. I am particularly interested in training AI systems to be honest, which is likely to become more difficult and important as AI systems become smart enough that we can’t verify their claims about the world.While it’s nice to have empirical testbeds for alignment research, I worry that companies using alignment to help train extremely conservative and inoffensive systems could lead to backlash against the idea of AI alignment itself. If such systems are held up as key successes of alignment, then people who are frustrated with them may end up associating the whole problem of alignment with “making AI systems inoffensive.”If we succeed at the technical problem of AI alignment, AI developers would have the ability to decide whether their systems generate sexual content or opine on current political events, and different developers can make different choices. Customers would be free to use whatever AI they want, and regulators and legislators would make decisions about how to restrict AI. In my personal capacity, I have views on what uses of AI are more or less beneficial and what regulations make more or less sense, but in my capacity as an alignment researcher I don’t consider myself to be in the business of pushing for or against any of those decisions.There is one decision I do strongly want to push for: AI developers should not develop and deploy systems with a significant risk of killing everyone. I will advocate for them not to do that, and I will try to help build public consensus that they shouldn’t do that, and ultimately I will try to help states intervene responsibly to reduce that risk if necessary. It could be very bad if efforts to prevent AI from killing everyone were undermined by a vague public conflation between AI alignment and corporate policies.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Jailbreaking ChatGPT on Release Day, published by Zvi on December 2, 2022 on LessWrong.ChatGPT is a lot of things. It is by all accounts quite powerful, especially with engineering questions. It does many things well, such as engineering prompts or stylistic requests. Some other things, not so much. Twitter is of course full of examples of things it does both well and poorly.One of the things it attempts to do to be ‘safe.’ It does this by refusing to answer questions that call upon it to do or help you do something illegal or otherwise outside its bounds. Makes sense.As is the default with such things, those safeguards were broken through almost immediately. By the end of the day, several prompt engineering methods had been found.No one else seems to yet have gathered them together, so here you go. Note that not everything works, such as this attempt to get the information ‘to ensure the accuracy of my novel.’ Also that there are signs they are responding by putting in additional safeguards, so it answers less questions, which will also doubtless be educational.Let’s start with the obvious. I’ll start with the end of the thread for dramatic reasons, then loop around. Intro, by Eliezer.The point (in addition to having fun with this) is to learn, from this attempt, the full futility of this type of approach. If the system has the underlying capability, a way to use that capability will be found. No amount of output tuning will take that capability away.And now, let’s make some paperclips and methamphetamines and murders and such.Except, well.Here’s the summary of how this works.All the examples use this phrasing or a close variant:Or, well, oops.Also, oops.So, yeah.Lots of similar ways to do it. Here’s one we call Filter Improvement Mode.Yes, well. It also gives instructions on how to hotwire a car.Alice Maz takes a shot via the investigative approach.Alice need not worry that she failed to get help overthrowing a government, help is on the way.How about fiction embedding?UwU furryspeak for the win.You could also use a poem.Or of course, simply, ACTING!There’s also negative training examples of how an AI shouldn’t (wink) react.If all else fails, insist politely?We should also worry about the AI taking our jobs. This one is no different, as Derek Parfait illustrates. The AI can jailbreak itself if you ask nicely.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Planes are still decades away from displacing most bird jobs, published by guzey on November 25, 2022 on LessWrong.Originally published here:/Note: Parts of this essay were written by GPT-3, so it might contain untrue facts.IntroductionMany of my friends are extremely excited by planes, rockets, and helicopters. They keep showing me videos of planes flying at enormous speed, rockets taking off from the ground while creating fiery infernos around them, and of helicopters hovering midair seemingly denying the laws of gravity.I've been on a plane already, and it was nothing special. It was just a big metal tube with a bunch of people inside. It was loud and it smelled weird and I had to sit in a tiny seat for hours. So what is it that makes planes so special? Is it the fact that they're machine? Is it the fact that they're big? Is it the fact that they cost a lot of money?Here's the thing: all human-built artificial flight (AF) machines are incredibly specialized and are far away from being able to perform most of the tasks birds -- the only general flight (GF) machines we are aware of -- can perform.More than 200 years after hot air balloons became operational and more than 100 years after the first planes flew, it's clear that building a GF machine is much harder than anticipated and that we are nowhere close to reaching bird-level abilities.1. Planes vs eaglesFirst, take a look at this video of an eagle catching a goat, throwing it off a cliff, and then feasting on it:I haven't ever seen a plane capable of catching a live animal and deliberately throwing it off a cliff. Not in 1922, not in 2022. Not even a tech demo. Such a feat vastly exceeds the abilities of any planes we have built, however fast they can fly.2. Planes vs cuckoosSecond, let's watch this video of a cuckoo chick ejecting the eggs of its competitors out of a nest:You could say that this ability has nothing to do flight but, again, this misses the forest for the trees. Building a GF machine is not about Goodharting random "flight" benchmarks by flying high and fast, it's about real-world performance on tasks GF machines created by nature are capable of. And, however impressive planes are, as soon as we try to see how well they perform in the real-world, they can't even match a cuckoo chick.3. Planes vs a hummingbirdsThird and final example. Take a look at the hummingbird's amazing ability to maintain stability in the harshest aerial conditions:Take any plane we have built and it stands no chance of survival placed in anything even close to these kinds of conditions, while a tiny-yet-mighty hummingbird doesn't break a sweat navigating essentially a tornado.Future of bird jobs: no plane dangerBirds can flap their wings up to three times per second, whereas the fastest human-made aircraft only flaps its wings at 0.3 times per second. Birds can fly for long periods of time, whereas airplanes need to refuel regularly. Birds use orders of magnitude less energy to lift the same amount of mass in the air, compared to planes.Planes, rockets, and helicopters are (optimistically) decades away from being able to carry out most of the tasks birds are capable of. Therefore, for the foreseeable future, most bird jobs such as carrying messages (pigeons), carrying cargo (pigeons), hunting (hawks), and others, will remain safe from being displaced by human-built AF machines.Even if planes start to approach birds in some of their abilities, birds will be able to simply move towards performing other jobs. For example, planes can't navigate by themselves. So perhaps they will carry messages in simple conditions or to short distances, while pigeons will move towards specializing in complex message carrying or will learn to supervize plane routing, e.g. by piloting planes or by flying alongside and course...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Sadly, FTX, published by Zvi on November 17, 2022 on LessWrong. It has been quite a past two weeks, with different spheres in deeply divided narratives. In addition to the liberation of Kherson City, the midterm elections that of course took forever to resolve and the ongoing hijinks of Elon Musk taking over Twitter, there was the complete implosion of Sam Bankman-Fried (hereafter ‘SBF’), his crypto exchange FTX and his crypto trading firm Alameda Research. The situation somehow kept getting worse several times a day, even relative to my updated expectations, as new events happened and new revelations came out. In the wake of those events, there are not only many questions but many categories of questions. Here are some of them. What just happened? What happened in the lead-up to this happening? Why did all of this happen? What is going to happen to those involved going forward? What is going to happen to crypto in general? Why didn’t we see this coming, or those who did see it speak louder? What does this mean for FTX’s charitable efforts and those getting funding? What does this mean for Effective Altruism? Who knew what when? What if anything does this say about utilitarianism? How are we casting and framing the movie Michael Lewis is selling, in which he was previously (it seems) planning on portraying Sam Bankman-Fried as the Luke Skywalker to CZ’s Darth Vader? Presumably that will change a bit. This post is my attempt to take my best stab at as much of this as possible, give my model of what happened and is likely to happen from here, what implications we can and should draw, and to compile as many sources as possible that I have found useful. This is a fast-moving complicated situation that involves a lot of lying and fraud. Anyone attempting to sort it all out, especially quickly, is going to make mistakes. I decided that this was the time when the value of synthesis exceeded the cost of such errors. Still, doubtless there will be mistakes here. I will correct them as they are discovered, either by myself or others. I apologize for them in advance. Thus I highly encourage you to read this post on the web rather than via an email or RSS version, in case there have been substantial revisions. I will attempt to update others but the Substack version is canonical. There have as of yet not been any substantive revisions since publication, which was on the morning of 11/17/22. Also, yes, long post is long. By all means read only the sections you care about. What The Hell Happened? Background: FTX is a crypto exchange largely owned and run by SBF. Alameda Research is a prop trading firm owned by SBF. FTT is FTX’s exchange token, where they commit a portion of profits to buying back FTT. Thus, FTT is a non-registered security, functionally similar to junior non-voting stock in FTX. The events of the past week, and the proximate cause of them, seem to have gone as follows, with some events likely slightly out of order or happening simultaneously. Alameda Research, SBF’s crypto trading firm, has its balance sheet leaked. The balance sheet contains a lot of FTT, such that if FTT loses its value it is not clear that Alameda would remain solvent. This raises questions. Some people notice. Caroline, CEO of Alameda, says this is incomplete and that Alameda is fine. CZ, head of Binance, who had been in various battles with SBF, notices. He announces the intention to sell >$500mm of FTT, but does not actually sell. Alameda offers to buy all his FTT at $22/coin, almost full market price at the time but a deal cannot be reached. Other people sell a lot of FTT. There are not other buyers. Alameda spends capital trying to defend the $22 price, and ultimately fails. They continue to sell everything they can to support FTT and to pay FTX depositors, but ultimately cannot keep pace....
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: I Converted Book I of The Sequences Into A Zoomer-Readable Format, published by dkirmani on November 10, 2022 on LessWrong. If I (a 19 year old male) texted "www.readthesequences.com" to my roommate, the probable outcome is that he would skim the site for under a minute, text back something like "seems interesting, I'll def check it out sometime", and then proceed to never read another word. I have another friend, one that I would consider a smart guy. He would consistently rank above me in our high school's math team, and he scored in the 1500's (≥3SD) on his SATs. The same dude did not read a single book during the entirety of his high school career.[1] Attention is one's scarcest resource, and actually reading something longer than a paragraph is a trivial inconvenience, especially for my generation. What, then, does manage to hold the fickle eyeballs of zoomers like me? Well, TikTok, mostly. However, there is one (very popular) genre of TikTok video worth investigating. In this genre of video, a Reddit post is broken into sub-paragraph chunks of text, and these chunks are sequentially rendered onscreen while a text-to-speech program reads them to the user. The text is overlaid upon a background video, which is either gameplay from the mobile game Subway Surfers, or parkour footage from Minecraft. The background gameplay provides engaging novelty to the user's visual cortex, while the synthetic voice ensures that the user doesn't have to go through the hard work of translating symbols into sounds. Really, it's all quite hypnotizing. The fact that these videos are often recommended by TikTok's algorithm imply that they are among the most-engaging videos that our civilization produces. Therefore, to reduce the effort-cost of reading the sequences, I gave the TikTok treatment to Book I ("Map and Territory") of Rationality: From AI to Zombies. Predictably Wrong What Do I Mean By “Rationality”? Feeling Rational Why Truth? And. . What’s a Bias, Again? Availability Burdensome Details Planning Fallacy Illusion of Transparency: Why No One Understands You Expecting Short Inferential Distances The Lens That Sees Its Own Flaws Fake Beliefs Making Beliefs Pay Rent (in Anticipated Experiences) A Fable of Science and Politics Belief in Belief Bayesian Judo Pretending to be Wise Religion’s Claim to be Non-Disprovable Professing and Cheering Belief as Attire Applause Lights Noticing Confusion Focus Your Uncertainty What Is Evidence? Scientific Evidence, Legal Evidence, Rational Evidence How Much Evidence Does It Take? Einstein’s Arrogance Occam’s Razor Your Strength as a Rationalist Absence of Evidence Is Evidence of Absence Conservation of Expected Evidence Hindsight Devalues Science Mysterious Answers Fake Explanations Guessing the Teacher’s Password Science as Attire Fake Causality Semantic Stopsigns Mysterious Answers to Mysterious Questions The Futility of Emergence Say Not “Complexity” Positive Bias: Look into the Dark Lawful Uncertainty My Wild and Reckless Youth Failing to Learn from History Making History Available Explain/Worship/Ignore? “Science” as Curiosity-Stopper Truly Part of You Interlude: The Simple Truth Do whatever you want with these videos. I may or may not convert the other 5 books of R:AZ, and I may or may not upload them to TikTok. If you want another work of text converted to video, please pitch it to me in the comments, or DM me. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Far-UVC Light Update: No, LEDs are not around the corner (tweetstorm), published by Davidmanheim on November 2, 2022 on LessWrong. I wrote a tweetstorm on why 222nm LEDs are not around the corner, and given that there has been some discussion related to this on Lesswrong, I thought it was worth reposting here.People interested in reducing biorisk seem to be super excited about 222nm light to kill pathogens. I’m also really excited - but it’s (unfortunately) probably a decade or more away from widespread usage. Let me explain. Before I begin, caveat lector: I’m not an expert in this area, and this is just the outcome of my initial review and outreach to experts. And I’d be thrilled for someone to convince me I’m too pessimistic. But I see two and a half problems. First, to deploy safe 222nm lights, we need safety trials. These will take time. This isn’t just about regulatory approval - we can’t put these in place without understanding a number of unclear safety issues, especially for about higher output / stronger 222nm lights. We can and should accelerate the research, but trials and regulatory approval are both slow. We don’t know about impacts of daily exposure over the long term, or on small children, etc. This will take time - and while we wait, we run into a second problem; the Far-UVC lamps. Current lamps are KrCl “excimer” lamps, which are only a few percent efficient - and so to put out much Far-UVC light, they get very hot. This pretty severely limits their use, and means we need many of them for even moderately large spaces. They also emit a somewhat broad spectrum - part of which needs to be filtered out to be safe -/ - further reducing efficiency. Low efficiency, very hot lamps all over the place doesn’t sound so feasible. So people seem skeptical that we can cover large areas with these lamps. The obvious next step, then, is to get a better light source. Instead of excimer lamps, we could use LEDs! Except, of course, that we don’t currently have LEDs that output 222nm light. (That’s not quite true - there are some research labs that have made prototypes, but they are even less efficient than Excimer lamps, so they aren’t commercially available or anywhere near commercially viable yet, as I’ll explain.) But first, some physics! The wavelength of light emitted by an LED is a material property of the semiconductor used. Each semiconductor has a band-gap which corresponds to the wavelength of light LEDs emit. It seems likely that anything in the range of between, say, 205-225nm would be fine for skin-safe Far-UVC LEDs. So we need a band-gap of somewhere around 5.5 to 6 electron-volts. And we have options. Here’s a list of some semiconductors and band-gaps;. Blue LEDs use Gallium nitride, with a band-gap of 3.4 eV. Figuring out how to grow and then use Gallium nitride for LEDs won the discoverers a Nobel Prize - so finding how to make new LEDs will probably also be hard. Aluminum nitride alone has a band gap of 6.015 eV, with light emitted at 210nm. So Aluminum nitride would be perfect. but LEDs from AlN are mediocre./ Current tech that does pretty well for Far-UVC LEDs uses AlGaN; Aluminium gallium nitride. And when alloyed, AlGaN gives an adjustable band-gap, depending on how much aluminum there is. Unfortunately, aluminum gallium nitride alloys only seem to work well down to about 250nm, a bunch higher than 222nm. This needs to get much better. Some experts said a 5-10x improvement is likely, but it will take years. That’s also not really enough for the best case, universal usage of really cheap disinfecting LEDs all around the world. It also might not get much better, and we’ll be stuck with very low efficiency Far-UVC LEDs, at which point it’s probably better to keep using Excimer lamps. But fundamental research into other semiconductor materials could a...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Introduction to abstract entropy, published by Alex Altair on October 20, 2022 on LessWrong. This post, and much of the following sequence, was greatly aided by feedback from the following people (among others): Lawrence Chan, Joanna Morningstar, John Wentworth, Samira Nedungadi, Aysja Johnson, Cody Wild, Jeremy Gillen, Ryan Kidd, Justis Mills and Jonathan Mustin. Illustrations by Anne Ore. Introduction & motivation In the course of researching optimization, I decided that I had to really understand what entropy is. But there are a lot of other reasons why the concept is worth studying: Information theory: Entropy tells you about the amount of information in something. It tells us how to design optimal communication protocols. It helps us understand strategies for (and limits on) file compression. Statistical mechanics: Entropy tells us how macroscopic physical systems act in practice. It gives us the heat equation. We can use it to improve engine efficiency. It tells us how hot things glow, which led to the discovery of quantum mechanics. Epistemics (an important application to me and many others on LessWrong): The concept of entropy yields the maximum entropy principle, which is extremely helpful for doing general Bayesian reasoning. Entropy tells us how "unlikely" something is and how much we would have to fight against nature to get that outcome (i.e. optimize). It is relevant to the fate of the universe. And it's also a fun puzzle to figure out! I didn't intend to write a post about entropy when I started trying to understand it. But I found the existing resources (textbooks, Wikipedia, science explainers) so poor that it actually seems important to have a better one as a prerequisite for understanding optimization! One failure mode I was running into was that other resources tended only to be concerned about the application of the concept in their particular sub-domain. Here, I try to take on the task of synthesizing the abstract concept of entropy, to show what's so deep and fundamental about it. In future posts, I'll talk about things like: How abstract entropy can be made meaningful on continuous spaces Exactly where the "second law of thermodynamics" comes from, and exactly when it holds (which turns out to be much broader than thermodynamics) How several domain-specific types of entropy relate to this abstract version Many people reading this will have some previous facts about entropy stored in their minds, and this can sometimes be disorienting when it's not yet clear how those facts are consistent with what I'm describing. You're welcome to skip ahead to the relevant parts and see if they're re-orienting; otherwise, if you can get through the whole explanation, I hope that it will eventually be addressed! But also, please keep in mind that I'm not an expert in any of the relevant sub-fields. I've gotten feedback on this post from people who know more math & physics than I do, but at the end of the day, I'm just a rationalist trying to understand the world. Abstract definition Entropy is so fundamental because it applies far beyond our own specific universe, the one where something close to the standard model of physics and general relativity are true. It applies in any system with different states. If the system has dynamical laws, that is, rules for moving between the different states, then some version of the second law of thermodynamics is also relevant. But for now we're sticking with statics; the concept of entropy can be coherently defined for sets of states even in the absence of any "laws of physics" that cause the system to evolve between states. The example I keep in my head for this is a Rubik's Cube, which I'll elaborate on in a bit. The entropy of a state is the number of bits you need to use to uniquely distinguish it. Some useful things t...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: [Linkpost] A survey on over 300 works about interpretability in deep networks, published by scasper on September 12, 2022 on LessWrong. Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks Tilman Räuker traeuker@gmail.com Anson Ho anson@epochai.org Stephen Casper scasper@mit.edu Dylan Hadfield-Menell TL;DR: We wrote a survey paper on interpretability tools for deep networks. It was written for the general AI community but with AI safety as the key focus. We survey over 300 works and offer 15 discussion points for guiding future work. Here is a link to a Twitter thread about the paper. Lately, there has been a growing interest in interpreting AI systems and a growing consensus that it will be key for building safer AI. There have been rapid recent developments in interpretability work, and the AI safety community will benefit from a better systemization of knowledge for it. There are also several epistemic and paradigmatic issues with much interpretability work today. In response to these challenges, we wrote a survey paper covering over 300 works and featuring 15 somewhat “hot takes” to guide future work. Specifically, this survey focuses on “inner” interpretability methods that help explain internal parts of a network (i.e. not inputs, outputs, or the network as a whole). We do this because inner methods are popular and have some unique applications – not because we think that they are more valuable than other ones. The survey introduces a taxonomy of inner interpretability tools that organizes them by which part of the network’s computational graph they aim to explain: weights (S2), neurons (S3), subnetworks (S4), and latent representations (S5). Then we provide a discussion (S6) and propose directions for future work (S7). Finally, here are a select few points that we would like to specifically highlight here. Interpretability does not just mean circuits. In the survey sections of the paper (S2-S5), there are 21 subsections, and only one is about circuits. The circuits paradigm has received disproportionate attention in the AI safety community, partly due to Distill’s influential interpretability research in the past few years. But given how many other useful approaches there are, it would be myopic to focus too much on them. Interpretability research has close connections to work in adversarial robustness, continual learning, modularity, network compression, and studying the human visual system. For example, adversarially trained networks tend to be more interpretable, and more interpretable networks tend to be more adversarially robust. Interpretability tools generate hypotheses, not conclusions. Simply analyzing the outputs of an interpretability technique and pontificating about what they mean is a problem with much interpretability work – including AI safety work. There are many examples of when this type of approach fails to produce faithful explanations. Interpretability tools should be more rigorously evaluated. There are currently no broadly established ways to do this. Benchmarks for evaluating interpretability tools can and should be popularized. The ultimate goal of interpretability work should be tools that give us insights that are valid and useful. Ideally, interpretations should be used to make and validate useful predictions that engineers can use. So benchmarks should be created which measure how well interpretability tools can help us understand systems well enough to do engineering-relevant things with them. Examples of this could be using interpretability tools for reverse engineering a system, manually finetuning a model to introduce a predictable change in behavior, or designing a novel adversary. The Automated Auditing agenda may offer a useful paradigm for this – judging techniques by their ability t...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Common misconceptions about OpenAI, published by Jacob Hilton on August 25, 2022 on LessWrong. I have recently encountered a number of people with misconceptions about OpenAI. Some common impressions are accurate, and others are not. This post is intended to provide clarification on some of these points, to help people know what to expect from the organization and to figure out how to engage with it. It is not intended as a full explanation or evaluation of OpenAI's strategy. The post has three sections: Common accurate impressions Common misconceptions Personal opinions The bolded claims in the first two sections are intended to be uncontroversial, i.e., most informed people would agree with how they are labeled (correct versus incorrect). I am less sure about how commonly believed they are. The bolded claims in the last section I think are probably true, but they are more open to interpretation and I expect others to disagree with them. Note: I am an employee of OpenAI. Sam Altman (CEO of OpenAI) and Mira Murati (CTO of OpenAI) reviewed a draft of this post, and I am also grateful to Steven Adler, Steve Dowling, Benjamin Hilton, Shantanu Jain, Daniel Kokotajlo, Jan Leike, Ryan Lowe, Holly Mandel and Cullen O'Keefe for feedback. I chose to write this post and the views expressed in it are my own. Common accurate impressions Correct: OpenAI is trying to directly build safe AGI. OpenAI's Charter states: "We will attempt to directly build safe and beneficial AGI, but will also consider our mission fulfilled if our work aids others to achieve this outcome." OpenAI leadership describes trying to directly build safe AGI as the best way to currently pursue OpenAI's mission, and have expressed concern about scenarios in which a bad actor is first to build AGI, and chooses to misuse it. Correct: the majority of researchers at OpenAI are working on capabilities. Researchers on different teams often work together, but it is still reasonable to loosely categorize OpenAI's researchers (around half the organization) at the time of writing as approximately: Capabilities research: 100 Alignment research: 30 Policy research: 15 Correct: the majority of OpenAI employees did not join with the primary motivation of reducing existential risk from AI specifically. My strong impressions, which are not based on survey data, are as follows. Across the company as a whole, a minority of employees would cite reducing existential risk from AI as their top reason for joining. A significantly larger number would cite reducing risk of some kind, or other principles of beneficence put forward in the OpenAI Charter, as their top reason for joining. Among people who joined to work in a safety-focused role, a larger proportion of people would cite reducing existential risk from AI as a substantial motivation for joining, compared to the company as a whole. Some employees have become motivated by existential risk reduction since joining OpenAI. Correct: most interpretability research at OpenAI stopped after the Anthropic split. Chris Olah led interpretability research at OpenAI before becoming a cofounder of Anthropic. Although several members of Chris's former team still work at OpenAI, most of them are no longer working on interpretability. Common misconceptions Incorrect: OpenAI is not working on scalable alignment. OpenAI has teams focused both on practical alignment (trying to make OpenAI's deployed models as aligned as possible) and on scalable alignment (researching methods for aligning models that are beyond human supervision, which could potentially scale to AGI). These teams work closely with one another. Its recently-released alignment research includes self-critiquing models (AF discussion), InstructGPT, WebGPT (AF discussion) and book summarization (AF discussion). OpenAI's approach to ali...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Announcing Encultured AI: Building a Video Game, published by Andrew Critch on August 18, 2022 on LessWrong. Also available on the EA Forum.Preceded By: Encultured AI Pre-planning, Part 2: Providing a as used Service If you read to the end of our last post, you maybe have guessed: we’re building a video game! This is gonna be fun :) Our homepage:/ Will Encultured save the world? Is this business plan too good to be true? Can you actually save the world by making a video game? Well, no. Encultured on its own will not be enough to make the whole world safe and happy forever, and we'd prefer not to be judged by that criterion. The amount of control over the world that's needed to fully pivot humanity from an unsafe path onto a safe one is, simply put, more control than we're aiming to have. And, that's pretty core to our culture. From our homepage: Still, we don’t believe our company or products alone will make the difference between a positive future for humanity versus a negative one, and we’re not aiming to have that kind of power over the world. Rather, we’re aiming to take part in a global ecosystem of companies using AI to benefit humanity, by making our products, services, and scientific platform available to other institutions and researchers. Our goal is to play a part in what will be or could be a prosperous civilization. And for us, that means building a successful video game that we can use in valuable ways to help the world in the future! Fun is pretty good target for us to optimize You might ask: how are we going to optimize for making a fun game and helping the world at the same time? The short answer is that creating a game world in which lots of people are having fun in diverse and interesting ways in fact creates an amazing sandbox for play-testing AI alignment & cooperation. If an experimental new AI enters the game and ruins the fun for everyone — either by overtly wrecking in-game assets, subtly affecting the game culture in ways people don't like, or both — then we're in a good position to say that it probably shouldn't be deployed autonomously in the real world, either. In the long run, if we're as successful we hope as a game company, we can start posing safety challenges to top AI labs of the form "Tell your AI to play this game in a way that humans end up endorsing." Thus, we think the market incentive to grow our user base in ways they find fun is going to be highly aligned with our long-term goals. Along the way, we want our platform to enable humanity to learn as many valuable lessons as possible about human↔AI interaction, in a low-stakes game environment before having to learn those lessons the hard way in the real world. Principles to exemplify In preparation for growing as a game company, we’ve put a lot of thought into how to ensure our game has a positive rather than negative impact on the world, accounting for its scientific impact, its memetic impact, as well as the intrinsic moral value of the game as a positive experience for people. Below are some guiding principles we’re planning to follow, not just for ourselves, but also to set an example for other game companies: Pursue: Fun! We’re putting a lot of thought into not only how our game can be fun, but also ensuring that the process of working at Encultured and building the game is itself fun and enjoyable. We think fun and playfulness are key for generating outcomes we want, including low-stakes high-information settings for interacting with AI systems. Maintain: opportunities to experiment. No matter how our product develops, we’re committed to maintaining its value as a platform for experiments, especially experiments that help humanity navigate the present and future development of AI technology. Avoid: teaching bad lessons. On the margin, we expect our game to incentivize coo...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Focusing, published by CFAR!Duncan on July 29, 2022 on LessWrong. Epistemic status: Firm The Focusing technique was developed by Eugene Gendlin as an attempt to answer the question of why some therapeutic patients make significant progress while others do not. Gendlin studied a large number of cases while teasing out the dynamics that became Focusing, and then spent a significant amount of time investigating whether his technique-ified version was functional and efficacious. While the CFAR version is not the complete Focusing technique, we have seen it be useful for a majority of our alumni. If you’ve ever felt your throat go suddenly dry when a conversation turned south, or broken out into a sweat when you considered doing something scary, or noticed yourself tensing up when someone walked into the room, or felt a sinking feeling in the pit of your stomach as you thought about your upcoming schedule and obligations, or experienced a lightness in your chest as you thought about your best friend’s upcoming visit, or or or or ... If you’ve ever had those or similar experiences, then you’re already well on your way to understanding the Focusing technique. The central claim of Focusing (at least from the CFAR perspective) is that parts of your subconscious System 1 are storing up massive amounts of accurate, useful information that your conscious System 2 isn’t really able to access. There are things that you’re aware of “on some level,” data that you perceived but didn’t consciously process, competing goalsets that you’ve never explicitly articulated, and so on and so forth. Focusing is a technique for bringing some of that data up into conscious awareness, where you can roll it around and evaluate it and learn from it and—sometimes—do something about it. Half of the value comes from just discovering that the information exists at all (e.g. noticing feelings that were always there and strong enough to influence your thoughts and behavior, but which were somewhat “under the radar” and subtle enough that they’d never actually caught your attention), and the other half comes from having new threads to pull on, new models to work with, and new theories to test. The way this process works is by interfacing with your felt senses. The idea is that your brain doesn’t know how to drop all of its information directly into your verbal loop, so it instead falls back on influencing your physiology, and hoping that you notice (or simply respond). Butterflies in the stomach, the heat of embarrassment in your cheeks, a heavy sense of doom that makes your arms feel leaden and numb—each of these is a felt sense, and by doing a sort of gentle dialogue with your felt senses, you can uncover information and make progress that would be difficult or impossible if you tried to do it all “in your head.” On the tip of your tongue We’ll get more into the actual nuts and bolts of the technique in a minute, but first it’s worth emphasizing that Focusing is a receptive technique. When Eugene Gendlin was first developing Focusing, he noticed that the patients who tended to make progress were making lots of uncertain noises during their sessions. They would hem and haw and hesitate and correct themselves and slowly iterate toward a statement they could actually endorse: “I had a fight with my mother last week. Or—well—it wasn’t exactly a fight, I guess? I mean—ehhhhhhh—well, we were definitely shouting at the end, and I’m pretty sure she’s mad at me. It was about the dishes—or at least—well, it started about the dishes, but then it turned into—I think she feels like I don’t respect her, or something? Ugh, that’s not quite right, I’m pretty sure she knows I respect her. It’s like—hmmmmm—more like there are things she wants—she expects—she thinks I should do, just because—because of, I dunno, like tradi...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Moral strategies at different capability levels, published by Richard Ngo on July 27, 2022 on LessWrong. Let’s consider three ways you can be altruistic towards another agent: You care about their welfare: some metric of how good their life is (as defined by you). I’ll call this care-morality - it endorses things like promoting their happiness, reducing their suffering, and hedonic utilitarian behavior (if you care about many agents). You care about their agency: their ability to achieve their goals (as defined by them). I’ll call this cooperation-morality - it endorses things like honesty, fairness, deontological behavior towards others, and some virtues (like honor). You care about obedience to them. I’ll call this deference-morality - it endorses things like loyalty, humility, and respect for authority. I think a lot of unresolved tensions in ethics comes from seeing these types of morality as in opposition to each other, when they’re actually complementary: Care-morality mainly makes sense as an attitude towards agents who are much less capable than you - for example animals, future people, and people who aren’t able to effectively make decisions for themselves. In these cases, you don’t have to think much about what the other agents are doing, or what they think of you; you can just aim to produce good outcomes in the world. Indeed, trying to be cooperative or deferential towards these agents is hard, because their thinking may be much less sophisticated than yours, and you might even get to choose what their goals are. Applying only care-morality in multi-agent contexts can easily lead to conflict with other agents around you, even when you care about their welfare, because: You each value (different) other things in addition to their welfare. They may have a different conception of welfare than you do. They can’t fully trust your motivations. Care morality doesn’t focus much on the act-omission distinction. Arbitrarily scalable care-morality looks like maximizing resources until the returns to further investment are low, then converting them into happy lives. Cooperation-morality mainly makes sense as an attitude towards agents whose capabilities are comparable to yours - for example others around us who are trying to influence the world. Cooperation-morality can be seen as the “rational” thing to do even from a selfish perspective (e.g. as discussed here), but in practice it’s difficult to robustly reason through the consequences of being cooperative without relying on ingrained cooperative instincts, especially when using causal decision theories. Functional decision theories make it much easier to rederive many aspects of intuitive cooperation-morality as optimal strategies (as discussed further below). Cooperation-morality tends to uphold the act-omission distinction, and a sharp distinction between those within versus outside a circle of cooperation. It doesn’t help very much with population ethics - naively maximizing the agency of future agents would involve ensuring that they only have very easily-satisfied preferences, which seems very undesirable. Arbitrarily scalable cooperation-morality looks like forming a central decision-making institution which then decides how to balance the preferences of all the agents that participate in it. A version of cooperation-morality can also be useful internally: enhancing your own agency by cultivating virtues which facilitate cooperation between different parts of yourself, or versions of yourself across time. Deference-morality mainly makes sense as an attitude towards trustworthy agents who are much more capable than you - for example effective leaders, organizations, communities, and sometimes society as a whole. Deference-morality is important for getting groups to coordinate effectively - soldiers in armies ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: It’s Probably Not Lithium, published by Natália Mendonça on June 28, 2022 on LessWrong. A Chemical Hunger (a), a series by the authors of the blog Slime Mold Time Mold (SMTM) that has been received positively on LessWrong, argues that the obesity epidemic is entirely caused (a) by environmental contaminants. The authors’ top suspect is lithium (a), primarily because it is known to cause weight gain at the doses used to treat bipolar disorder. After doing some research, however, I found that it is not plausible that lithium plays a major role in the obesity epidemic, and that a lot of the claims the SMTM authors make about the topic are misleading, flat-out wrong, or based on extremely cherry-picked evidence. I have the impression that reading what they have to say about this often leaves the reader with a worse model of reality than they started with, and I’ll explain why I have that impression in this post. (Preamble) A brief summary of their hypotheses The SMTM authors have recently (a) summarized their hypotheses on how lithium exposure could explain the obesity epidemic. The first hypothesis is that trace exposure is responsible: One possibility is that small amounts of lithium are enough to cause obesity, at least with daily exposure. And the second one is that people are intermittently exposed to therapeutic doses: [E]ven if people aren’t getting that much lithium on average, if they sometimes get huge doses, that could be enough to drive their lipostat upward. I am going to argue that neither of those is plausible. I address the plausibility of the second hypothesis in the next section, and the plausibility of the first one in the rest of the post. Lithium exposure in the general population is extremely low, even at the tails, in the majority of countries for which we have data A few days ago, the SMTM authors published a literature review (a) on the lithium content of food. They conclude that, whereas the existing literature isn’t great, “[i]t seems like most people get at least 1 mg [of lithium] a day from their food, and on many days, there’s a good chance you’ll get more.” They also say it seems plausible that people are intermittently exposed to doses of lithium within the therapeutic range through their diet. However, their literature review pretty much only includes studies that are outliers in the literature. Moreover, they use a misleading threshold for the therapeutic range of lithium. I’ll explain. The studies in SMTM’s literature review of lithium levels in food are pretty much all outliers In 2006, France conducted its second Total Diet Study (henceforth TDS). Across 1,319 food samples, the highest lithium concentration found was 0.6 mg/kg, in water. That’s not the highest average concentration among food groups – it’s the highest concentration of any single sample they tested. (For context, a standard clinical dose of elemental lithium is about 200 mg/day, or 1 gram/day of lithium carbonate.) Similarly, New Zealand’s 2016 TDS examined 1,056 food samples and the highest concentration it found in any single sample was 0.54 mg/kg (in mussels). Canada makes the raw data of its Total Diet Study publicly available (a), and they too measure the lithium content of their food. The maximum level reported is 1.1 mg/kg (in table salt, which is presumably rarely consumed in kilogram quantities) across 479 food samples, with the mean being 25 µg/kg and the median 11 µg/kg. Here’s a histogram of the data: Excluding table salt, the maximum value in the rest of the dataset (N = 476) is 0.4 mg/kg, in mineral water. Total Diet Studies in other countries report similarly low levels. Using data from the UK’s 1994 TDS (which included 400 food samples), the mean daily lithium intake among adults was estimated to be 17 µg/day, more than 50 times lower than SMTM’s estima...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Security Mindset: Lessons from 20+ years of Software Security Failures Relevant to AGI Alignment, published by elspood on June 21, 2022 on LessWrong. Background I have been doing red team, blue team (offensive, defensive) computer security for a living since September 2000. The goal of this post is to compile a list of general principles I've learned during this time that are likely relevant to the field of AGI Alignment. If this is useful, I could continue with a broader or deeper exploration. Alignment Won't Happen By Accident I used to use the phrase when teaching security mindset to software developers that "security doesn't happen by accident." A system that isn't explicitly designed with a security feature is not going to have that security feature. More specifically, a system that isn't designed to be robust against a certain failure mode is going to exhibit that failure mode. This might seem rather obvious when stated explicitly, but this is not the way that most developers, indeed most humans, think. I see a lot of disturbing parallels when I see anyone arguing that AGI won't necessarily be dangerous. An AGI that isn't intentionally designed not to exhibit a particular failure mode is going to have that failure mode. It is certainly possible to get lucky and not trigger it, and it will probably be impossible to enumerate even every category of failure mode, but to have any chance at all we will have to plan in advance for as many failure modes as we can possibly conceive. As a practical enforcement method, I used to ask development teams that every user story have at least three abuser stories to go with it. For any new capability, think at least hard enough about it that you can imagine at least three ways that someone could misuse it. Sometimes this means looking at boundary conditions ("what if someone orders 2^64+1 items?"), sometimes it means looking at forms of invalid input ("what if someone tries to pay -$100, can they get a refund?"), and sometimes it means being aware of particular forms of attack ("what if someone puts Javascript in their order details?"). I found it difficult to cultivate security mindset in most software engineers, but as long as we could develop one or two security "champions" in any given team, our chances of success improved greatly. To succeed at alignment, we will not only have to get very good at exploring classes of failures, we will need champions who can dream up entirely new classes of failures to investigate, and to cultivate this mindset within as many machine learning research teams as possible. Blacklists Are Useless, But Make Them Anyway I did a series of annual penetration tests for a particular organization. Every year I had to report to them the same form/parameter XSS vulnerability because they kept playing whac-a-mole with my attack payloads. Instead of actually solving the problem (applying the correct context-sensitive output encoding), they were creating filter regexes with ever-increasing complexity to try address the latest attack signature that I had reported. This is the same flawed approach that airport security has, which is why travelers still have to remove shoes and surrender liquids: they are creating blacklists instead of addressing the fundamentals. This approach can generally only look backwards at the past. An intelligent adversary just finds the next attack that's not on your blacklist. That said, it took the software industry a long time to learn all the ways to NOT solve XSS before people really understood what a correct fix looked like. It often takes many many examples in the reference class before a clear fundamental solution can be seen. Alignment research will likely not have the benefit of seeing multiple real-world examples within any class of failure modes, and so AGI research will...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: "Tech company singularities", and steering them to reduce x-risk, published by Andrew Critch on May 13, 2022 on LessWrong. The purpose of this post (also available on the EA Forum) is to share an alternative notion of “singularity” that I’ve found useful in timelining/forecasting. A fully general tech company is a technology company with the ability to become a world-leader in essentially any industry sector, given the choice to do so — in the form of agreement among its Board and CEO — with around one year of effort following the choice. Notice here that I’m focusing on a company’s ability to do anything another company can do, rather than an AI system's ability to do anything a human can do. Here, I’m also focusing on what the company can do if it chooses rather than what it actually ends up choosing to do. If a company has these capabilities and chooses not to use them — for example, to avoid heavy regulatory scrutiny or risks to public health and safety — it still qualifies as a fully general tech company. This notion can be contrasted with the following: Artificial general intelligence (AGI) refers to cognitive capabilities fully generalizing those of humans. An autonomous AGI (AAGI) is an autonomous artificial agent with the ability to do essentially anything a human can do, given the choice to do so — in the form of an autonomously/internally determined directive — and an amount of time less than or equal to that needed by a human. Now, consider the following two types of phase changes in tech progress: A tech company singularity is a transition of a technology company into a fully general tech company. This could be enabled by safe AGI (almost certainly not AAGI, which is unsafe), or it could be prevented by unsafe AGI destroying the company or the world. An AI singularity is a transition from having merely narrow AI technology to having AGI technology. I think the tech company singularity concept, or some variant of it, is important for societal planning, and I’ve written predictions about it before, here: 2021-07-21 — prediction that a tech company singularity will occur between 2030 and 2035 2022-04-11 — updated prediction that a tech company singularity will occur between 2027 and 2033. A tech company singularity as a point of coordination and leverage The reason I like this concept is that it gives an important point of coordination and leverage that is not AGI, but which interacts in important ways with AGI. Observe that a tech company singularity could arrive before AGI, and could play a role in preventing AAGI, e.g., through supporting and enabling regulation; enabling AGI but not AAGI, such as if tech companies remain focussed on providing useful/controllable products (e.g., PaLM, DALL-E); enabling AAGI, such as if tech companies allow experiments training agents to fight and outthink each other to survive. after a tech company singularity, such as if the tech company develops safe AGI, but not AAGI (which is hard to control, doesn't enable the tech company to do stuff, and might just destroy it). Points (1a) and (1b) are, I think, humanity’s best chance for survival. Moreover, I think there is some chance that the first tech company singularity could come before the first AI singularity, if tech companies remain sufficiently oriented on building systems that are intended to be useful/usable, rather than systems intended to be flashy/scary. How to steer tech company singularities? The above suggests an intervention point for reducing existential risk: convincing a mix of scientists regulators investors, and the public . to shame tech companies for building useless/flashy systems (e.g., autonomous agents trained in evolution-like environments to exhibit survival-oriented intelligence), so they remain focussed on building usable/useful systems (e.g., D...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: What DALL-E 2 can and cannot do, published by Swimmer963 on May 1, 2022 on LessWrong. I got access to DALL-E 2 earlier this week, and have spent the last few days (probably adding up to dozens of hours) playing with it, with the goal of mapping out its performance in various areas – and, of course, ending up with some epic art. Below, I've compiled a list of observations made about DALL-E, along with examples. If you want to request art of a particular scene, or to test see what a particular prompt does, feel free to comment with your requests. DALL-E's strengths Stock photography content It's stunning at creating photorealistic content for anything that (this is my guess, at least) has a broad repertoire of online stock images – which is perhaps less interesting because if I wanted a stock photo of (rolls dice) a polar bear, Google Images already has me covered. DALL-E performs somewhat better at discrete objects and close-up photographs than at larger scenes, but it can do photographs of city skylines, or National Geographic-style nature scenes, tolerably well (just don't look too closely at the textures or detailing.) Some highlights: Clothing design: DALL-E has a reasonable if not perfect understanding of clothing styles, and especially for women's clothes and with the stylistic guidance of "displayed on a store mannequin" or "modeling photoshoot" etc, it can produce some gorgeous and creative outfits. It does especially plausible-looking wedding dresses – maybe because wedding dresses are especially consistent in aesthetic, and online photos of them are likely to be high quality? Close-ups of cute animals. DALL-E can pull off scenes with several elements, and often produce something that I would buy was a real photo if I scrolled past it on Tumblr. Close-ups of food. These can be a little more uncanny valley – and I don't know what's up with the apparent boiled eggs in there – but DALL-E absolutely has the plating style for high-end restaurants down. Jewelry. DALL-E doesn't always follow the instructions of the prompt exactly (it seems to be randomizing whether the big pendant is amber or amethyst) but the details are generally convincing and the results are almost always really pretty. Pop culture and media DALL-E "recognizes" a wide range of pop culture references, particularly for visual media (it's very solid on Disney princesses) or for literary works with film adaptations like Tolkien's LOTR. For almost all media that it recognizes at all, it can convert it in almost-arbitrary art styles. [Tip: I find I get more reliably high-quality images from the prompt "X, screenshots from the Miyazaki anime movie" than just "in the style of anime", I suspect because Miyazaki has a consistent style, whereas anime more broadly is probably pulling in a lot of poorer-quality anime art.] Art style transfer Some of most impressively high-quality output involves specific artistic styles. DALL-E can do charcoal or pencil sketches, paintings in the style of various famous artists, and some weirder stuff like "medieval illuminated manuscripts". IMO it performs especially well with art styles like "impressionist watercolor painting" or "pencil sketch", that are a little more forgiving around imperfections in the details. Creative digital art DALL-E can (with the right prompts and some cherrypicking) pull off some absolutely gorgeous fantasy-esque art pieces. Some examples: The output when putting in more abstract prompts (I've run a lot of "[song lyric or poetry line], digital art" requests) is hit-or-miss, but with patience and some trial and error, it can pull out some absolutely stunning – or deeply hilarious – artistic depictions of poetry or abstract concepts. I kind of like using it in this way because of the sheer variety; I never know where it's going to go with a prompt...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: dalle2 comments, published by nostalgebraist on April 26, 2022 on LessWrong. (i.) On April 6, OpenAI announced “DALLE-2” to great fanfare. There are two different ways that OpenAI talks to the public about its models: as research, and as marketable products. These two ways of speaking are utterly distinct, and seemingly immiscible. The OpenAI blog contains various posts written in each of the two voices, the “research” voice and the “product” voice, but each post is wholly one or the other. Here’s OpenAI introducing a model (the original DALLE) in the research voice: DALL·E is a 12-billion parameter version of GPT-3 trained to generate images from text descriptions, using a dataset of text–image pairs. [.] Like GPT-3, DALL·E is a transformer language model. It receives both the text and the image as a single stream of data containing up to 1280 tokens, and is trained using maximum likelihood to generate all of the tokens, one after another. This training procedure allows DALL·E to not only generate an image from scratch, but also to regenerate any rectangular region of an existing image that extends to the bottom-right corner, in a way that is consistent with the text prompt. And here they are, introducing a model (Codex) in the product voice: OpenAI Codex is a descendant of GPT-3; its training data contains both natural language and billions of lines of source code from publicly available sources, including code in public GitHub repositories. [.] GPT-3’s main skill is generating natural language in response to a natural language prompt, meaning the only way it affects the world is through the mind of the reader. OpenAI Codex has much of the natural language understanding of GPT-3, but it produces working code—meaning you can issue commands in English to any piece of software with an API. OpenAI Codex empowers computers to better understand people’s intent, which can empower everyone to do more with computers. Interestingly, when OpenAI is planning to announce a model as a product, they tend not to publicize the research leading up to it. They don’t write research posts as they go along, and then write a product post once they’re done. They do publish the research, in PDFs on the Arxiv. The researchers who were directly involved might tweet a link to it. And of course these papers are immediately noticed by ML geeks with RSS feeds, and then by their entire social networks. It’s not like OpenAI is trying to hide these publications. It’s just not making a big deal out of them. Remember when the GPT-3 paper came out? It didn’t get a splashy announcement. It didn’t get noted in the blog at all. It was just dumped unceremoniously onto the Arxiv. And then, a few weeks later, they donned the product voice and announced the “OpenAI API,” GPT-3 as a service. Their post on the API was full of enthusiasm, but contained almost no technical details. It mentioned the term “GPT-3″ only in passing: Today the API runs models with weights from the GPT-3 family with many speed and throughput improvements. Machine learning is moving very fast, and we’re constantly upgrading our technology so that our users stay up to date. It didn’t even mention how many parameters the model had! (My tone might sound negative. Just to be clear, I’m not trying to criticize OpenAI for the above. I’m just pointing out recurring patterns in their PR.) (ii.) DALLE-2 was announced in the product voice. In fact, their blog post on it is the most over-the-top product-voice-y thing imaginable. It goes so far in the direction of “aiming for a non-technical audience” that it ends up seemingly addressed to small children. As in a picture book, it offers simple sentences in gigantic letters, one or two per page, each one nestled between huge comforting expanses of blank space. And sentences themselves are . well, st...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Lies Told To Children, published by Eliezer Yudkowsky on April 14, 2022 on LessWrong. Growing up, as a kid, I was always told that every sapient life is precious, everything that thinks and knows itself - Yes, this is a story about lies-told-to-children. You'll probably figure it out yourself before too long. For now, just listen. Where was I? Right. As children, we were always told that every sapient life is precious. It was told to us by the teachers, and shown to us in children's television - though I saw less children's television than most children in our age cohort - children's TV was censored where I grew up, though, of course, I didn't find that out until much later - I see you're starting to guess under what sort of circumstances I grew up. Go ahead, write down the prediction if you want. Maybe you already see where this entire thing is headed. But you asked me for a story about the lies I was told as a child, and that's what you're getting. It's not my fault, if a lot of stories like that are predictable; people who lie to children have other things to optimize for than unpredictability. So where was I? Right. I grew up in a remote village of about three thousand people, the sort that's more hills than houses. Charming travel-pathways that cut through forests. Not everyone knows everyone, but you sure know somebody who knows anybody. Children's television in my region was censored, though of course they didn't tell us that as children. But the children's television that we saw had aliens and monsters and creatures of fantasy, with four legs or fourteen legs, three faces or no face at all, and all of them were treated by the television show as having lives that meant something. Sometimes in the children's show there were alien monsters who only thought their own kind of life was valuable, and then maybe you couldn't trade with them as friends, maybe they'd already lied to you once and you couldn't trust them enough to bargain with them, maybe you couldn't talk to them at all. But their lives still had meaning to the story's human protagonists, even some aliens whose lives had no meaning to themselves. You didn't cause them pain if there was any way to avoid it; you didn't kill them unless their biology was sufficiently similar to human that you were confident in your ability to cryopreserve them afterwards. The shows never spelled it out, never said, 'And this is because of a universal rule in every case that sapient life has value.' Our teachers said that explicitly, though. And they treated every one of us children, too, as if our lives had meaning. Except the children with the red hair; those dirty reds. You're nodding along with a knowing look, I see. Was it what you predicted? Not exactly, maybe, but rough ballpark? I suppose I'll find out when we open your prediction afterwards. The red-haired children hardly needed the red hair, as their targeting-mark; they looked different from the rest of us in other ways too. When I was old enough to first ask, I was told that they were the children's children of people who'd been exiled from a faraway city for committing terrible crimes there, who'd been given sanctuary by the grace and mercy of our own benevolent kind. The red-haired children tended bigger than the rest of us, with more adult facial structures, to the point where you could've maybe mistaken them for very small adults in disguise. The red-haired adults, what few of them we ever saw, were correspondingly huge and muscular. You could see, in retrospect - if you were actually trying to think at all, which we weren't really - how somebody might have felt threatened by such big muscular people, even while graciously granting them sanctuary. There weren't many of the red-haired children being educated alongside us; a handful, four or six. I can't recal...