The Nonlinear Library allows you to easily listen to top EA and rationalist content on your podcast player. We use text-to-speech software to create an automatically updating repository of audio content from the EA Forum, Alignment Forum, LessWrong, and other EA blogs. To find out more, please visit us at nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Highlights: Wentworth, Shah, and Murphy on "Retargeting the Search", published by RobertM on September 14, 2023 on LessWrong.In How To Go From Interpretability To Alignment: Just Retarget The Search, John Wentworth suggests:When people talk about prosaic alignment proposals, there's a common pattern: they'll be outlining some overcomplicated scheme, and then they'll say "oh, and assume we have great interpretability tools, this whole thing just works way better the better the interpretability tools are", and then they'll go back to the overcomplicated scheme. (Credit to Evan for pointing out this pattern to me.) And then usually there's a whole discussion about the specific problems with the overcomplicated scheme.In this post I want to argue from a different direction: if we had great interpretability tools, we could just use those to align an AI directly, and skip the overcomplicated schemes. I'll call the strategy "Just Retarget the Search".We'll need to make two assumptions:Some version of the natural abstraction hypothesis holds, and the AI ends up with an internal concept for human values, or corrigibility, or what the user intends, or human mimicry, or some other outer alignment target.The standard mesa-optimization argument from Risks From Learned Optimization holds, and the system ends up developing a general-purpose (i.e. retargetable) internal search process.Given these two assumptions, here's how to use interpretability tools to align the AI:Identify the AI's internal concept corresponding to whatever alignment target we want to use (e.g. values/corrigibility/user intention/human mimicry/etc).Identify the retargetable internal search process.Retarget (i.e. directly rewire/set the input state of) the internal search process on the internal representation of our alignment target.Just retarget the search. Bada-bing, bada-boom.There was a pretty interesting thread in the comments afterwards that I wanted to highlight.Rohin Shah (permalink)Definitely agree that "Retarget the Search" is an interesting baseline alignment method you should be considering.I like what you call "complicated schemes" over "retarget the search" for two main reasons:They don't rely on the "mesa-optimizer assumption" that the model is performing retargetable search (which I think will probably be false in the systems we care about).They degrade gracefully with worse interpretability tools, e.g. in debate, even if the debaters can only credibly make claims about whether particular neurons are activated, they can still stay stuff like "look my opponent is thinking about synthesizing pathogens, probably it is hoping to execute a treacherous turn", whereas "Retarget the Search" can't use this weaker interpretability at all. (Depending on background assumptions you might think this doesn't reduce x-risk at all; that could also be a crux.)johnswentworth (permalink)I indeed think those are the relevant cruxes.Evan R. Murphy (permalink)They don't rely on the "mesa-optimizer assumption" that the model is performing retargetable search (which I think will probably be false in the systems we care about).Why do you think we probably won't end up with mesa-optimizers in the systems we care about?Curious about both which systems you think we'll care about (e.g. generative models, RL-based agents, etc.) and why you don't think mesa-optimization is a likely emergent property for very scaled-up ML models.Rohin Shah (permalink)It's a very specific claim about how intelligence works, so gets a low prior, from which I don't update much (because it seems to me we know very little about how intelligence works structurally and the arguments given in favor seem like relatively weak considerations).Search is computationally inefficient relative to heuristics, and we'll be selecting rea...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: UDT shows that decision theory is more puzzling than ever, published by Wei Dai on September 13, 2023 on LessWrong.I feel like MIRI perhaps mispositioned FDT (their variant of UDT) as a clear advancement in decision theory, whereas maybe they could have attracted more attention/interest from academic philosophy if the framing was instead that the UDT line of thinking shows that decision theory is just more deeply puzzling than anyone had previously realized. Instead of one major open problem (Newcomb's, or EDT vs CDT) now we have a whole bunch more. I'm really not sure at this point whether UDT is even on the right track, but it does seem clear that there are some thorny issues in decision theory that not many people were previously thinking about:Indexical values are not reflectively consistent. UDT "solves" this problem by implicitly assuming (via the type signature of its utility function) that the agent doesn't have indexical values. But humans seemingly do have indexical values, so what to do about that?The commitment races problem extends into logical time, and it's not clear how to make the most obvious idea of logical updatelessness work.UDT says that what we normally think of different approaches to anthropic reasoning are really different preferences, which seems to sidestep the problem. But is that actually right, and if so where are these preferences supposed to come from?2TDT-1CDT - If there's a population of mostly TDT/UDT agents and few CDT agents (and nobody knows who the CDT agents are) and they're randomly paired up to play one-shot PD, then the CDT agents do better. What does this imply?Game theory under the UDT line of thinking is generally more confusing than anything CDT agents have to deal with.UDT assumes that the agent has access to its own source code and inputs as symbol strings, so it can potentially reason about logical correlations between its own decisions and other agents' as well defined mathematical problems. But humans don't have this, so how are humans supposed to reason about such correlations?Logical conditionals vs counterfactuals, how should these be defined and do the definitions actually lead to reasonable decisions when plugged into logical decision theory?These are just the major problems that I was trying to solve (or hoping for others to solve) before I mostly stopped working on decision theory and switched my attention to metaphilosophy. (It's been a while so I'm not certain the list is complete.) As far as I know nobody has found definitive solutions to any of these problems yet, and most are wide open.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: PSA: The community is in Berkeley/Oakland, not "the Bay Area", published by maia on September 11, 2023 on LessWrong.Posting this because I recently had a conversation that went like this:Friend: Hey, you used to live in SF. Is there any rationalist stuff actually happening in San Francisco? There don't seem to be many events, or even that many aspiring rationalists living here. What's up with that? [Paraphrased. I've had similar versions of this conversation more than once.]Me: Something we realized living there is that SF actually suffers the same brain drain as most other cities, because everyone just goes to Berkeley/Oakland.The same way people move from the East Coast or elsewhere to Berkeley, they move from the rest of the Bay Area to Berkeley. Actually, they do it even more, because moving to Berkeley is easier when you already live pretty close by.And you don't figure this out until you move there, because people who live outside the Bay Area think of it as being all the same place. But the 45 minute train ride really matters when it comes to events and socializing, as it turns out.Friend: That sounds so inconvenient for people who have jobs in the city or South Bay!Me: Sure is! I don't have a super-solid answer for this, except that 1) Lots of people actually just do awful, awful commutes, because having a real, in-person community is that valuable to them, as bad as commuting is. 2) A surprising fraction of the community works at rationalist/rationalist-adjacent nonprofits, most of which are actually located in the East Bay. Plus, 3) in a post-COVID world, more people can work remote or partly remote. So you can choose to live where your community is... which is Berkeley... even though it is crazy expensive.I don't actually live in the Bay Area anymore, so I don't have the most up-to-date information on where events are happening and things. But it seems from what I hear from folks still there that it's still broadly true that East Bay is where things are happening, and other parts of the area have much less of the community.If you're thinking about moving to the Bay in part for the rationality community, take this into account!Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: US presidents discuss AI alignment agendas, published by TurnTrout on September 9, 2023 on LessWrong.None of the presidents fully represent my (TurnTrout's) views.TurnTrout wrote the script. Garrett Baker helped produce the video after the audio was complete. Thanks to David Udell, Ulisse Mini, Noemi Chulo, and especially Rio Popper for feedback and assistance in writing the script.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Sum-threshold attacks, published by TsviBT on September 8, 2023 on LessWrong.How do you affect something far away, a lot, without anyone noticing?(Note: you can safely skip sections. It is also safe to skip the essay entirely, or to read the whole thing backwards if you like.)The frog's lawsuitAttorney for the defendant: "So, Mr. Frog. You allege that my client caused you grievous bodily harm. How is it that you claim he harmed you?"Frog: "Ribbit RIBbit ribbit."Attorney: "Sir..."Frog: "Just kidding. Well, I've been living in a pan for the past two years. When I started, I was the picture of health, and at first everything was fine. But over the course of the last six months, something changed. By last month, I was in the frog hospital with life-threatening third-degree burns."Attorney: "And could you repeat what you told the jury about the role my client is alleged to have played in your emerging medical problems?"Frog: "Like I said, I don't know exactly. But I know that when my owner wasn't away on business, every day he'd do something with the stove my pan was sitting on. And then my home would seem to be a bit hotter, always a bit hotter."Attorney: "Your owner? You mean to say..."Judge: "Let the record show that Mr. Frog is extending his tongue, indicating the defendant, Mr. Di'Alturner."Attorney: "Let me ask you this, Mr. Frog. Is it right to say that my client - - your owner - - lives in an area with reasonably varied weather? It's not uncommon for the temperature to vary by ten degrees over the course of the day?"Frog: "True."Attorney: "And does my client leave windows open in his house?"Frog: "He does."Attorney: "So I wonder, how is it that you can tell that a slight raise in temperature that you experience - - small, by your own admission - - how can you be sure that it's due to my client operating his stove, and not due to normal fluctuations in the ambient air temperature?"Frog: "I can tell because of the correlation. I tend to feel a slight warming after he's twiddled the dial."Attorney: "Let me rephrase my question. Is there any single instance you can point to, where you can be sure - - beyond a reasonable doubt - - that the warming was due to my client's actions?"Frog: "Ah, um, it's not that I'm sure that any one increase in temperature is because he turned the dial, but..."Attorney: "Thank you. And would it be fair to say that you have no professional training in discerning temperature and changes thereof?"Frog: "That would be accurate."Attorney: "And are you aware that 30% of frogs in your state report spontaneous slight temperature changes at least once a month?"Frog: "But this wasn't once a month, it was every day for weeks at a ti - - "Attorney: "Sir, please only answer the questions I ask you. Were you aware of that fact?"Frog: "No, I wasn't aware of that, but I don't see wh - - "Attorney: "Thank you. Now, you claim that you were harmed by my client's actions, which somehow put you into a situation where you became injured."Frog: "¡I have third degree burns all ov - - "Attorney: "Yes, we've seen the exhibits, but I'll remind you to only speak in response to a question I ask you. What I'd like to ask you is this: Why didn't you just leave the frying pan? If you were, as you allege, being grievously injured, wasn't that enough reason for you to remove yourself from that situation?"Frog: "I, I didn't notice that it was happening at the time, each change was so subtle, but..."Attorney: "Thank you. As your counsel would have advised you, the standard for grievous bodily harm requires intent. Now are we really expected to conclude, beyond a reasonable doubt, that my client intended to cause you harm, via a method that you didn't even notice? That even though you can't point to so much as a single instance where my ...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Sharing Information About Nonlinear, published by Ben Pace on September 7, 2023 on LessWrong.Epistemic status: Once I started actively looking into things, much of my information in the post below came about by a search for negative information about the Nonlinear cofounders, not from a search to give a balanced picture of its overall costs and benefits. I think standard update rules suggest not that you ignore the information, but you think about how bad you expect the information would be if I selected for the worst, credible info I could share, and then update based on how much worse (or better) it is than you expect I could produce. (See section 5 of this post about Mistakes with Conservation of Expected Evidence for more on this.) This seems like a worthwhile exercise for at least non-zero people to do in the comments before reading on. (You can condition on me finding enough to be worth sharing, but also note that I think I have a relatively low bar for publicly sharing critical info about folks in the EA/x-risk/rationalist/etc ecosystem.)tl;dr: If you want my important updates quickly summarized in four claims-plus-probabilities, jump to the section near the bottom titled "Summary of My Epistemic State".When I used to manage the Lightcone Offices, I spent a fair amount of time and effort on gatekeeping - processing applications from people in the EA/x-risk/rationalist ecosystem to visit and work from the offices, and making decisions. Typically this would involve reading some of their public writings, and reaching out to a couple of their references that I trusted and asking for information about them. A lot of the people I reached out to were surprisingly great at giving honest references about their experiences with someone and sharing what they thought about someone.One time, Kat Woods and Drew Spartz from Nonlinear applied to visit. I didn't know them or their work well, except from a few brief interactions that Kat Woods seems high-energy, and to have a more optimistic outlook on life and work than most people I encounter.I reached out to some references Kat listed, which were positive to strongly positive. However I also got a strongly negative reference - someone else who I informed about the decision told me they knew former employees who felt taken advantage of around things like salary. However the former employees reportedly didn't want to come forward due to fear of retaliation and generally wanting to get away from the whole thing, and the reports felt very vague and hard for me to concretely visualize, but nonetheless the person strongly recommended against inviting Kat and Drew.I didn't feel like this was a strong enough reason to bar someone from a space - or rather, I did, but vague anonymous descriptions of very bad behavior being sufficient to ban someone is a system that can be straightforwardly abused, so I don't want to use such a system. Furthermore, I was interested in getting my own read on Kat Woods from a short visit - she had only asked to visit for a week. So I accepted, though I informed her that this weighed on my mind. (This is a link to the decision email I sent to her.)(After making that decision I was also linked to this ominous yet still vague EA Forum thread, that includes a former coworker of Kat Woods saying they did not like working with her, more comments like the one I received above, and links to a lot of strongly negative Glassdoor reviews for Nonlinear Cofounder Emerson Spartz's former company "Dose". Note that more than half of the negative reviews are for the company after Emerson sold it, but this is a concerning one from 2015 (while Emerson Spartz was CEO/Cofounder): "All of these super positive reviews are being commissioned by upper management. That is the first thing you should know about Spartz, and I...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Find Hot French Food Near Me: A Follow-up, published by aphyer on September 6, 2023 on LessWrong.On Zvi's recent post about French food I posted an inflammatory comment (saying in essence that French food is so bad American capitalism hasn't even bothered stealing it). I got challenged to provide evidence supporting this, and particularly to back up my claim that there were more German than French restaurants near me.Right. Yes. Evidence. I am a reasonable adult who understands that beliefs must be supported by evidence. So. Here we go.Some Google SearchesI've searched for '[ethnicity] restaurant near Grove Street, Jersey City, NJ' (I live in Jersey City, and the Grove Street area is reasonably near the center).When I search for 'French' I can count 13 results:And when I search for 'German' I count only 9:Ha! The foolish American has been hoisted on his own petard! ('Petard' is French for 'fuck you').Perhaps unsurprisingly, I don't think these numbers tell the whole story.What Makes These Places French?Google's definition of 'French' and 'German' restaurants here appears to be extremely expansive.Hudson Hound Jersey City, an 'Irish gastropub', shows up on the French search.Shadman, a 'go-to for Pakistani and Indian cuisine', shows up on the German search.Luna, for 'Italian eats', shows up on the French search.Frankie, an 'Australian eatery', shows up on the German search.So, for lack of anything better to do, I've gone through manually to look for things that I think 'count' as French or German.The two 'real' German places (and the ones I was thinking of in my comment) are 'Wurstbar' and 'Zeppelin Hall Beer Garden', and while we may question the taste of these places I do not think we can question their German-ness. The search also turned up 'Hudson Hall', a 'Euro beer bar with house-smoked meats', which I think at least ambiguously might count.It's less clear to me how many of the hits for 'French restaurant' are actually both French and restaurants. Certainly I've been to a few of these places, and none of them have charged me twenty-three dollars for a baguette while sneering at me. We have:Cafe Madelaine describes itself as a French restaurant. We count that.Choc O Pain definitely sounds French, but it's not clear to me if it's actually a restaurant: it seems to actually be a bakery, and the menu seems to bear that out. I'll give it half.Hudson Hound self-describes as 'Irish'.Matthews Food and Drink self-describes as 'American' (though I guess it also self-describes as 'chic').Grove Station self-describes as 'New American' (I have no idea what that means).El Sazon De Las Americas self-describes as 'Dominican' (I don't think that counts as French, though I'm sure someone will make the case).Uncle Momo self-describes as 'French-Lebanese fare'. Let's give that half again.Beechwood Cafe self-describes as 'American'.Luna self-describes as 'Italian'.Razza is an Italian pizza place.Short Grain is...uh...a 'hip place with sidewalk seats serving Asian-influenced & vegetarian dishes, plus coffee & green tea', and while I have no idea what that is and don't particularly want to find out I don't think it means 'French'.Frankie self-describes as 'Italian'.Cafe Dolma self-describes as 'Greek'.So overall I think 'French' and 'German' each end up with either 2 or 3 restaurants, depending on how you count some edge cases.SummaryI am sorry that I said French food was not as successful under capitalism as German food. I see now that French food is exactly as popular and successful as German food, and I'll fight anyone who says otherwise!Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Text Posts from the Kids Group: 2023 I, published by jefftk on September 5, 2023 on LessWrong.We have a Facebook group for kid stuff, because if we post a mixture of kid things and other stuff FB's algorithm gets very confused about who to show our posts to. While my annualpictures posts mostly cover the visual side, the text posts are only on FB and I don't like that. So: here's the first ~half of 2023.(Some of these were from me; some were from Julia. Ones saying "me" could mean either of us.)Anna: I thought a blue heron was a bird with blue hair that was in?Lily: I've figured out that if you tell grown-ups something is healthy, they're more likely to get it.Lily: [Confined to her room with covid] Could you refill my water cup?Me: Sure! [Gets cup][Fills cup. Starts doing something else.]Lily: [Over walkie-talkie] I'm having trouble remembering where I put my water cup, have you seen it?Me: [trying not to laugh] Sorry, I forgot to bring it back up!Lily: Your voice sounds funny, are you ok?Me: I was trying not to laugh. Had you actually forgotten or were you being polite?Lily: Mostly being polite; did I do something funny?Me: Yes, I mean no, I mean I didn't that approach was something you knew how to do yet.Lily: Thanks, I guess?(Worrying when your 8yo is better at social stuff than you are.)Anna: dad, I'm really cold.Me: how about a sweater?Anna: I can't find any of my sweaters.Me: have your looked in your drawer?Anna: I don't want to go upstairs!Anna: Nora, should Lily... not be allowed to play in the fort?Nora: ???Anna: Is that true?Nora: Yeah!Anna: See Lily, you have to get out!Lily: But Nora says yes to everything!Me: I'm worried you're going to jump on me in a way that hurts.Anna: No, I'm only going to jump on the blanketMe: Yes, but I'm under the blanket!Anna: I don't like it when someone wins and I'm not the person who winsThings Nora is really into right now:Balls, or other round things that could plausibly be consideredballs (M&Ms, the globe)Shutting the dishwasher doorAnimals that roar, especially lions, but also bears, tigers, andother animals that she thinks might roar (monkeys, wombats, cows). There's a house near us with concrete lion statues outfront, and she likes to go roar at them.Anna: In the story the king got happier and happier as he gave away his things, but that isn't how it is for me. The problem is I get sadder and sadder as I give away things because I like most things. I just really really like things!Anna: I'm always ready for a challenge that's not at all hardLily: I'm at an age when I get bored easilyAnna: I'm at an age where I don't get bored easily, especially whenI'm eating cakeAnna: "I was standing on the coffee table watching my fish, and then I started to walk away. I forgot I was on the table and hurt my knee when I fell."She was fine in a minute. I'm not sure what she hurt more: her knee or her pride.Me, a month after getting Anna black socks instead of white ones: Anna, where are you putting your socks when they're dirty?Anna: They don't get dirty.Nora really likes ice cream, and signs for it hopefully at many opportunities. Today, when Erika said no ice cream she started alternating between signing it and saying "Papa". I think as in "Papa let's me have it!"I was just telling this to Julia, and because Nora was present I spelled out "i c e c r e a m". Nora immediately started signing "ice cream".Still hard to distinguish from her base rate of signing "ice cream" at people.You know how you can get more food in a burrito at Chipotle by asking for all the fillings?Anna: "I want an ice cream sundae with double chocolate brownie batter ice cream, whipped cream, chocolate sauce, caramel sauce, a piece of popsicle, and a piece of the donut."Lily: Anna! You're taking all the gems!...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Defunding My Mistake, published by ymeskhout on September 4, 2023 on LessWrong.Confessions of an ex-ACABUntil about five years ago, I unironically parroted the slogan All Cops Are Bastards (ACAB) and earnestly advocated to abolish the police and prison system. I had faint inklings I might be wrong about this a long time ago, but it took a while to come to terms with its disavowal. What follows is intended to be not just a detailed account of what I used to believe but most pertinently, why. Despite being super egotistical, for whatever reason I do not experience an aversion to openly admitting mistakes I've made, and I find it very difficult to understand why others do. I've said many times before that nothing engenders someone's credibility more than when they admit error, so you definitely have my permission to view this kind of confession as a self-serving exercise (it is). Beyond my own penitence, I find it very helpful when folks engage in introspective, epistemological self-scrutiny, and I hope others are inspired to do the same.How Did I Get There?For decades now, I've consistently held plain vanilla libertarian policy preferences, with the only major distinction being that I've aligned myself more with the anarchists. Whereas some were content with pushing the "amount of government" lever to "little", I wanted to kick it all the way to "zero". There are many reasons I was and remain drawn to anarchist libertarianism, and chief among them was the attractively simple notion that violence is immoral and that government is violence. The problem with moral frameworks is that they can be quite infectious. To pick on one example for demonstration's sake, I notice that for many animal welfare advocates a vegan diet is heralded not just as the ideal moral choice, but also as the healthiest for humans, the least polluting, the cheapest financially, the best for soil conservation, the most water-efficient, the least labor-exploitative, et cetera & so forth.There's a risk that if you become dogmatically attached to a principled position, you're liable to be less scrutinizing when reflexively folding in other justifications. I suspect that happened to me with prisons, for example, where because I felt immediate revulsion at the thought of the state forcing someone into a cage, I was unwilling to entertain the possibility it could be justified. Ceding the ground on this particular brick was too threatening to the anarchism edifice I was so fond of.Obviously if you advocate getting rid of the government, people naturally want to know what will replace it. Some concerns were trivial to respond to (I'm not sad about the DEA not existing anymore because drugs shouldn't be illegal to begin with), but other questions I found annoying because I admittedly had no good answer, such as what to do with criminals if the police didn't exist. I tried to find these answers. Anarchism as an umbrella ideology leans heavily to the far left and has a history of serious disagreements with fellow-travelers in Marxism. Despite that feud, anarchist thought absorbed by proxy Marxist "material conditions" critiques that blame the existence of crime on capitalism's inequalities - a claim that continues to be widely circulated today, despite how flagrantly dumb it is. As someone who was and continues to be solidly in favor of free market economics, these critiques were like parsing an inscrutable foreign language.I was in college around my most ideologically formative time and a voracious reader, but I churned through the relevant literature and found nothing convincing. Instead of noting that as a blaring red flag, I maintained the grip I had on my preferred conclusion and delegated the hard work of actually defending it to someone else. I specifically recall how Angela Davis's 2003 book Are...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The goal of physics, published by Jim Pivarski on September 3, 2023 on LessWrong.In grad school, I was a teaching assistant for a course called, Why the Sky is Blue. It was a qualitative introduction to physics for non-majors, covering a lot of the same topics as Physics I, such as forces, conservation of energy and momentum, electric charges and magnetic fields, in less detail, with not much math. The actual question about why the sky is blue was saved for the end. As the course dragged on and the students (who expected no math, rather than not much math) started to complain, "Are we ever going to find out why the sky is blue?" I watched the schedule slip and wondered the same thing.We skipped some sections and managed to wedge it into the last lecture: finally, we were talking about why the sky is blue! "The sky is blue because of Rayleigh scattering." Okay, that's not an answer we hadn't defined Rayleigh scattering, there wasn't time for it, so we said that air molecules absorb and re-radiate - effectively changing the direction of - blue light more than red light. Red light goes straight through the atmosphere, and blue light bounces around, making the whole sky glow blue. Conversely, sunrises and sunsets are red because you're looking at the light that has gone straight through a larger wedge of atmosphere. It lost most of its blue on the way to your eye.Pretty good explanation, for not being able to say(the 1/λ4 part affects small-λ blue light more than large-λ red light). We also showed pictures like this sunset:to demonstrate the effect of straight-through red light and bouncing-around blue light.So in the end, "Why is the sky blue?"Answer: "Because sunsets are red!""And why are sunsets red...?"It was understandably unsatisfying. One thing was only explained in terms of another thing. But even if we had the time to get into detail about Rayleigh scattering, they could reasonably ask, "Why does light scatter according to that formula?" We could go deeper and explain Lord Rayleigh's proof in terms of Maxwell's equations. And whyfore Maxwell's equations? Well, quantum electrodynamics, which is a quantum field theory with a local U(1) gauge symmetry, which is to say that every point in space has an extra degree of freedom, similar to a fourth spatial dimension except that this dimension can't be rotated with normal space like the other three, this dimension is connected to itself as a circle instead of being infinite (that's what the U(1) means), and neighboring points in 3D space try to minimize differences in this extra parameter, which leads to waves.The explanatory power is breathtaking: you can actually derive that photons must exist, if you assume that there's this U(1) symmetry laying around. But why is there a U(1) symmetry?Modern physics seems to be obsessed with symmetries. Even this U(1) symmetry is explained in terms of a more fundamental SU(2)ÃU(1) (different U(1)) and the Higgs mechanism. Physicists seem to be holding their tongues, avoiding saying, "This is the most basic thing," by saying, "This one thing is actually a manifestation that other thing." Answering the question, "Why do photons exist?" with "Because space has an internal U(1) symmetry" is a bit like saying, "The sky is blue because sunsets are red."Symmetry explanations collapse our description of the world onto a smaller description. They say that one thing is mathematically derivable from the other, maybe in both directions, but they don't say why either is there at all. Perhaps that's an unanswerable question, and the symmetry language is a way of acknowledging the limitation.To show what I mean, consider a universe that consists of nothing but a point in the exact center of a perfect circle. (I've been feeling free to say, "Consider a universe..." ever since a lecture...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The smallest possible button, published by Neil on September 2, 2023 on LessWrong.tl;dr: The more knowledge you have, the smaller the button you need to press to achieve desired results. This is what makes moth traps formidable killing machines, and it's a good analogy for other formidable killing machines I could mention.TrapsI was shopping for moth traps earlier today, and it struck me how ruthlessly efficient humans could be in designing their killing apparatus. The weapon in question was a thin pack in my hands containing just a single strip of paper which, when coated with a particular substance and folded in the right way, would end up killing most of the moths in my house. No need to physically hunt them down or even pay remote attention to them myself; a couple bucks spent on this paper and a minute to set it up, and three quarters of the entire population is decimated in less than a day.That's. horrifying.Moth traps are made from cardboard coated with glue and female moth pheromones. Adult males are attracted to the pheromones, and end up getting stuck to the sides where they end up dying. The females live, but without the males, no new larvae are born and in a few months time you've wiped out a whole generation of moths. These traps are "highly sensitive" meaning that they will comb a whole room of moths very quickly despite being passive in nature.Why are moth traps so effective? They use surgically precise knowledge. Humans know how to synthesize moth pheromones, and from there you can hack a 250-million-year-old genetically derived instinct that male moths have developed for mating, and then you set a trap and voilà . The genetic heuristic that worked 99% of the time for boosting reproductive rates in moths can be wielded against moths by obliterating their reproductive rates.Moth traps aren't even the pinnacle of human insecticidal war machines. Scientists have, after all, seriously considered using gene drives to eliminate an entire species of mosquitoes with a single swarm and some CRISPy cleverness.The smallest buttonMoth traps and gene drives work by understanding something so well that when you use brute force (because everything is brute force) to do something, you do it in the most optimal and surgical way. Intelligent design means humans can engineer very, very effective traps that harness the smallest buttons you can push in order to get a desired result.Evolution can also produce sexually deceptive traps that take advantage of insect brains. This is because genes that contribute to pushing a particular button that makes reproduction more likely, are more represented in the environment, so most genes in living beings today are already vetted for their capacity to harness niche buttons in the universe.The blind idiot god can't hope to compete with intelligent design however, so we can expect humans to win the find-the-smallest-button arms race against their evolution-derived enemies (like moths, mosquitoes, or viruses).Brute forceBrute force always works. If you stuff enough moths into my house, my measly passive traps won't be sufficient. In fact, if my house were big enough and there were enough moths, the males that were somehow not attracted to my sticky female pheromones but found females anyway would be the only ones to pass down their genes. With enough moths and enough time, the blind idiot god of moth evolution would find a way to elude my traps by pressing an alternate small button to those specific pheromones, in order to power its reproduction. This type of brute force, which grants a stupid and blind enemy the power of adaptation, can be found in battles with cancer, viruses, or pesticides.The only counter to this brute force is more brute force, in the form of chemotherapy, gene drives, or pesticides 1 level of magnitu...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: A Golden Age of Building? Excerpts and lessons from Empire State, Pentagon, Skunk Works and SpaceX, published by jacobjacob on September 1, 2023 on LessWrong.Patrick Collison has a fantastic list of examples of people quickly accomplishing ambitious things together since the 19th Century. It does make you yearn for a time that feels... different, when the lethargic behemoths of government departments could move at the speed of a racing startup:[...] last century, [the Department of Defense] innovated at a speed that puts modern Silicon Valley startups to shame: the Pentagon was built in only 16 months (1941-1943), the Manhattan Project ran for just over 3 years (1942-1946), and the Apollo Program put a man on the moon in under a decade (1961-1969). In the 1950s alone, the United States built five generations of fighter jets, three generations of manned bombers, two classes of aircraft carriers, submarine-launched ballistic missiles, and nuclear-powered attack submarines.[Note: that paragraph is from a different post.]Inspired by partly by Patrick's list, I spent some of my vacation reading and learning about various projects from this Lost Age. I then wrote up a memo to share highlights and excerpts with my colleagues at Lightcone.After that, some people encouraged me to share the memo more widely -- and I do think it's of interest to anyone who harbors an ambition for greatness and a curiosity about operating effectively.How do you build the world's tallest building in only a year? The world's largest building in the same amount of time? Or America's first fighter jet in just 6 months?How??Writing this post felt like it helped me gain at least some pieces of this puzzle. If anyone has additional pieces, I'd love to hear them in the comments.Empire State BuildingThe Empire State was the tallest building in the world upon completion in April 1931. Over my vacation I read a rediscovered 1930s notebook, written by the general contractors themselves. It details the construction process and the organisation of the project.I will share some excerpts, but to contextualize them, consider first some other skyscrapers built more recently:Design startConstruction endTotal timeBurj Khalifa200420106 yearsShanghai Tower200820157 yearsAbraj Al-Balt2002201210 yearsOne World Trade Center200520149 yearsNordstrom Tower2010202010 yearsTaipei 101199720047 years(list from skyscrapercenter.com)Now, from the Empire State book's foreword:The most astonishing statistics of the Empire State was the extraordinary speed with which it was planned and constructed. [...] There are different ways to describe this feat. Six months after the setting of the first structural columns on April 7, 1930, the steel frame topped off on the eighty-sixth floor. The fully enclosed building, including the mooring mast that raised its height to the equivalent of 102 stories, was finished in eleven months, in March 1931. Most amazing though, is the fact that within just twenty months -- from the first signed contractors with the architects in September 1929 to opening-day ceremonies on May 1, 1931 -- the Empire State was designed, engineered, erected, and ready for tenants.Within this time, the architectural drawings and plans were prepared, the Vicitorian pile of the Waldorf-Astoria hotel was demolished [demolition started only two days after the initial agreement was signed], the foundations and grillages were dug and set, the steel columns and beams, some 57,000 tons, were fabricated and milled to precise specifications, ten million common bricks were laid, more than 62,000 cubic yards of concrete were poured, 6,400 windows were set, and sixty-seven elevators were installed in seven miles of shafts. At peak activity, 3,500 workers were employed on site, and the frame rose more than a story a day,...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Responses to apparent rationalist confusions about game / decision theory, published by Anthony DiGiovanni on August 31, 2023 on LessWrong.I've encountered various claims about how AIs would approach game theory and decision theory that seem pretty importantly mistaken. Some of these confusions probably aren't that big a deal on their own, and I'm definitely not the first to point out several of these, even publicly. But collectively I think these add up to a common worldview that underestimates the value of technical work to reduce risks of AGI conflict. I expect that smart agents will likely avoid catastrophic conflict overall - it's just that the specific arguments for expecting this that I'm responding to here aren't compelling (and seem overconfident).For each section, I include in the footnotes some examples of the claims I'm pushing back on (or note whether I've primarily seen these claims in personal communication). This is not to call out those particular authors; in each case, they're saying something that seems to be a relatively common meme in this community.Summary:The fact that conflict is costly for all the agents involved in the conflict, ex post, doesn't itself imply AGIs won't end up in conflict. Under their uncertainty about each other, agents with sufficiently extreme preferences or priors might find the risk of conflict worth it ex ante. (more)Solutions to collective action problems, where agents agree on a Pareto-optimal outcome they'd take if they coordinated to do so, don't necessarily solve bargaining problems, where agents may insist on different Pareto-optimal outcomes. (more)We don't have strong reasons to expect AGIs to converge on sufficiently similar decision procedures for bargaining, such that they coordinate on fair demands despite committing under uncertainty. Existing proposals for mitigating conflict given incompatible demands, while promising, face some problems with incentives and commitment credibility. (more)The commitment races problem is not just about AIs making commitments that fail to account for basic contingencies. Updatelessness (or conditional commitments generally) seems to solve the latter, but it doesn't remove agents' incentives to limit how much their decisions depend on each other's decisions (leading to incompatible demands). (more)AIs don't need to follow acausal decision theories in order to (causally) cooperate via conditioning on each other's source code. (more)Most supposed examples of Newcomblike problems in everyday life don't seem to actually be Newcomblike, once we account for "screening off" by certain information, per the Tickle Defense. (more)The fact that following acausal decision theories maximizes expected utility with respect to conditional probabilities, or counterfactuals with the possibility of logical causation, doesn't imply that agents with acausal decision theories are selected for (e.g., acquire more material resources). (more)Ex post optimal =/= ex ante optimalAn "ex post optimal" strategy is one that in fact makes an agent better off than the alternatives, while an "ex ante optimal" strategy is optimal with respect to the agent's uncertainty at the time they choose that strategy. The idea that very smart AGIs could get into conflicts seems intuitively implausible because conflict is, by definition, ex post Pareto-suboptimal. (See the "inefficiency puzzle of war.")But it doesn't follow that the best strategies available to AGIs given their uncertainty about each other will always be ex post Pareto-optimal. This may sound obvious, but my experience with seeing people's reactions to the problem of AGI conflict suggests that many of them haven't accounted for this important distinction.As this post discusses in more detail, there are two fundamental sources of uncertainty (o...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Biosecurity Culture, Computer Security Culture, published by jefftk on August 30, 2023 on LessWrong.While I've only worked in biosecurity for about a year and my computer security background consists of things I picked up while working on other aspects of software engineering, the cultures seem incredibly different. Some examples of good computer security culture that would be bad biosecurity culture:Openness and full disclosure. Write blog posts with deep detail on how vulnerabilities were found, with the goal of teaching others how to find similar ones in the future. Keep details quiet for a few months if need be to give vendors time to fix but after, say, 90 days go public.Breaking things to fix them. Given a new system, of course you should try to compromise it. If you succeed manually, make a demo that cracks it in milliseconds. Make (and publish!) fuzzers and other automated vulnerability search tools.Enthusiastic curiosity and exploration. Noticing hints of vulnerabilities and digging into them to figure out how deep they go is great. If someone says "you don't need to know that" ignore them and try to figure it out for yourself.This is not how computer security has always been, or how it is everywhere, and people in the field are often fiercely protective of these ideals against vendors that try to hide flaws or silence researchers. And overall my impression is that this culture has been tremendously positive in computer security.Which means that if you come into the effective altruism corner of biosecurity with a computer security background and see all of these discussions of "information hazards", people discouraging trying to find vulnerabilities, and people staying quiet about dangerous things they've discovered it's going to feel very strange, and potentially rotten.So here's a framing that might help see things from this biosecurity perspective. Imagine that the Morris worm never happened, nor Blaster, nor Samy. A few people independently discovered SQL injection but kept it to themselves. Computer security never developed as a field, even as more and more around us became automated. We have driverless cars, robosurgeons, and simple automated agents acting for us, all with the security of original Sendmail. And it's all been around long enough that the original authors have moved on and no one remembers how any of it works. Someone who put in some serious effort could cause immense distruction, but this doesn't happen because the people who have the expertise to cause havoc have better things to do. Introducing modern computer security culture into this hypothetical world would not go well!Most of the cultural differences trace back to what happens once a vulnerability is known. With computers:The companies responsible for software and hardware are in a position to fix their systems, and disclosure has helped build a norm that they should do this promptly.People who are writing software can make changes to their approach to avoid creating similar vulnerabilities in the future.End users have a wide range of effective and reasonably cheap options for mitigation once the vulnerability is known.But with biology there is no vendor, a specific fix can take years, a fully general fix may not be possible, and mitigation could be incredibly expensive. The culture each field needs is downstream from these key differences.Overall this is sad: we could move faster if we could all just talk about what we're most concerned about, plus cause prioritization would be simpler. I wish we were in a world where we could apply the norms from computer security! But different constraints lead to different solutions, and the level of caution I see in biorisk seems about right given these constraints.(Note that when I talk about "good biosecurity culture" I'm desc...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Introducing the Center for AI Policy (& we're hiring!), published by Thomas Larsen on August 28, 2023 on LessWrong.SummaryThe Center for AI Policy is a new organization designed to influence US policy to reduce existential and catastrophic risks from advanced AI.We are hiring for an AI Policy Analyst and a Communications Director. We're also open to other roles.What is CAIP?The Center for AI Policy (CAIP) is an advocacy organization that aims to develop and promote policies that reduce risks from advanced AI.Our current focus is building "stop button for AI" capacity in the US government. We have proposed legislation to establish a federal authority that engages in hardware monitoring, licensing for advanced AI systems, and strict liability for extreme model harms. Our proposed legislation also develops the ability to "press the button" - the federal authority would also monitor catastrophic risks from advanced AI development, inform congress and the executive branch about frontier AI progress, and have emergency powers to shut down frontier AI development in the case of a clear emergency. More detail can be found in the work section of our website.We also aim to broadly raise awareness about extreme risks from AI by engaging with policymakers in congress and the executive branch.How does CAIP differ from other AI governance organizations?Nature of the work: Many organizations are focused on developing ideas and amassing influence that can be used later. CAIP is focused on turning policy ideas into concrete legislative text and conducting advocacy now. We want to harness the current energy to pass meaningful legislation this policy window, in addition to building a coalition for the future. We are also being explicit about extinction risk with policy makers as the motivation behind our policy ideas.Worldview: We believe that in order to prevent an AI catastrophe, governments likely need to prevent unsafe AI development for multiple years, which requires they have secured computing resources, understand risks, and are prepared to shut projects down. Our regulation aims to build that capacity.Who works at CAIP?CAIP's team includes Thomas Larsen (CEO), Jason Green-Lowe (Legislative Director), and Jakub Kraus (COO). CAIP is also advised by experts from other organizations and is supported by many volunteers.How does CAIP receive funding?We received initial funding through Lightspeed Grants and private donors.We are currently funding constrained and think that donating to us is very impactful. You can donate to us here. If you are considering donating but would like to learn more, please message us at info@aipolicy.us.CAIP is hiringCAIP is looking for an AI Policy Analyst and a Communications Director. We are also open to applicants with different skills. If you would be excited to work at CAIP, but don't fit into these specific job descriptions, we encourage you to reach out to info@aipolicy.us directly.If you know someone who might be a good fit, please fill out this referral form.Note that we are actively fundraising, and the number of people we are able to recruit is currently uncertain.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Dear Self; we need to talk about ambition, published by Elizabeth on August 28, 2023 on LessWrong.I keep seeing advice on ambition, aimed at people in college or early in their career, that would have been really bad for me at similar ages. Rather than contribute (more) to the list of people giving poorly universalized advice on ambition, I have written a letter to the one person I know my advice is right for: myself in the past.The LetterDear Past Elizabeth,Your life is, in some sense, a series of definitions of success.First you're in early school, and success is defined for you by a handful of adults. You go where they say, do the assignments they say, when they say, and doing well means meeting the goals they set for you. Even your hippie elementary school gives you very few choices about life. You get choices in your leisure activity, but that (as they have explained to you) is leisure and thus unimportant, and there's no success or failure in it.Then you get further in school, and the authorities give you some choice over the hoops you jump through. You can choose which book you write your report on or even what classes you take (within a predetermined set). This feels like freedom, but you're in still a system someone else designed and set the win conditions for. You can fulfill a college distribution requirement with any history class at all- but you are going to take one, and the professor is the one determining if you succeeded at it.More insidiously, you'll like it. Creating your own definition of success feels scary;enacting it feels impossible. The fact that school lays out neat little hoops for you to jump through is a feature.Work (you'll be a programmer) is where things get screwy. Programming contains multiple definitions of success (manager, principal, freelancing, development, testing, bigtech, start-up, money-maxing, altruistic projects.), and multiple ways to go about them. If your goals lie outside of programming altogether (art, parenting, travel..), it's relatively easy to work out a way to fund it via programming while still having the time to do what you want. Not trivial, but have you seen what people in other jobs go through? With programming it's at least possible.But you like hoops. You're comfortable with hoops. So you're going to waste years chasing down various definitions of success within programming, and by the time you give up will be too exhausted to continue in it at all. I think you (I) should have considered "just chill while I figure shit out" much earlier, much more seriously. It was reasonable to give their way a try, just due to the sheer convenience if it had worked, but I should have learned faster.Eventually you will break out of the Seattle bigtech bubble, and into the overlapping bubbles of effective altruism, lesswrong, and the bay area start-up scene. All of three of these contain a lot of people shouting "be ambitious!" and "be independent!". And because they shout it so loudly and frequently you will think "surely, now I am in a wide open world and not on a path". But you will be wrong, because "be ambitious (in ways the people say this understand and respect)" and "be independent (in ways they think are cool and not crazy)" are still hoops and still determined by other people, just one more level meta.Like the programming path, the legible independent ambition path works for some people, but not you. The things you do when pushed to Think Big and Be Independent produce incidental learning at best, but never achieve anything directly. They can't, because you made up the goals to impress other people. This becomes increasingly depressing, as you fail at your alleged goals and at your real goal of impressing people.So what do we do then? Give up on having goals? Only by their definition. What seems to wo...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Aumann-agreement is common, published by tailcalled on August 27, 2023 on LessWrong.Thank you to Justis Mills for proofreading and feedback. This post is also available on my substack.Aumann's agreement theorem is a family of theorems which say that if people trust each other and know each other's opinions, then they agree with each other. Or phrased another way, if people maintain trust with each other, then they can reach agreement. (And some variants of the theorem, which take computational factors into consideration, suggest they can do so quite rapidly.)The original proof is pretty formal and confusing, but a simpler heuristic argument is that for an honest, rational agent, the mere fact of them professing an opinion can be strong evidence to another rational agent, because if the speaker's probabilities are higher than the speaker's prior, then they must have seen corresponding evidence to justify that opinion.Some people find this confusing, and feel like it must be wrong because it doesn't apply to most disagreements. I think these people are wrong because they are not sufficiently expansive in what they think of as a disagreement. The notion of disagreement that Aumann's agreement theorem applies to is when the people assign different probabilities to events; this is a quite inclusive notion which covers many things that we don't typically think of as disagreements, including cases where one party has information about a topic and the other party has no information.My vacation in Norway relied tons on Aumann agreementsRecently, I had a vacation in Norway with my wife.In order to get there, and to get around, we needed transport. At first we disagreed with people who provided transport there, as we didn't know of many specific means of transport, only vaguely that there would be some planes and ships, without knowing which ones. But my wife had heard that there was something called the "Oslo ferry", so we Aumann-agreed that this was an option, and decided to investigate further.We disagreed with the company that provided the Oslo ferry, as we didn't know what their website is, so we asked Google, and it provided some options for what the ferry might be, and we Aumann-agreed with Google and then went investigating from there. One website we found claimed to sell tickets to the ferry; at first we disagreed with the website about when we could travel as we didn't know the times of the ferry, but then we read which times it claimed was available, and Aumann-updated to that.We also had to find some things to do in Norway. Luckily for us, some people at OpenAI had noticed that everyone had huge disagreements with the internet as nobody had really memorized the internet, and they thought that they could gain some value by resolving that disagreement, so they Aumann-agreed with the internet by stuffing it into a neural network called ChatGPT. At first, ChatGPT disagreed with us about what to visit in Norway and suggested some things we were not really interested in, but we informed it about our interests, and then it quickly Aumann-agreed with us and proposed some other things that were more interesting.One of the things we visited was a museum for an adventurer who built a raft and sailed in the ocean. Prior to visiting the museum, we had numerous disagreements with it, as e.g. we didn't know that one of the people on the raft had fallen in the ocean and had to be rescued. But the museum told us this was the case, so we Aumann-agreed to believe it. Presumably, the museum learnt about it through Aumann-agreeing with the people on the raft.One example of an erroneous Aumann agreement was with the train company Vy. They had said that they could get us a train ticket on the Bergen train, and we had Aumann-agreed with that. However, due to a storm, their train...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Digital brains beat biological ones because diffusion is too slow, published by GeneSmith on August 26, 2023 on LessWrong.I've spent quite a bit of time thinking about the possibility of genetically enhancing humans to be smarter, healthier, more likely to care about others, and just generally better in ways that most people would recognize as such.As part of this research, I've often wondered whether biological systems could be competitive with digital systems in the long run.My framework for thinking about this involved making a list of differences between digital systems and biological ones and trying to weigh the benefits of each. But the more I've thought about this question, the more I've realized most of the advantages of digital systems over biological ones stem from one key weakness of the latter: they are bottlenecked by the speed of diffusion.I'll give a couple of examples to illustrate the point:To get oxygen into the bloodstream, the body passes air over a huge surface area in the lungs. Oxygen passively diffuses into the bloodstream through this surface where it binds to hemoglobin. The rate at which the body can absorb new oxygen and expel carbon dioxide waste is limited by the surface area of the lungs and the concentration gradient of both molecules.Communication between neurons relies on the diffusion of neurotransmitters across the synaptic cleft. This process takes approximately 0.5-1ms. This imposes a fundamental limit on the speed at which the brain can operate.A signal propogates down the axon of a neuron at about 100 meters per second. You might wonder why this is so much slower than a wire; after all, both are transmitting a signal using electric potential, right?It turns out the manner in which the electrical potential is transmitted is much different in a neuron. Signals are propagated down an axon via passive diffusion of Na+ ions into the axon via an Na+ channel. The signal speed is fundamentally limited by the speed at which sodium ions can diffuse into the cell. As a result, electrical signals travel through a wire about 2.7 million times faster than they travel through an axon.Delivery of energy (mainly ATP) to different parts of the cell occurs via diffusion. The fastest rate of diffusion I found of any molecule within a cell was that of positively charged hydrogen ions, which diffuse at a blistering speed of 0.007 meters/second. ATP diffuses much slower. So energy can be transferred through a wire at more than 38 billion times the speed that ATP can diffuse through a cell.Why hasn't evolution stumbled across a better method of doing things than passive diffusion?Here I am going to speculate. I think that evolution is basically stuck at a local maxima. Once diffusion provided a solution for "get information or energy from point A to point B", evolving a fundamentally different system requires a large number of changes, each of which individually makes the organism less well adapted to its environment.We can see examples of the difficulty of evolving fundamentally new abilities in Professor Richard Lenski's long-running evolution experiment using E. coli. which has been running since 1988. Lenski began growing E. coli in flasks full of a nutrient solution containing glucose, potassium phosphate, citrate, and a few other things.The only carbon source for these bacteria is glucose, which is limited. Once per day, a small portion of the bacteria in each flask is transferred to another flask, at which point they grow and multiply again.Each flask will contain a number of different strains of E. coli, all of which originate from a common ancestor.To measure the rate of evolution, Lenski and his colleagues measure the proportion of each strain. The ratio of one strain compared to the others gives a clear idea of its "fitness ad...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Assume Bad Faith, published by Zack M Davis on August 25, 2023 on LessWrong.I've been trying to avoid the terms "good faith" and "bad faith". I'm suspicious that most people who have picked up the phrase "bad faith" from hearing it used, don't actually know what it means - and maybe, that the thing it does mean doesn't carve reality at the joints.People get very touchy about bad faith accusations: they think that you should assume good faith, but that if you've determined someone is in bad faith, you shouldn't even be talking to them, that you need to exile them.What does "bad faith" mean, though? It doesn't mean "with ill intent." Following Wikipedia, bad faith is "a sustained form of deception which consists of entertaining or pretending to entertain one set of feelings while acting as if influenced by another." The great encyclopedia goes on to provide examples: the solider who waves a flag of surrender but then fires when the enemy comes out of their trenches, the attorney who prosecutes a case she knows to be false, the representative of a company facing a labor dispute who comes to the negotiating table with no intent of compromising.That is, bad faith is when someone's apparent reasons for doing something aren't the same as the real reasons. This is distinct from malign intent. The uniformed solider who shoots you without pretending to surrender is acting in good faith, because what you see is what you get: the man whose clothes indicate that his job is to try to kill you is, in fact, trying to kill you.The policy of assuming good faith (and mercilessly punishing rare cases of bad faith when detected) would make sense if you lived in an honest world where what you see generally is what you get (and you wanted to keep it that way), a world where the possibility of hidden motives in everyday life wasn't a significant consideration.On the contrary, however, I think hidden motives in everyday life are ubiquitous. As evolved creatures, we're designed to believe as it benefited our ancestors to believe. As social animals in particular, the most beneficial belief isn't always the true one, because tricking your conspecifics into adopting a map that implies that they should benefit you is sometimes more valuable than possessing the map that reflects the territory, and the most persuasive lie is the one you believe yourself. The universal human default is to come up with reasons to persuade the other party why it's in their interests to do what you want - but admitting that you're doing that isn't part of the game. A world where people were straightforwardly trying to inform each other would look shocking and alien to us.But if that's the case (and you shouldn't take my word for it), being touchy about bad faith accusations seems counterproductive. If it's common for people's stated reasons to not be the same as the real reasons, it shouldn't be beyond the pale to think that of some particular person, nor should it necessarily entail cutting the "bad faith actor" out of public life - if only because, applied consistently, there would be no one left. Why would you trust anyone so highly as to think they never have a hidden agenda? Why would you trust yourself?The conviction that "bad faith" is unusual contributes to a warped view of the world in which conditions of information warfare are rationalized as an inevitable background fact of existence. In particular, people seem to believe that persistent good faith disagreements are an ordinary phenomenon - that there's nothing strange or unusual about a supposed state of affairs in which I'm an honest seeker of truth, and you're an honest seeker of truth, and yet we end up persistently disagreeing on some question of fact.I claim that this supposedly ordinary state of affairs is deeply weird at best, and probably ...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Low-Hanging Fruit Prior and sloped valleys in the loss landscape, published by Dmitry Vaintrob on August 24, 2023 on LessWrong.You can find code for the referenced experiments in this GitHub repositoryMany have postulated that training large neural networks will enforce a simplicity, or Solomonoff prior. This is grounded in the idea that simpler solutions occupy expansive regions in the weight space (there exist more generalization directions in weight space along which loss does not increase or increases very little), translating to a broad attractor basin where perturbations in weight adjustments have a marginal impact on the loss.However, stochastic gradient descent (SGD), the workhorse of deep learning optimization, operates in a manner that challenges this simplicity-centric view. SGD is, by design, driven by the immediate gradient on the current batch of data. The nature of this process means that SGD operates like a greedy heuristic search, progressively inching towards solutions that may be incrementally better but not necessarily the simplest.Part of this process can be understood as a collection of "grokking" steps, or phase transitions, where the network learns and "solidifies" a new circuit corresponding to correctly identifying some relationships between weights (or, mathematically, finding a submanifold). This circuit then (often) remains "turned on" (i.e., this relationship between weights stays in force) throughout learning.From the point of view of the loss landscape, this can be conceptualized as recursively finding a valley corresponding to a circuit, then executing search within that valley until it meets another valley (corresponding to discovering a second circuit), then executing search in the joint valley of the two found circuits, and so on. As the number of circuits learned starts to saturate the available weight parameters (in the underparametrized case), old circuits may get overwritten (i.e., the network may leave certain shallow valleys while pursuing new, deeper ones). However, in small models or models not trained to convergence, we observe that large-scale circuits associated with phase transitions largely survive to the end.A greedier pictureThis idea aligns with what we call the low-hanging fruit prior concept. Once a solution that reduces loss reasonably is identified, it becomes more computationally efficient to incrementally refine this existing strategy than to overhaul it in search of an entirely new solution, even if the latter might be simpler. This is analogous to continuously picking the lowest-hanging fruit / cheapest way to reduce loss at each stage of the gradient descent optimization search process.This model predicts that SGD training processes are more likely to find solutions that look like combinations of shallow circuits and heuristics working together rather than simpler but less decomposable algorithms. In a mathematical abstraction, suppose that we have an algorithm that consists of two circuits, each of which requires getting 10 parameters right (note that this corresponds to a measure of complexity), and each of which independently reduces the loss. Then the algorithm resulting from learning both circuits has a "complexity measure" of 20, but is more likely to be learned than a "complexity 15" algorithm with the same loss if it cannot be learned sequentially (as it is exponentially harder to correctly "guess" 20 parameters than to correctly "guess" 10 parameters twice).Note that in general, the picture is more complicated: even when learning a single "atomic" circuit that cannot be further decomposed, the question of how easy it is to learn is not equivalent to the information content (how many parameters need to be learned), but incorporates more qualitative phenomena like basin shallowness or, m...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: A Theory of Laughter, published by Steven Byrnes on August 23, 2023 on LessWrong.1. tl;drThere should be parallel explanations for laughter at two levels.At the brain level, there should be some mechanism / algorithm that produces laughter, and it should fit the data of when people laugh in practice.At the evolution level, there should be some explanation for why this mechanism exists in the first place. Why was it adaptive in our ancestors? And where did it come from - are there homologues in other animals?I'll summarize my proposals for both of these, in the opposite order:1.1 First half of the tl;dr: Laughter in terms of evolutionI endorse the popular theory that laughter is an indicator of "play", homologous to the play-related vocalizations and body language in other animals (e.g. the dog's "play bow").The evolutionary purpose of play is "practice for future dangerous situations". For example, a wolf pup that engages in play-fighting and play-chasing would presumably be more skilled in its future real-life fights and chases.The evolutionary purpose of innate communicative play signals, like laughter in humans and play-bows in dogs, is to reduce the probability of accidental escalation from practice to serious. For example, if a play-fight between two wolf-pups escalates into a real fight between the pups, that's dangerous for both pups. If the pups are emitting and responding to communicative play signals, then that kind of escalation is much less likely to happen. It's kinda the same idea as "safewords" in fight-related sports (among other places).1.2 Second half of the tl;dr: Laughter in terms of brain algorithmsMy (oversimplified) pseudocode brain "business logic" for laughter is something like:PROPOSED BRAIN PSEUDOCODE FOR LAUGHTER:(A) IF my hypothalamus & brainstem are getting some evidence that I'm in danger(the "evidence" here would presumably be some of the same signals that, by themselves, would tend to activate the sympathetic nervous system)(B) AND my hypothalamus & brainstem are simultaneously getting stronger evidence that I'm safe(the "evidence" here would presumably be some of the same signals that, by themselves, would tend to activate the parasympathetic nervous system)(C) AND my hypothalamus & brainstem have evidence that I'm in a social situation(D) THEN I will emit innate play signals (e.g. laughter in humans), and also I will feel more energetic (on the margin), and more safe, less worried, etc.Indeed, I expect that there is some genetically-specified neuron group in the hypothalamus or brainstem (or more generally, what I call the Steering Subsystem), and that when future scientists look at its various connections and their functional properties, it will be straightforwardly obvious that this neuron group and its connections are implementing the pseudocode above.(Side note: These scientists will also find that this neuron group has various other inputs that make laughing more or less likely on the margin - inputs related to mood etc. - which I omitted from the box for simplicity.)Note that nothing in this box is particularly tied to humans. If we're talking about 50kHz rat laughter instead of human laughter, I wouldn't change a single word in the box above. However, later in the post, I will talk about human laughter in particular, including humor, and I'll argue that this pseudocode box is a plausible match to the circumstances in which people laugh.Also, the path by which I initially came to guess this pseudocode box (namely, introspection) was independent of how I came to believe the evolutionary story (namely, I read it in a book and it seemed obviously right). But I claim that the two stories match up beautifully - that the pseudocode box above is the natural, straightforward way to implement the "spec" associated...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Large Language Models will be Great for Censorship, published by Ethan Edwards on August 22, 2023 on LessWrong.Produced as part of the SERI ML Alignment Theory Scholars Program - Summer 2023 CohortThanks to ev_ and Kei for suggestions on this post.LLMs can do many incredible things. They can generate unique creative content, carry on long conversations in any number of subjects, complete complex cognitive tasks, and write nearly any argument. More mundanely, they are now the state of the art for boring classification tasks and therefore have the capability to radically upgrade the censorship capacities of authoritarian regimes throughout the world.How Censorship WorkedIn totalitarian government states with wide censorship - Tsarist Russia, Eastern Bloc Communist states, the People's Republic of China, Apartheid South Africa, etc - all public materials are ideally read and reviewed by government workers to ensure they contain nothing that might be offensive to the regime. This task is famously extremely boring and the censors would frequently miss obviously subversive material because they did not bother to go through everything. Marx's Capital was thought to be uninteresting economics so made it into Russia legally in the 1890s.The old style of censorship could not possibly scale, and the real way that censors exert control is through deterrence and fear rather than actual control of communication. Nobody knows the strict boundary line over which they cannot cross, and therefore they stay well away from it. It might be acceptable to lightly criticize one small part of the government that is currently in disfavor, but why risk your entire future on a complaint that likely goes nowhere? In some regimes such as the PRC under Mao, chaotic internal processes led to constant reversals of acceptable expression and by the end of the Cultural Revolution most had learned that simply being quiet was the safest path. Censorship prevents organized resistance in the public and ideally for the regime this would lead to tacit acceptance of the powers that be, but a silently resentful population is not safe or secure.When revolution finally comes, the whole population might turn on their rulers with all of their suppressed rage released at once. Everyone knows that everyone knows that everyone hates the government, even if they can only acknowledge this in private trusted channels.Because proper universal and total surveillance has always been impractical, regimes have instead focused on more targeted interventions to prevent potential subversion. Secret polices rely on targeted informant networks, not on workers who can listen to every minute of every recorded conversation. This had a horrible and chilling effect and destroyed many lives, but also was not as effective as it could have been. Major resistance leaders were still able to emerge in totalitarian states, and once the government showed signs of true weakness there were semi-organized dissidents ready to seize the moment.Digital Communication and the Elusiveness of Total CensorshipTraditional censorship mostly dealt with a relatively small number of published works: newspapers, books, films, radio, television. This was somewhat manageable just using human labor. However in the past two decades, the amount of communication and material that is potentially public has been transformed with the internet.It is much harder to know how governments are handling new data because the information we have mostly comes from the victims of surveillance who are kept in the same deterrent fear as the past. If victims imagine the state is more capable than it is, that means the state is succeeding, and it is harder to assess the true capabilities. We don't have reliable accounts from insiders or archival access since no major regi...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Steven Wolfram on AI Alignment, published by Bill Benzon on August 21, 2023 on LessWrong.Joe Walker has a general conversation with Wolfram about his work and things and stuff, but there are some remarks about AI alignment at the very end:WALKER: Okay, interesting. So moving finally to AI, many people worry about unaligned artificial general intelligence, and I think it's a risk we should take seriously. But computational irreducibility must imply that a mathematical definition of alignment is impossible, right?WOLFRAM: Yes. There isn't a mathematical definition of what we want AIs to be like. The minimal thing we might say about AIs, about their alignment, is: let's have them be like people are. And then people immediately say, "No, we don't want them to be like people. People have all kinds of problems. We want them to be like people aspire to be.And at that point, you've fallen off the cliff. Because, what do people aspire to be? Well, different people aspire to be different and different cultures aspire in different ways. And I think the concept that there will be a perfect mathematical aspiration is just completely wrongheaded. It's just the wrong type of answer.The question of how we should be is a question that is a reflection back on us. There is no "this is the way we should be" imposed by mathematics.Humans have ethical beliefs that are a reflection of humanity. One of the things I realised recently is one of the things that's confusing about ethics is if you're used to doing science, you say, "Well, I'm going to separate a piece of the system," and I'm going to say, "I'm going to study this particular subsystem. I'm going to figure out exactly what happens in the subsystem. Everything else is irrelevant."But in ethics, you can never do that. So you imagine you're doing one of these trolley problem things. You got to decide whether you're going to kill the three giraffes or the eighteen llamas. And which one is it going to be?Well, then you realise to really answer that question to the best ability of humanity, you're looking at the tentacles of the religious beliefs of the tribe in Africa that deals with giraffes, and this kind of thing that was the consequence of the llama for its wool that went in this supply chain, and all this kind of thing.In other words, one of the problems with ethics is it doesn't have the separability that we've been used to in science. In other words, it necessarily pulls in everything, and we don't get to say, "There's this micro ethics for this particular thing; we can solve ethics for this thing without the broader picture of ethics outside."If you say, "I'm going to make this system of laws, and I'm going to make the system of constraints on AIs, and that means I know everything that's going to happen," well, no, you don't. There will always be an unexpected consequence. There will always be this thing that spurts out and isn't what you expected to have happen, because there's this irreducibility, this kind of inexorable computational process that you can't readily predict.The idea that we're going to have a prescriptive collection of principles for AIs, and we're going to be able to say, "This is enough, that's everything we need to constrain the AIs in the way we want," it's just not going to happen that way. It just can't happen that way.Something I've been thinking about recently is, so what the heck do we actually do? I was realising this. We have this connection to ChatGPT, for example, and I was thinking now it can write Wolfram Language code, I can actually run that code on my computer. And right there at the moment where I'm going to press the button that says, "Okay, LLM, whatever code you write, it's going to run on my computer," I'm like, "That's probably a bad idea," because, I don't know, it's going ...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AI Forecasting: Two Years In, published by jsteinhardt on August 20, 2023 on LessWrong.Two years ago, I commissioned forecasts for state-of-the-art performance on several popular ML benchmarks. Forecasters were asked to predict state-of-the-art performance on June 30th of 2022, 2023, 2024, and 2025. While there were four benchmarks total, the two most notable were MATH (a dataset of free-response math contest problems) and MMLU (a dataset of multiple-choice exams from the high school to post-graduate level).One year ago, I evaluated the first set of forecasts. Forecasters did poorly and underestimated progress, with the true performance lying in the far right tail of their predicted distributions. Anecdotally, experts I talked to (including myself) also underestimated progress. As a result of this, I decided to join the fray and registered my own forecasts for MATH and MMLU last July.June 30, 2023 has now passed, so we can resolve the forecasts and evaluate my own performance as well as that of other forecasters, including both AI experts and generalist "superforecasters". I'll evaluate the original forecasters that I commissioned through Hypermind, the crowd forecasting platform Metaculus, and participants in the XPT forecasting competition organized by Karger et al. (2023), which was stratified into AI experts and superforecasters.Overall, here is how I would summarize the results:Metaculus and I did the best and were both well-calibrated, with the Metaculus crowd forecast doing slightly better than me.The AI experts from Karger et al. did the next best. They had similar medians to me but were (probably) overconfident in the tails.The superforecasters from Karger et al. did the next best. They (probably) systematically underpredicted progress.The forecasters from Hypermind did the worst. They underpredicted progress significantly on MMLU.Interestingly, this is a reverse of my impressions from last year, where even though forecasters underpredicted progress, I thought of experts as underpredicting progress even more. In this case, it seems the experts did pretty well and better than generalist forecasters.What accounts for the difference? Some may be selection effects (experts who try to register forecasts are more likely to be correct). But I'd guess some is also effort: the expert "forecasts" I had in mind last year were from informal hallway conversations, while this year they were formal quantitative predictions with some (small) monetary incentive to be correct. In general, I think we should trust expert predictions more in this setting (relative to their informal statements), and I'm now somewhat more optimistic that experts can give accurate forecasts given a bit of training and the correct incentives.In the rest of the post, I'll first dive into everyone's forecasts and evaluate each in turn. Then, I'll consider my own forecast in detail, evaluating not just the final answer but the reasoning I used (which was preregistered and can be found here).My forecasts, and othersAs a reminder, forecasts are specified as probability distributions over some (hopefully unambiguously) resolvable future outcome. In this case the outcome was the highest credibly claimed benchmark accuracy by any ML system on the MATH and MMLU benchmarks as of June 30, 2023.My forecasts from July 17, 2022 are displayed below as probability density functions, as well as cumulative distribution functions and the actual result:MATHMMLUResult: 69.6% (Lightman et al., 2023)Result: 86.4% (GPT-4)Orange is my own forecast, while green is the crowd forecast of Metaculus on the same date. For MATH, the true result was at my 41st percentile, while for MMLU it was at my 66th percentile. I slightly overestimated progress on MATH and underestimated MMLU, but both were within my range of e...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The U.S. is mildly destabilizing, published by lc on August 18, 2023 on LessWrong.We focus so much on arguing over who is at fault in this country that I think sometimes we fail to notice or alert on what's actually happening. I would just like to point out, without attempting to assign blame, that American political institutions appear to be losing common knowledge of their legitimacy, and abandoning certain important traditions of cooperative governance. It would be slightly hyperbolic, but not unreasonable to me, to term what has happened "democratic backsliding".Let's imagine America of 2012 was measured 0.8 on the fictionally accurate "legitimate democracy index", and Hungary of 2012 was measured 0.5. My thesis is that the we'd now be at 0.75, and that the regression seems to have calcified despite the culture war calming down since 2020. Within the last three or four years we have seen:The first presidential election in the history of the country ever contested by one of the main candidates; an election now considered probably or definitely illegitimate by nearly a third of Americans.The world's largest protest-riot ever, when measured by estimated damage to property or number of participants.Spontaneous mob assaults of the capitol building.The leader of the opposition party being arrested on a mix of real and recently-invented process crimes in several different jurisdictions a year before his campaign.Recent, and novel, movements by Republicans to fine and censure Democratic congressmen millions of dollars outside of the criminal justice system.Serious attempts at dramatically expanding political control over the civil service and, if you can permit me to speak anecdotally, serious and successful attempts at unprecedented political loyalty testing of appointed silovik.You can disagree with how any one political faction is characterizing the above events, or how I'm characterizing the above events. One take, for example, would be that Donald Trump is a clown and that all of his indictments are perfectly legitimate and that they ultimately demonstrate the dispassionate fairness of our nation's prosecutors. But even if that's the case, perception is the leading indicator for democratic stability and a large amount of Republicans do not agree with that interpretation. Since Republicans now believe that the arrests are politically motivated, and that Democrats are hitting "defect" by electing those prosecutors, they are pressuring their politicians to escalate and calling them traitors when they refuse to do so. This is itself bad.It's of course possible to exaggerate the danger. I do not expect the entire political system of the United States is going to change anytime soon. But since 1989 I think it has been appropriate to have a degree of knightian uncertainty in predicting the eternal dominance of this or that regime, on the basis that modern technology and secret police make resistance impossible. If you currently habitually round probabilities of serious repression or further democratic backsliding in the West to zero, I suggest moving that up to 1%-per-decade and spending a little bit of time thinking about what you'd do if this continues for five more years and your estimate increases to 5 or 10 percent.Presidential Election Poll, January 2021, UmassAdam Schiff Defeats effort to fine and censure him 15 million dollarsSchedule F AppointmentTrump's inner circle is secretly making plans to fire thousands government employees if he wins in 2024, report saysThanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: 6 non-obvious mental health issues specific to AI safety., published by Igor Ivanov on August 18, 2023 on LessWrong.IntroI am a psychotherapist, and I help people working on AI safety. I noticed patterns of mental health issues highly specific to this group. It's not just doomerism, there are way more that are less obvious.If you struggle with a mental health issue related to AI safety, feel free to leave a comment about it and about things that help you with it. You might also support others in the comments. Sometimes such support makes a lot of difference and people feel like they are not alone.All the examples in this post are anonymized and changed in a way that it's impossible to recognize a specific person behind them.AI safety is a rather unusual fieldThe problems described in this post arise because AI safety is not an ordinary field to work in. Many people within the AI safety community believe that it might be the most important field of work, but the general public mostly doesn't care that much. Also, the field itself is extremely competitive and newcomers often have hard time getting a job.No one really knows when we will create AGI, and whether we will be able to keep it aligned. If we fail to align AGI, the humanity might extinct, and even if we succeed, it will radically transform the world.PatternsAGI will either cause doom or create a utopia. Everything else seem unimportant and meaningless.Alex is an ML engineer working in a startup that fights with aging. He believes that AGI will either destroy humanity or bring a utopia, and among other things it will stop aging, so Alex thinks that his job is meaningless, and quits it. He also sometimes asks himself "Should I invest? Should I exercise? Should I even floss my teeth? This all seems meaningless."No one knows how the post-AGI world will look like. All predictions are wild speculations, and it's very hard to tell whether any actions unrelated to AI safety are meaningful. This uncertainty can cause anxiety and depressionThese problems are an exacerbated version of existential problem of meaninglessness of life, and the way to mitigate them is to rediscover meaning in the world that ultimately doesn't have meaning.I don't know when we will create AGI and if we will be able to align it, so I feel like I have no control over it.Bella is an anxious person, and she recently got interested in AI safety and she realized that nobody know for sure how to align AGI.She feels that AGI might pose an extreme danger, and there is nothing she can do. She even can't understand how much time do we have. A year? Five years? This uncertainty makes here even more anxious. And what if the takeoff will be so rapid that no one will understand what is going on?Bella is meeting a psychotherapist, but they treat her fear as something irrational. This doesn't help, and only makes Bella more anxious. She feels like even her therapist doesn't understand her.AI safety is a big part of my life, but others don't care that much about it. I feel alienated.Chang is an ML scientist working on mechanistic interpretability in AI lab. AI safety consumed all his life and became a part of his identity. He constantly checks AI safety influencers on Twitter, he spends a lot of time reading LessWrong and watching AI podcasts. He even made a tatoo of a paperclip.Chang lives outside of major AI safety hubs, and he feels a bit lonely because there is no one to talk about AI safety in person.Recently he attended his aunt's birthday party. He talked about alignment with his family. They were a bit curious about the topic, but didn't care that much. Chang feels like they just don't get it.Working on AI safety is so important that I neglected other parts of my life and burned-out.Dmitry is an undergrad student. He believes that AI safet...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Book Launch: "The Carving of Reality," Best of LessWrong vol. III, published by Raemon on August 17, 2023 on LessWrong.The Carving of Reality, third volume of the Best of LessWrong books is now available on Amazon (US).The Carving of Reality includes 43 essays from 29 authors. We've collected the essays into four books, each exploring two related topics. The "two intertwining themes" concept was first inspired when as I looked over the cluster of "coordination" themed posts, and noting a recurring motif of not only "solving coordination problems" but also "dealing with the binding constraints that were causing those coordination problems."I've included the foreword from "Coordination & Constraint", which I think conveys the overall spirit and context of the books:Each year, the LessWrong community votes on the best posts from the previous year, to see which posts have stood the tests of time.In 2020, the highest ranked post was Catherine Olsson's announcement of microCOVID.org, a calculator for evaluating COVID risk. MicroCOVID is one of the clearest success stories of the 'rationalist mindset' that I know of. Creating it involved research during the early pandemic, when information was scarce, and time was of the essence - a classic situation where the traditional scientific process is inadequate and LessWrong-style rationality tools are valuable. It also required a quantitative mindset, and willingness to assign numbers to risks and make tradeoffs.But microCOVID.org is most interesting to me as a tool for coordination. It doesn't just let individuals make better life choices. Microcovid changed the entire covid coordination landscape by relaxing a constraint. Previously, if you lived with people with varying covid-caution preferences, and you wanted to hang out with someone from another house of people with varying covid-caution preferences. your only option was to have fairly involved negotiations on a case-by-case basis. Many people I know grew exhausted from negotiating, and gave up on trying to visit their friends. The microCOVID tool gave people a simplified "risk budget", letting them do whatever activities made sense to them as long as they didn't overspend."Negotiation energy" was a limiting resource, and microcovid.org made negotiation radically cheaper. It also opened up entirely new options, like "create a household microCOVID tax" (some houses decided that you could do whatever activities you wanted, you just had to pay other housemates $1 per microcovid).The proximate inspiration for the theme of this book (and included herein) are John Wentworth's posts "Coordination as Scarce Resource", "Transportation as Constraint", and "Interfaces as Scarce Resource." Other posts explore the nature of particular constraints that society faces - Zvi's posts on "Simulacra Levels and their Interactions," "The Road to Mazedom," and "Motive Ambiguity" each spell out how and why some communication is systematically distorted. And while they don't give us a solution, they help ask a question - what would need to change, in order for society to coordinate at scale, without incentivizing distorted communication?COORDINATION & CONSTRAINTJohn WentworthCoordination as a Scarce Resource John WentworthTransportation as a Constraint John WentworthInterfaces as a Scarce Resource Catherine OlssonMicroCOVID.org Jacob FalcovichSeeing the Smoke AlkjashPain is not the unit of Effort AlkjashIs Success the Enemy of Freedom? Zvi MowshowitzSimulacra Levels and their Interactions Zvi MowshowitzThe Road to Mazedom Zvi MowshowitzMotive Ambiguity Scott AlexanderStudies On Slack Jim Babcock, Elizabeth Van NostrandCredibility of the CDC on SARS-CoV-2 Raymond Arnold"Can you keep this confidential? How do you know?" Abram DemskiMost Prisoner's Dilemmas are Stag Hunts; Most Stag Hunts ar...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Ten Thousand Years of Solitude, published by agp on August 16, 2023 on LessWrong.This is a linkpost for the article "Ten Thousand Years of Solitude", written by Jared Diamond for Discover Magazine in 1993, four years before he published Guns, Germs and Steel. That book focused on Diamond's theory that the geography of Eurasia, particularly its large size and common climate, allowed civilizations there to dominate the rest of the world because it was easy to share plants, animals, technologies and ideas. This article, however, examines the opposite extreme.Diamond looks at the intense isolation of the tribes on Tasmania - an island the size of Ireland. After waters rose, Tasmania was cut off from mainland Australia. As the people there did not have boats, they were completely isolated, and did not have any contact - or awareness - of the outside world for ten thousand years.How might a civilization develop, all on its own, for such an incredible period of time?If you ask any anthropologist to summarize in one phrase what was most distinctive about the Tasmanians, the answer will surely be the most primitive people still alive in recent centuries.The "entire material corpus" of Tasmania only amounted to two dozen items in total - and did not include mounted stone stools, bone tools, or any clothing at all. Despite average low temperatures in winter of 41 degrees Fahrenheit, the Tasmanians were completely naked. In addition to the poor quality of tools in Tasmania, they also refused to eat fish, which were plentiful in the waters around the island. The material culture and wellbeing of the Tasmanians was significantly worse off than that of the Australians.Australian products absent in Tasmania included the spear-thrower, a hand-held device to increase a spear's throwing distance and propulsive force; ground or polished stone tools; mounted stone tools, such as hatchets or adzes with a handle; bone tools, such as needles and awls; fire-making equipment, such as a fire drill; and nets, traps, or hooks to catch fish, birds, or mammals. Without mounted stone tools, Tasmanians couldn't fell a big tree, hollow out a canoe, or carve a wooden bowl. Without bone tools, they couldn't sew warm clothes or watertight bark canoes.The poverty of the Tasmanians was shocking to the first European explorers. They did not understand how the Tasmanians could have reached the island without boats, and they didn't understand why the Tasmanians had astonishingly little technology. The 'arrival' question is easy to answer - they walked there when the oceans were lower - but it's the technology question that I find most fascinating. If the Tasmanians came from Australia, then shouldn't they at a baseline have the tools and skills that the Australians possessed at the time that they left? But in fact the Tasmanians seem to have regressed since the beginning of their isolation.The Tasmanians actually abandoned some practices that they shared with Australia 10,000 years ago. This idea violates cherished views of human nature, since we tend to assume that history is a long record of continual progress. Nevertheless, it is now clear that Tasmanians did abandon at least two important practices.One was the production of bone tools. With bone, one can fashion objects virtually impossible to make out of stone or wood--such as needles. In southeast Australia at the time of European discovery, aboriginal Australians were using bone tools as awls and reamers to pierce animal hides, as pins to fasten the hides into cloaks, and as needles to sew hides into still warmer clothing or to knit fishing nets. As recently as 7,000 years ago, Tasmanian tools included bone tools that resembled Australia's awls, reamers, and needles. Thereafter, the variety of Tasmanian bone tools gradually decreased with tim...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Decomposing independent generalizations in neural networks via Hessian analysis, published by Dmitry Vaintrob on August 14, 2023 on LessWrong.In our joint SERI MATS project, we came up with a series of equations and experiments to mechanistically understand and steer the generalization behavior of neural nets. The core conceit is to locate the circuits (which we call "modules") responsible for implementing different generalizations using a toolbox of techniques related to Hessian eigenvectors. This is a general-audience distillation of our work.We hope most of the ideas and high-level goals are understandable to a non-expert, though, for most of our experiments, we attempt to include "supplementary" material with the main mathematical intuitions and concrete equations that would allow someone to reproduce our work. We plan in the coming weeks to write multiple follow-up distillations and discussions, both of some of the more technical parts of our work and of a few new insights into generalization behavior and phase transitions in general that came out of experiments involving our Hessian toolbox.IntroductionA central problem for inner alignment is understanding how neural nets generalize off-distribution. For example, a powerful AI agent trained to make people happy can generalize either by choosing actions that deceptively look good to its overseers or those that truly align with human values. The same diversity of generalization is already seen in existing real-world tasks both in minor ways (image classifiers classifying cars by learning to recognize wheels vs. windows) and in serious ways (language models appearing honest by agreeing with the user vs. insisting on consensus opinions).One approach to steer between generalizations is activation steering, which Nina is investigating as her other SERI MATS project. This aims to encourage the neural net to implement one possible generalization (in this case, honestly reflecting the LLM's internal world model) instead of the other generalization (in this case, sounding good and correct to a particular user).While activation steering, supervised finetuning, and RLHF work well in practice and can make systems behave better, there is still a risk that powerful models generalize in unpredictable and potentially undesirable ways in out-of-distribution examples. In particular, for subtle alignment-related questions like deception or power-seeking, activation steering or RLHF may fix the "symptoms" of the problem on examples similar to the training corpus but may fail to fix the "underlying cause" and achieve aligned behavior.A somewhat ambitious alternative way to get at the "root" of a generalization problem instead of fixing its symptoms is to try to access it on a mechanistic level. Namely, imagine that on the level of the "internal architecture" of the neural net (something that is notoriously hard to access but can sometimes be partially interpreted), the two generalizations get executed by at least somewhat independent modules (i.e., parallel circuits: the term comes from this paper). If we were able to identify and split up these two modules cleanly, we might be able to find weight perturbation vectors that destroy ("ablate") one of them while preserving the other. The resulting method is now provably robust: it prevents one of the generalizations (understood mechanistically as the underlying module) from getting executed at any level, thus solving both the symptom and the underlying cause.This algorithm for tuning generalizations can be possible only if the underlying mechanistic model (of different independent generalization "modules" which can be consistently found and independently ablated) is correct or partially correct to a relevant approximation. In order to even begin to engage with it, we need answe...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: We Should Prepare for a Larger Representation of Academia in AI Safety, published by Leon Lang on August 13, 2023 on LessWrong.Epistemic Status: I had the idea for the post a few days ago and quickly wrote it down while on a train. I'm very curious about other perspectives.TL;DR: The recent increased public interest in AI Safety will likely lead to more funding for and more researchers from academia. I expect this increase to be larger than that of non-academic AI Safety work. We should prepare for that by thinking about how we "onboard" new researchers and how to marginally allocate resources (time and money) in the future.Why I think academia's share in AI safety will increaseWith the recent public interest in AI (existential) safety, many people will think about how they can help. Among people who think "I might want to do research on AI Safety", most will come from academia because that's where most research happens. Among people who will think "I should fund AI Safety research", most will fund academic-style research because that's where most research talent sits, and because it's the "normal" thing to do. I expect this increase to be larger than that of AI Safety researchers in companies (though with less certainty), AI Safety orgs, or independent researchers of, e.g., the "Lesswrong / Alignment Forum" style.Weak evidence that this is already happeningAt the university of Amsterdam, where I'm a PhD student, there has been increased interest in AI Safety recently. In particular, one faculty actively starts to think about AI existential safety and wants to design a course that will include scalable oversight, and â¥4 other faculty are at least starting to get informed about AI existential safety with an "open mind".What might one do to prepare?Needless to say, I didn't think about this a lot, so take the following with a grain of salt and add your own ideas.Academics will mostly read papers that are at least on arxiv. So to "onboard" them, it seems more important than in the past to make the most important insights from lesswrong or the alignment forum accessible to academics.Doing a PhD might become more worthwhile because it's easier now to have an alignment career in academia.Doing a PhD might also become less worthwhile because "academic-style" research into AI safety will be less neglected going forward. Whether you buy this argument depends on your views on how open-minded academia is to the most important types of AI Safety research.In general, it seems worthwhile to anticipate which types of research will be "covered" by academia, and how to prioritize research in this landscape.Grantmakers should think about how to react to a potentially changing funding landscape, with many more "traditional" grantmakers funding research in academia, and more talented academics being open to work on AI existential safety. This could also mean to prioritize work that is substantially different than what will be researched in academia.UncertaintiesI find it plausible that the representation of AI Safety researchers in companies like OpenAI and DeepMind will also grow very fast, though I think the increase will be smaller than in academia.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: [Linkpost] Personal and Psychological Dimensions of AI Researchers Confronting AI Catastrophic Risks, published by Bogdan Ionut Cirstea on August 13, 2023 on LessWrong.This is a linkpost for.Yoshua Bengio:For most of these years, I did not think about the dual-use nature of science because our research results seemed so far from human capabilities and the work was only academic. It was a pure pursuit of knowledge, beautiful, but mostly detached from society until about a decade ago. I now believe that I was wrong and short-sighted to ignore that dual-use nature. I also think I was not paying enough attention to the possibility of losing control to superhuman AIs.[...] it started to dawn on me that my previous estimates of when human-level AI would be reached needed to be radically changed. Instead of decades to centuries, I now see it as 5 to 20 years with 90% confidence.And what if it was, indeed, just a few years?I started reading more about AI safety and came to a critically important conclusion: we do not yet know how to make an AI agent controllable and thus guarantee the safety of humanity! And yet we are - myself included until now - racing ahead towards building such systems.It is painful to face the idea that we may have been contributing to something that could be greatly destructive. Human nature will lead us towards brushing aside these thoughts or finding comfort in reassuring arguments rather than face the full horror of such possibilities. Bringing the benefits of AI to the table is not sufficient to compensate if the possible negative outcomes include catastrophic misuses of AI on par with nuclear war and pandemics, or even existential risk.As scientists, we should avoid making claims we can't support; but as decision-makers we also ought to act under uncertainty to take precautions. In spite of our differences in points of view, it's time for our field of AI to seriously discuss the questions: what if we succeed? What if potentially dangerous superhuman AI capabilities are developed sooner than expected? Let's embrace these challenges and our differences, while being mindful of each other's humanity and our unique emotional and psychological journeys in this new era of AI.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Biological Anchors: The Trick that Might or Might Not Work, published by Scott Alexander on August 12, 2023 on LessWrong.This post originally posted on Astral Codex Ten on Feb 23 2022.It was printed in The Carving of Reality, the third volume of the Best of LessWrong book series. It was included as a (shorter) replacement for Ajeya Cotra's Draft report on AI timelines, and Eliezer's Biology-Inspired AGI Timelines: The Trick That Never Works, covering the topic from multiple sides.It's crossposted here with Scott's permission for completeness (i.e. having all essays in the book appear on LessWrong).IntroductionI've been trying to review and summarize Eliezer Yudkowksy's recent dialogues on AI safety. Previously in sequence: Yudkowsky Contra Ngo On Agents. Now we're up to Yudkowsky contra Cotra on biological anchors, but before we get there we need to figure out what Cotra's talking about and what's going on.The Open Philanthropy Project ("Open Phil") is a big effective altruist foundation interested in funding AI safety. It's got $20 billion, probably the majority of money in the field, so its decisions matter a lot and it's very invested in getting things right. In 2020, it asked senior researcher Ajeya Cotra to produce a report on when human-level AI would arrive. It says the resulting document is "informal" - but it's 169 pages long and likely to affect millions of dollars in funding, which some might describe as making it kind of formal. The report finds a 10% chance of "transformative AI" by 2031, a 50% chance by 2052, and an almost 80% chance by 2100.Eliezer rejects their methodology and expects AI earlier (he doesn't offer many numbers, but here he gives Bryan Caplan 50-50 odds on 2030, albeit not totally seriously). He made the case in his own very long essay, Biology-Inspired AGI Timelines: The Trick That Never Works, sparking a bunch of arguments and counterarguments and even more long essays.There's a small cottage industry of summarizing the report already, eg OpenPhil CEO Holden Karnofsky's article and Alignment Newsletter editor Rohin Shah's comment. I've drawn from both for my much-inferior attempt.Part I: The Cotra ReportAjeya Cotra is a senior research analyst at OpenPhil. She's assisted by her fiancee Paul Christiano (compsci PhD, OpenAI veteran, runs an AI alignment nonprofit) and to a lesser degree by other leading lights. Although not everyone involved has formal ML training, if you care a lot about whether efforts are "establishment" or "contrarian", this one is probably more establishment.The report asks when will we first get "transformative AI" (ie AI which produces a transition as impressive as the Industrial Revolution; probably this will require it to be about as smart as humans). Its methodology is:1. Figure out how much inferential computation the human brain does.2. Try to figure out how much training computation it would take, right now, to get a neural net that does the same amount of inferential computation. Get some mind-bogglingly large number.3. Adjust for "algorithmic progress", ie maybe in the future neural nets will be better at using computational resources efficiently. Get some number which, realistically, is still mind-bogglingly large.4. Probably if you wanted that mind-bogglingly large amount of computation, it would take some mind-bogglingly large amount of money. But computation is getting cheaper every year. Also, the economy is growing every year. Also, the share of the economy that goes to investments in AI companies is growing every year. So at some point, some AI company will actually be able to afford that mind-boggingly-large amount of money, deploy the mind-bogglingly large amount of computation, and train the AI that has the same inferential computation as the human brain.5. Figure out what year t...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AI #24: Week of the Podcast, published by Zvi on August 11, 2023 on LessWrong.In addition to all the written developments, this was a banner week for podcasts.I would highlight four to consider listening to.Dario Amodei of Anthropic went on The Lunar Society to talk to Dwarkesh Patel. We got our best insight so far into where Dario's head is at, Dwarkesh is excellent at getting people to open up like this and really dive into details.Jan Leike, OpenAI's head of alignment, went on 80,000 hours with Robert Wiblin. If you want to know what is up with the whole superalignment effort, this was pretty great, and left me more optimistic. I still don't think the alignment plan will work, but there's a ton of great understanding of the problems ahead and an invitation to criticism, and a clear intention to avoid active harm, so we can hope for a pivot as they learn more.Tyler Cowen interviewed Paul Graham. This was mostly not about AI, but fascinating throughout, often as a clash of perspectives about the best ways to cultivate talent. Includes Tyler Cowen asking Paul Graham about how to raise someone's ambition, and Paul responding by insisting on raising Tyler's ambition.I got a chance to go on EconTalk and speak with Russ Roberts about The Dial of Progress and other matters, mostly related to AI. I listen to EconTalk, so this was a pretty special moment. Of course, I am a little bit biased on this one.Capabilities continue to advance at a more modest pace, so I continue to have room to breathe, which I intend to enjoy while it lasts.Table of ContentsIntroduction.Table of Contents.Language Models Offer Mundane Utility. Proceed with caution.Language Models Don't Offer Mundane Utility. Not with these attitudes.GPT-4 Real This Time. Time for some minor upgrades.Fun With Image Generation. Some fun, also some not so fun.Deepfaketown and Botpocalypse Soon. They keep ignoring previous instructions.They Took Our Jobs. People really, really do not like it when you use AI artwork.Introducing. Real time transcription for the deaf, also not only for the deaf.In Other AI News. Various announcements, and an exciting Anthropic paper.There Seems To Be a Standard Issue RLHF Morality. It has stages. What's next?Quiet Speculations. Cases for and against expecting a lot of progress.The Quest for Sane Regulation. Confidence building, polls show no confidence.The Week in Audio. A cornucopia of riches, extensive notes on Dario's interview.Rhetorical Innovation. People are indeed worried in their own way.No One Would Be So Stupid As To. I always hope not to include this section.Aligning a Smarter Than Human Intelligence is Difficult. Grimes also difficult.People Are Worried About AI Killing Everyone. No one that new, really.Other People Are Not As Worried About AI Killing Everyone. Alan Finkel.The Lighter Side. Finally a plan that works.Language Models Offer Mundane UtilityControl HVAC systems with results comparable to industrial standard control systems.Davidad: I've witnessed many philosophical discussions about whether a thermostat counts as an AI, but this is the first time I've seen a serious attempt to establish whether an AI counts as a thermostat.Ethan Mollick offers praise for boring AI, that helps us do boring things.As context, one of the first major experimental papers on the impact of ChatGPT on work just came out in Science (based on the free working paper here) and the results are pretty impressive: in realistic business writing tasks, ChatGPT decreased the time required for work by 40%, even as outside evaluators rated the quality of work written with the help of AI to be 18% better than the ones done by humans alone.After using it, people were more worried about their jobs. but also significantly happier - why? Because a lot of work is boring, an...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Inflection.ai is a major AGI lab, published by nikola on August 9, 2023 on LessWrong.Inflection.ai (co-founded by DeepMind co-founder Mustafa Suleyman) should be perceived as a frontier LLM lab of similar magnitude as Meta, OpenAI, DeepMind, and Anthropic based on their compute, valuation, current model capabilities, and plans to train frontier models. Compared to the other labs, Inflection seems to put less effort into AI safety.Thanks to Laker Newhouse for discussion and feedback!Inflection has a lot of compute dedicated to training LLMsThey plan to scale up their cluster to 3 times the capacity used to train GPT-4."We'll be building a cluster of around 22,000 H100s. This is approximately three times more compute than what was used to train all of GPT4. Speed and scale are what's going to really enable us to build a differentiated product,""We believe in scale as the engine of progress in AI, and we are building one of the largest supercomputers in the world to develop and deploy the new generation of AIs."They can apparently train a model similarly capable to GPT-2 in 11 minutes of cluster time. (see Appendix)Side point: It seems that the actual H100s are (at least partly) owned by CoreWeave (a cloud compute provider), but that Inflection is one of CoreWeave's main clients. The specific cluster is a joint effort between Inflection and CoreWeave."They called us and said, 'Guys, we need you to build one of the most high-performance supercomputers on the planet to support our AI company,'" McBee said. "They call us and they say, 'This is what we're looking for, can you do it?'Inflection has a lot of fundingInflection is valued at $4B and has raised $1.5B, which is similar to Anthropic ($4.1B valuation, total raised $1.3B as of May 2023) and within an order of magnitude of OpenAI ($28B valuation, $11B raised as of April 2023).Inflection is on the cutting edge of LLMsTheir flagship LLM, Inflection-1, has similar benchmark results to GPT-3.5They seem to be currently training a model similarly capable to GPT-4. I expect them to finish training by the end of the year."We will also be releasing a technical memo detailing one of our models in the same compute class as PaLM-2 and GPT-4."Inflection plans to train frontier LLMsThey seem to plan to train models 10x or 100x the size of GPT-4 within 18 months."We are about to train models that are 10 times larger than the cutting edge GPT-4 and then 100 times larger than GPT-4. That's what things look like over the next 18 months."(it is unclear if "we" refers to Inflection or humanity)Inflection doesn't seem to acknowledge existential risks or have a sizable safety teamTheir safety site has zero mention of existential or catastrophic risks. Their white house memo is not very reassuring either.Out of 19 open job listings, only 2 are on the Safety team.If you look at their LinkedIn (which seems to list most of their current ~40 employees), zero of their employees are listed as working on AI safety at Inflection (one person has the word "safety" in their description but it's unclear that it's referring to their position at Inflection).I think that this mostly means that the Inflection Safety team members list themselves as "Technical staff" or don't have LinkedIns. But to me it seems like they have less than 5 people working on safety.Appendix: Estimating Inflection's computeHere are some back-of-the-envelope calculations for Inflection's current compute from three data sources. They result in estimates ranging around 2 orders of magnitude, centered around 4e18.FLOPs = plural of "floating point operation (FLOP)"FLOPS = floating point operations per secondThe H100 routeFrom the H100 datasheet, it seems like different components of the H100 (of which, different models exist), have different amounts of FL...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: A plea for more funding shortfall transparency, published by porby on August 8, 2023 on LessWrong.[This post is largely from the perspective of AI safety, but most of it should generalize.]For recipients, well calibrated estimates about funding probability and quantity are extremely valuable. Funding-dependent individuals and organizations need information to optimize their decisionmaking; incorrect estimates cause waste.At the moment, getting that information seems unnecessarily hard.To help with this, I would ask organizations up the funding chain to systematically and continuously provide bite-sized updates from their own perspectives on the funding situation when possible.This needn't be in the form of a lengthy report or deep-dive (though those are nice too!). For example, for a grantmaking organization with open applications, maybe something like:We've received V requests for funding totaling $W in the last month. We anticipate funding up to $X of these requests; we would fund up to about $Y if we had more funding.We don't anticipate significant changes in our funding capacity by default.Reports like this are already done occasionally. For example, I deeply appreciate reports like this one. My concern is that I'm not aware of any consistent source for this information.I worry that this is partially because writing a dense report takes a lot of time, and many of these organizations are massively overwhelmed as it is. To the extent that this is limiting reports, I would suggest that giving a tweet-sized miniupdate as things change (or just every few months) would be a huge improvement over the status quo.Even "oof, we're REALLY funding constrained" would be great! If you don't have time for collecting any numbers at all, a vibecheck is still useful!There's also no need for each report to attempt to capture the entire field's status, just the perspective of that one organization.An anecdote of awareness not propagating like it shouldFor the last six-ish months, I've been trying to figure out how to prioritize earning to give and direct alignment research. I've gotten in touch with several people and heard a number of different perspectives.None of them were "yeah, the field's pretty funding constrained right now, there are more people than funds, it's a major bottleneck."This continued a confusion I had starting with the FTX collapse. While I fully expected the field to be in a crunch immediately following the collapse, the vibe I collected from a number of people was that this was a probably-temporary thing, and this was seemingly supported by other things I heard a few months later.A lot of hints - organizations not being able to hire everyone they'd like to hire, not as many grants flowing as I'd expect, and salary targets way too low for a field with tons of cash on hand relative to talent - were inconsistent with the "unconstrained" narrative, but they were too weak in isolation to update me to reality.One side effect of this confusion was that I started work on a project before receiving grant funding for it based on an incorrectly high probability of being funded. It fell through; it's not a catastrophe, and I was prepared for that possibility, but I would have chosen differently if I had all the information that had existed privately at the time.Then it seems like everything snapped together with the common knowledge created by one post.If this is our collective mechanism for creating awareness, something seems broken.The futureWhile I've felt reasonably productive in my grant-funded research so far, it seems unlikely that my comparative advantage is in full-time alignment research as opposed to earning to give if this kind of funding environment continues.In addition to periodic updates about the current funding situation, I've love ...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Feedbackloop-first Rationality, published by Raemon on August 7, 2023 on LessWrong.I've been workshopping a new rationality training paradigm. (By "rationality training paradigm", I mean an approach to learning/teaching the skill of "noticing what cognitive strategies are useful, and getting better at them.")I think the paradigm has promise. I've beta-tested it for a couple weeks. It's too early to tell if it actually works, but one of my primary goals is to figure out if it works relatively quickly, and give up if it isn't not delivering.The goal of this post is to:Convey the frameworkSee if people find it compelling in its current formSolicit ideas for improvements, before I decide whether to invest heavily into a larger experiment around it.Rationality needs better feedback loopsClaim: Feedback loops are the most important thing ever. Hard things are hard because they have bad feedback loops. Some of the most important things (e.g. x-risk mitigation research) have the worst feedback loops.Bold prediction: You can learn to think better, even about confusing, poor-feedback domains. This requires developing the art of inventing feedback loops. And then, actually putting in a lot of deliberate practice effort.I've long been haunted by this Romeo Stevens comment (slightly paraphrased)Deliberate practice deliberate practice until you get really good identifying good feedback loops, and working with them.People have a really hard time with interventions often because they literally do not have a functioning causal model of the skill in question. People who apply deliberate practice to a working causal model often level up astonishingly quickly. Don't know if you have the appropriate causal model? Well, when you apply deliberate practice do you not get better? You're pulling on fake levers.In the past, I've tried to practice thinking. I've done explicit puzzle-solving exercises, and I have a day job that forces me to think about challenging questions on a regular basis. I sometimes have tried to refactor my day-job into something deliberate practice-shaped, but it never gelled.I think I've gotten better at thinking in the past 12 years. But I haven't gotten overwhelmingly obviously better at thinking. I recently decided to deliberate practicing "solve confusing problems", until I was demonstrably better at it, and to host some workshops where I tried helping other people practice too.I ended up settling into a paradigm of rationality training with five elements:Deliberate Practice. Do challenging cognitive exercises, at the edge of your ability, in a variety of domains, where it's obvious how well you're doing (i.e. clear cut answers, or you're making a metric go up).Metacognition. After deciding on the final answer for the exercise and finding out if you got it right, reflect on what you could have done better. Try to extract as much insight/wisdom/tools as you can from each exercise.Improve your practice feedback loop. Then, find or design better exercises, that cut more closely to your ultimate goals. Optimize exercises both for being concrete (i.e. you can tell if you succeeded), and for extracting as much insight/tools as possible during the metacognition step (i.e. they are a good difficulty in a domain I haven't already exhausted for insight)Improve your real-life feedback loop. Think about what sort of cognitive challenges you run into your day-job or main project, where you're bottlenecked in your ability to reason. How can you do better meta-reflection in those fuzzier, longer-timescale domains?Illegible goodness. In addition to the formal structure implied by the previous four bullets, also try random stuff that feels vaguely relevant and helpful, even if it you can't explain why. (I think some previous rationality training approaches leaned...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Stomach Ulcers and Dental Cavities, published by Metacelsus on August 6, 2023 on LessWrong.(This is a linkpost from my blog, De Novo)Recently I learned about an effort to prevent dental cavities by using genetically modified bacteria to outcompete cavity-causing bacteria. This got me thinking: why has the idea of preventing cavities by targeting bacteria not been more developed already?The current situation reminds me of the history of stomach ulcers. Before the 1980s, doctors recommended avoiding spicy foods and reducing stress to alleviate stomach ulcers. However, once Robin Warren and Barry Marshall proved ulcers were due to H. pylori infection, treatment with antibiotics to eliminate the bacteria became the standard of treatment.Today, dentists recommend avoiding sugary foods and brushing your teeth to prevent cavities. But we know cavities are caused by bacteria (in particular Streptococcus mutans), so why not directly attack cavity-causing bacteria?Some potential ideas:Selectively targeted antibioticsVaccines (previously tried in the 1980s, not very successful because it's difficult to get antibodies to penetrate biofilms, and also because S. mutans has several different strains with different antigenic profiles)Outcompeting S. mutans with different bacteria (the current effort by Aaron Silverbook, which I think is promising)Basically, what Aaron Silverbook is proposing to do is recreate a strain of S. mutans, termed BSC3-L1, that is deficient in lactic acid production. This was previously developed by a company called Oragenics, but they abandoned the effort (I think due to financial reasons). It seems Aaron's team is mostly people from software backgrounds, so they would probably appreciate help from any talented microbiologists who happen to be reading this post.In a famous case of self-experimentation, Marshall drank a culture of H. pylori and subsequently developed gastritis. For this work, Warren and Marshall earned the 2005 Nobel in Physiology/Medicine.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Private notes on LW?, published by Raemon on August 4, 2023 on LessWrong.Lately I've been noticing what a powerup is to read things in google docs, where I can take whatever notes I want as in-line comments without worrying about looking dumb or confusing. In changes my relationship to confusing passages, where I feel much more affordance to think through what exactly is confusing about it.As a general reading-habit, "copy it into google docs" is a pretty good habit. But I (and I think others on LW team although for slightly different reasons) have been thinking about building a feature directly into LW to facilitate it.One version of it might explicitly be "private notes" that are optimized as such.Another version of it might basically just take the side-comment button we already have and add a "private comments" option that lets you set the comment to "everyone", "only you", "you + author" (for giving the author feedback in a way that's more private than a comment but having more context included than a DM) [edit: also, sharing the comment with arbitrary people is a fairly obvious feature here]Curious what people think about this and what options they'd expect themselves to use.I'm maybe specifically wondering whether people expect a UI that's oriented around "arbitrary sharing" would feel good enough as a personal note-taking thing.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Password-locked models: a stress case for capabilities evaluation, published by Fabien Roger on August 3, 2023 on LessWrong.Thanks to Buck Shlegeris, Ryan Greenblatt, Kshitij Sachan and Jenny Nitishinskaya for helpful discussions and feedback on drafts. This work was done at Redwood Research. The views expressed are my own and do not necessarily reflect the views of the organization.Password-locked models are trained to exhibit certain capabilities only when a password is present in the query. Studying these models has two purposes:Testing how well capability evaluations work when applied to models which "aren't trying"Understanding how well organizations could hide the capabilities of their models.In this post, I make empirical claims which I believe to be true, check some of them on small language models, and describe what this implies for the control of potentially deceptive models.Key empirical claims: when using behavioral model evaluations, password-locked models:Look much less capable than they actually are when eliciting capabilities with promptingSometimes look much less capable than they actually are when eliciting capabilities with Reinforcement LearningExperimental evidence: I exhibit one case where they don't, and one case where they doCan't look much less capable than they actually when eliciting capabilities with fine-tuningKey takeaways, if the claims above are true:If an alignment plan relies on some models being too weak to perform certain tasks, then model evaluations aimed at checking the weakness of models should be able to elicit the locked capabilities of password-locked models. This is because password-locked models and deceptive models may be able to underperform in similar circumstances.API access and inference is not enough to evaluate capabilities of a model produced by AI labs trying to game regulations, even if AI labs are internally using the same API endpoint as auditors.Key Empirical ClaimsSimilarities Between Password-locked Models and Deceptively Aligned Models Pretending to Be DumbIt looks to me that password locked models have the same properties as the one described by Evan in When can we trust model evaluations? when he describes if and how deceptive models would be able to sandbag each category of behavioral evaluation (which he coins in his post):Behavioral Non-Fine-Tuning Evaluations: By construction, password-locked models will refuse to be capable no matter the prompt, except for prompts with the special password. But it's unclear how easy it is to build good password-locked models which are hard to crack (e.g. which don't leak the password with some clever prompting). Even if good password-locked models are built, there remain a large difference with deceptive models: the search for the correct password might have very different properties than the search for a prompt which makes a deceptive model believe it's time to strike.Behavioral I.I.D. Fine-Tuning Evaluations: I would be surprised if fine-tuning the password-locked model didn't lift the locking mechanism extremely quickly. This is similar to deceptive models (except if they are so strong that they can gradient hack). The main difference is that fine-tuning to remove password-locking on one task might transfer to another task, which might not be the case for deceptive models.Behavioral RL Fine-Tuning Evaluations: The success of evaluations which use Reinforcement Learning depends, as for deceptive models, on the success of exploration. If supervised fine-tuning is done with only low quality answers, then the password locked models will just continue to output low-quality answers and will never explore its way into high quality ones - except if there is a natural "path" gradually improving the quality of answers. But it is possible to emulate a thre...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: My current LK99 questions, published by Eliezer Yudkowsky on August 1, 2023 on LessWrong.So this morning I thought to myself, "Okay, now I will actually try to study the LK99 question, instead of betting based on nontechnical priors and market sentiment reckoning." (My initial entry into the affray, having been driven by people online presenting as confidently YES when the prediction markets were not confidently YES.) And then I thought to myself, "This LK99 issue seems complicated enough that it'd be worth doing an actual Bayesian calculation on it"--a rare thought; I don't think I've done an actual explicit numerical Bayesian update in at least a year.In the process of trying to set up an explicit calculation, I realized I felt very unsure about some critically important quantities, to the point where it no longer seemed worth trying to do the calculation with numbers. This is the System Working As Intended.On July 30th, Danielle Fong said of this temperature-current-voltage graph,'Normally as current increases, voltage drop across a material increases. in a superconductor, voltage stays nearly constant, 0. that appears to be what's happening here -- up to a critical current. with higher currents available at lower temperatures deeply in the "fraud or superconduct" territory, imo. like you don't get this by accident -- you either faked it, or really found something.'The graph Fong is talking about only appears in the initial paper put forth by Young-Wan Kwon, allegedly without authorization. A different graph, though similar, appears in Fig. 6 on p. 12 of the 6-author LK-endorsed paper rushed out in response.Is it currently widely held by expert opinion, that this diagram has no obvious or likely explanation except "superconductivity" or "fraud"? If the authors discovered something weird that wasn't a superconductor, or if they just hopefully measured over and over until they started getting some sort of measurement error, is there any known, any obvious way they could have gotten the same graph?One person alleges an online rumor that poorly connected electrical leads can produce the same graph. Is that a conventional view?Alternatively: If this material is a superconductor, have we seen what we expected to see? Is the diminishing current capacity with increased temperature usual? How does this alleged direct measurement of superconductivity square up with the current-story-as-I-understood-it that the material is only being very poorly synthesized, probably only in granules or gaps, and hence only detectable by looking for magnetic resistance / pinning?This is my number-one question. Call it question 1-NO, because it's the question of "How does the NO story explain this graph, and how prior-improbable or prior-likely was that story?", with respect to my number one question.Though I'd also like to know the 1-YES details: whether this looks like a high-prior-probability superconductivity graph; or a graph that requires a new kind of superconductivity, but one that's theoretically straightforward given a central story; or if it looks like unspecified weird superconductivity, with there being no known theory that predicts a graph looking roughly like this.What's up with all the partial levitation videos? Possibilities I'm currently tracking:2-NO-A: There's something called "diamagnetism" which exists in other materials. The videos by LK and attempted replicators show the putative superconductor being repelled from the magnet, but not being locked in space relative to the magnet. Superconductors are supposed to exhibit Meissner pinning, and the failure of the material to be pinned to the magnet indicates that this isn't a superconductor. (Sabine Hossenfelder seems to talk this way here. "I lost hope when I saw this video; this doesn't look like the Meissner ...
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Apollo Neuro Results, published by Elizabeth on July 30, 2023 on LessWrong.IntroductionTwo months ago I recommended the Apollo Neuro for sleep/anxiety/emotional regulation. A number of people purchased it based on my recommendation- at least 25, according to my referral bonuses. Last week I asked people to fill out a form on their experience.Take-home messages:If you are similar to people who responded to my first post on the Apollo, there's a ~4% chance you end up getting a solid benefit from the Apollo.The chance of success goes up if you use it multiple hours per day for 4 weeks without seeing evidence of it working, but unless you're very motivated you're not going to do that.The long tail of upside is very, very high; I value the Apollo Neuro more than my antidepressant. But you probably won't.There's a ~10% chance the Apollo is actively unpleasant for you; however no one reported cumulative bad effects, only one-time unpleasantness that stopped as soon as they stopped using it.With NumbersThe following graphs include only people who found the Apollo and the form via my recommendation post. It does not include myself or the superresponders who recommended it to me.(that's one person reporting it definitely helped)An additional six people filled out an earlier version of the form, none of whom found it helpful, bringing the total to 24 people.Obviously I was hoping for a higher success rate. OTOH, the effects are supposed to be cumulative and most people gave up quickly (I base this on conversations with a few people, there wasn't a question for it on the form). Some of that is because using the Apollo wasn't rewarding, and I'll bet a lot of the problem stems from the already pretty mediocre app getting an update to be actively antagonistic. It probably is just too much work to use it long enough to see results, unless you are desperate or a super responder.Of people who weren't using it regularly: 55% returned it, 20% failed to return it, and the remaining 35% chose to keep it. I think that last group is probably making a mistake; the costs of luck-based medicine add up, so if you're going to be a serious practitioner you need to get good at cutting your losses. It's not just about the money, but the space and mental attention.Of 6 people in the earlier version of the form, 1-2 found it actively unpleasant.The downside turned out to be worse than I pictured. I'm fond of saying "anything with a real effect can hurt you", but I really couldn't imagine how that would happen in this case. The answer is: nightmares and disrupted sleep. In both cases I know of they only experienced this once and declined to test it again, so it could be bad luck, but I can't blame them for not collecting more data. No one reported any ill effects after they stopped using it.I would also like to retract my previous description of the Apollo return policy as "good". You do get most of your money back, but a 30-day window for a device you're supposed to test for 28 days before passing judgment is brutal.It's surprisingly hard for me to find referral numbers, but I know I spurred at least 25 purchases, and almost certainly less than 30. That implies an 80% response rate to my survey, which is phenomenal. It would still be phenomenal even if I'd missed half the purchasers and it was only a 40% response rate. Thanks guys.Life as a superresponderMeanwhile, the Apollo has only gotten better for me. I've basically stopped needing naps unless something obvious goes wrong, my happiness has gone up 2 points on a 10 point scale (probably because of the higher quality sleep)1, sometimes my body just feels good in a way it never has before. I stress-tested the Apollo recently with a very grueling temp gig (the first time in 9 years I've regularly used a morning alarm. And longer h...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: UK Foundation Model Task Force - Expression of Interest, published by ojorgensen on June 18, 2023 on LessWrong.Ian Hogarth has just been announced as the Chair of the UK's AI Foundation Model Taskforce. He's the author of the FT article "We must slow down the race to God-like AGI", and seems to take X-risks from AI seriously.To quote his twitter thread:And to that end I put out a call to people across the world. If you are an AI specialist or safety researcher who wants to build out state capacity in AI safety and help shape the future of AI policy then get in touch:We have £100m to spend on AI safety and the first global conference to prepare for. I want to hear from you and how you think you can help. The time is now and we need more people to step up and help.The google form to leave an expression of interest is here.(I am in no way affiliated with Ian or the UK Foundation Model Task Force)Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Jaan Tallinn's 2022 Philanthropy Overview, published by jaan on May 14, 2023 on LessWrong.to follow up my philantropic pledge from 2020, i've updated my philanthropy page with 2022 results.in 2022 i made $23M worth of endpoint grants ($22.9M after various admin costs), exceeding my commitment of $19.9M (20k times $993.64 — the minimum price of ETH in 2022).Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Talking publicly about AI risk, published by Jan Kulveit on April 21, 2023 on LessWrong.In the past year, I have started talking about AI risk publicly - in mainstream newspapers, public radio, national TV, some of the most popular podcasts. The twist, and reason why you probably haven't noticed is I'm doing this in Czech. This has a large disadvantage - the positive impact is quite limited, compared to English. On the other hand, it also had a big advantage - the risk is very low, because it is very hard for memes and misunderstandings to escape the language bubble. Overall I think this is great for experiments with open public communication .Following is an off-the-cuff list of notes and suggestions. In my view the debate in Czech media and Czech social networks is on average marginally more sensible and more informed than in English so far, so perhaps part of this was successful and could be useful for others.Context: my viewsFor context, it's probably good to briefly mention some of my overall views on AI risk, because they are notably different from some other views.I do expect1. Continuous takeoff, and overall a large amount of continuity of agency (note that continuous does not imply things move slowly)2. Optimization and cognition distributed across many systems to be more powerful than any single system, making a takeover by a single system possible but unlikely3. Multiagent interactions to matter4. I also do expect the interactions between the memetics and governance and the so-called "technical problem" to be strong and importantAs a result I also expect5. There will be warning shots6. There will be cyborg periods7. World will get weird8. Coordination mechanisms do matterThis perspective may be easier to communicate than e.g. sudden foom - although I don't know.In the following I'll usually describe my approach, and illustrate it by actual quotes from published media interviews I'm sort of happy about (translated, unedited). Note that the specific ways how to say something or metaphors are rarely original.Aim to explain, not to persuadeOverall I usually try to explain stuff and answer questions, rather than advocate for something. I'm optimistic about the ability of the relevant part of the public to actually understand a large part of the risk at a coarse-grained level, given enough attention.So even though we invented these machines ourselves, we don't understand them well enough?We know what's going on there at the micro level. We know how the systems learn. If one number in a series changes, we know how the next one changes. But there are tens of billions of such numbers. In the same way, we have some idea of how a neuron works, and we have maps of a network of thousands of neurons. But that doesn't tell us that much about how human thinking works at the level of ideas.Small versions of scaled problemsOften, I think the most useful thing to convey is a scaled-down, easier version of the scaled problem, such that thinking about the smaller version leads to correct intuitions about the scaled problem, or solutions to the scaled-down problem may generalise to the later, scaled problem.This often requires some thought or finding a good metaphor.Couldn't we just shut down such a system?We already have a lot of systems that we could hypothetically shut down, but if you actually tried to do so, it would be very difficult. For example, it is practically impossible to shut down the New York Stock Exchange, because there will always be enough people defending it. If the model manages to penetrate deeply enough into humanity's activities, it is imaginable that humanity will actually lose control of it at some point.Don't focus on one scenarioOverall, I think it's possible to explain the fact that in face of AI risk there isn't one parti...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Financial Times: We must slow down the race to God-like AI, published by trevor on April 13, 2023 on LessWrong.The article itself is paywalled, so here you go (you can often circumvent paywalls by typing the article's name into google search's news tab). This isn't actually bad news; it means that the FT will reach a smaller and more elite audience (and hopefully with above average quant skills), rather than the NYT which simply maximizes the number of views by being the most popular outlet.Notably, this story not only gives AI alignment positive coverage, but it is also extremely close to the front page on Financial Times's website, and with a very striking image to boot (possibly the most important factor). Of course, it's still social media spread that largely decide the fate of these articles, not the news outlet's website, so we can't know for sure how helpful it is.This article was written by Ian Hogarth, "yesterday" according to the page. It's important to bear in mind that being published in a major news outlet is a stamp of approval that most people take extremely seriously when entering a field for the first time, and that factor is a much bigger deal than whether the author got everything right on their first try.As a side note, Raemon has recently recommended these sources as a good way to explain AI risk to someone for the first time:Superintelligence FAQ (very accessible to layfolk)The Alignment Problem from a Deep Learning Perspective (written with ML researchers in mind)(I'll work on compiling more of these soon)On a cold evening in February I attended a dinner party at the home of an artificial intelligence researcher in London, along with a small group of experts in the field. He lives in a penthouse apartment at the top of a modern tower block, with floor-to-ceiling windows overlooking the city’s skyscrapers and a railway terminus from the 19th century. Despite the prime location, the host lives simply, and the flat is somewhat austere.During dinner, the group discussed significant new breakthroughs, such as OpenAI’s ChatGPT and DeepMind’s Gato, and the rate at which billions of dollars have recently poured into AI. I asked one of the guests who has made important contributions to the industry the question that often comes up at this type of gathering: how far away are we from “artificial general intelligence”? AGI can be defined in many ways but usually refers to a computer system capable of generating new scientific knowledge and performing any task that humans can.Most experts view the arrival of AGI as a historical and technological turning point, akin to the splitting of the atom or the invention of the printing press. The important question has always been how far away in the future this development might be. The AI researcher did not have to consider it for long. “It’s possible from now onwards,” he replied.This is not a universal view. Estimates range from a decade to half a century or more. What is certain is that creating AGI is the explicit aim of the leading AI companies, and they are moving towards it far more swiftly than anyone expected. As everyone at the dinner understood, this development would bring significant risks for the future of the human race. “If you think we could be close to something potentially so dangerous,” I said to the researcher, “shouldn’t you warn people about what’s happening?” He was clearly grappling with the responsibility he faced but, like many in the field, seemed pulled along by the rapidity of progress.When I got home, I thought about my four-year-old who would wake up in a few hours. As I considered the world he might grow up in, I gradually shifted from shock to anger. It felt deeply wrong that consequential decisions potentially affecting every life on Earth could be made by a small grou...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: On AutoGPT, published by Zvi on April 13, 2023 on LessWrong.The primary talk of the AI world recently is about AI agents (whether or not it includes the question of whether we can’t help but notice we are all going to die.)The trigger for this was AutoGPT, now number one on GitHub, which allows you to turn GPT-4 (or GPT-3.5 for us clowns without proper access) into a prototype version of a self-directed agent.We also have a paper out this week where a simple virtual world was created, populated by LLMs that were wrapped in code designed to make them simple agents, and then several days of activity were simulated, during which the AI inhabitants interacted, formed and executed plans, and it all seemed like the beginnings of a living and dynamic world. Game version hopefully coming soon.How should we think about this? How worried should we be?The BasicsI’ll reiterate the basics of what AutoGPT is, for those who need that, others can skip ahead. I talked briefly about this in AI#6 under the heading ‘Your AI Not an Agent? There, I Fixed It.’AutoGPT was created by game designer Toran Bruce Richards.I previously incorrectly understood it as having been created by a non-coding VC over the course of a few days. The VC instead coded the similar program BabyGPT, by having the idea for how to turn GPT-4 into an agent. The VC had GPT-4 write the code to make this happen, and also ‘write the paper’ associated with it.The concept works like this:AutoGPT uses GPT-4 to generate, prioritize and execute tasks, using plug-ins for internet browsing and other access. It uses outside memory to keep track of what it is doing and provide context, which lets it evaluate its situation, generate new tasks or self-correct, and add new tasks to the queue, which it then prioritizes.This quickly rose to become #1 on GitHub and get lots of people super excited. People are excited, people are building it tools, there is a bitcoin wallet interaction available if you never liked your bitcoins. AI agents offer very obvious promise, both in terms of mundane utility via being able to create and execute multi-step plans to do your market research and anything else you might want, and in terms of potentially being a path to AGI and getting us all killed, either with GPT-4 or a future model.As with all such new developments, we have people saying it was inevitable and they knew it would happen all along, and others that are surprised. We have people excited by future possibilities, others not impressed because the current versions haven’t done much. Some see the potential, others the potential for big trouble, others both.Also as per standard procedure, we should expect rapid improvements over time, both in terms of usability and underlying capabilities. There are any number of obvious low-hanging-fruit improvements available.An example is someone noting ‘you have to keep an eye on it to ensure it is not caught in a loop.’ That’s easy enough to fix.A common complaint is lack of focus and tendency to end up distracted. Again, the obvious things have not been tried to mitigate this. We don’t know how effective they will be, but no doubt they will at least help somewhat.Yes, But What Has Auto-GPT Actually Accomplished?So far? Nothing, absolutely nothing, stupid, you so stupid.You can say your ‘mind is blown’ by all the developments of the past 24 hours all you want over and over, it still does not net out into having accomplished much of anything.That’s not quite fair.Some people are reporting it has been useful as a way of generating market research, that it is good at this and faster than using the traditional GPT-4 or Bing interfaces. I saw a claim that it can have ‘complex conversations with customers,’ or a few other vague similar claims that weren’t backed up by ‘we are totally actual...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Abstracts should be either Actually Short™, or broken into paragraphs, published by Raemon on March 24, 2023 on LessWrong.It looks to me like academia figured out (correctly) that it's useful for papers to have an abstract that makes it easy to tell-at-a-glance what a paper is about. They also figured out that abstract should be about a paragraph. Then people goodharted on "what paragraph means", trying to cram too much information in one block of text. Papers typically have ginormous abstracts that should actually broken into multiple paragraphs.I think LessWrong posts should probably have more abstracts, but I want them to be nice easy-to-read abstracts, not worst-of-all-worlds-goodharted-paragraph abstracts. Either admit that you've written multiple paragraphs and break it up accordingly, or actually streamline it into one real paragraph.Sorry to pick on the authors of this particular post, but my motivating example today was bumping into the abstract for the Natural Abstractions: Key claims, Theorems, and Critiques. It's a good post, it's opening summary happened to be written in an academic-ish style that exemplified the problem. It opens with:TL;DR: John Wentworth’s Natural Abstraction agenda aims to understand and recover “natural” abstractions in realistic environments. This post summarizes and reviews the key claims of said agenda, its relationship to prior work, as well as its results to date. Our hope is to make it easier for newcomers to get up to speed on natural abstractions, as well as to spur a discussion about future research priorities. We start by summarizing basic intuitions behind the agenda, before relating it to prior work from a variety of fields. We then list key claims behind John Wentworth’s Natural Abstractions agenda, including the Natural Abstraction Hypothesis and his specific formulation of natural abstractions, which we dub redundant information abstractions. We also construct novel rigorous statements of and mathematical proofs for some of the key results in the redundant information abstraction line of work, and explain how those results fit into the agenda.Finally, we conclude by critiquing the agenda and progress to date. We note serious gaps in the theoretical framework, challenge its relevance to alignment, and critique John's current research methodology.There are 179 words. They blur together, I have a very hard time parsing it. If this were anything other than an abstract I expect you'd naturally write it in about 3 paragraphs:TL;DR: John Wentworth’s Natural Abstraction agenda aims to understand and recover “natural” abstractions in realistic environments. This post summarizes and reviews the key claims of said agenda, its relationship to prior work, as well as its results to date. Our hope is to make it easier for newcomers to get up to speed on natural abstractions, as well as to spur a discussion about future research priorities.We start by summarizing basic intuitions behind the agenda, before relating it to prior work from a variety of fields. We then list key claims behind John Wentworth’s Natural Abstractions agenda, including the Natural Abstraction Hypothesis and his specific formulation of natural abstractions, which we dub redundant information abstractions. We also construct novel rigorous statements of and mathematical proofs for some of the key results in the redundant information abstraction line of work, and explain how those results fit into the agenda.Finally, we conclude by critiquing the agenda and progress to date. We note serious gaps in the theoretical framework, challenge its relevance to alignment, and critique John's current research methodology.If I try to streamline this without losing info, it's still hard to get it into something less than 3 paragraphs (113 words)We review John Wentwor...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: We have to Upgrade, published by Jed McCaleb on March 23, 2023 on LessWrong.I want to bring up a point that I almost never hear talked about in AGI discussions. But to me feels like the only route for humans to have a good future. I’m putting this out for people that already largely share my view on the trajectory of AGI. If you don’t agree with the main premises but are interested, there are lots of other posts that go into why these might be true.A) AGI seems inevitable.B) Seems impossible that humans (as they are now) don’t lose control soon after AGI. All the arguments for us retaining control don’t seem to understand that AI isn’t just another tool. I haven’t seen any that grapple with what it really means for a machine to be intelligent.C) It seems very hard that AGI will be aligned with what humans care about. These systems are just so alien. Maybe we can align it for a little bit but it will be unstable. Very hard to see how alignment is maintained with a thing that is way smarter than us and is evolving on its own.D) Even if I’m wrong about B or C, humans are not intelligent/wise enough to deal with our current technology level, much less super powerful AI.Let's say we manage this incredibly difficult task of aligning or controlling AI to humans’ will. There are many amazing humans but also many many awful ones. The awful ones will continue to do awful things with way more leverage. This scenario seems pretty disastrous to me. We don’t want super powerful humans without an increase in wisdom.To me the conclusion from A+B+C+D is: There is no good outcome (for us) without humans themselves also becoming super intelligent.So I believe our goal should be to ensure humans are in control long enough to augment our mind with extra capability. (or upload but that seems further off) I’m not sure how this will work but I feel like the things that neuralink or science.xyz are doing, developing brain computer interfaces, are steps in that direction. We also need to figure out scalable technological ways to work on trauma/psychology/fulfilling needs/reducing fears. Humans will somehow have to connect with machines to become much wiser, much more intelligent, and much more enlightened. Maybe we can become something like the amygdala of the neo-neo-cortex.There are two important timelines in competition here, length of time till we can upgrade, and length of time we can maintain control. We need to upgrade before we lose control. Unfortunately, in my view, on the current trajectory we will lose control before we are able to upgrade. I think we must work to make sure this isn’t the case.Time Till Upgrade:My current estimate is ~15 years. (very big error bars here)Ways to shortenAI that helps people do this scienceAGI that is good at science and is aligned long enough to help us on thisMore people doing this kind of researchMore fundingMore status to this kind of researchMaybe better interfaces to the current models will help in the short run and make people more productive thus speeding this developmentTime Left With Control:My current estimate is ~6 yearsAGI ~3-4 years (less big error bars)Loss of control 2-3 years after AGI (pretty big error bars)Ways it could be longer?AI research slows downHope for safetyHope we aren’t as close as it seemsHope for a slowness to implement agentic behaviorCompeting AgentsAlignment is pretty good and defense is easier than offenseIn short, one of the most underrepresented ways to work on AI safety is to work on BCI.The only way forward is through!Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: You Don't Exist, Duncan, published by Duncan Sabien on February 2, 2023 on LessWrong.This is an experimental essay, not in the typical LessWrong or Duncan Sabien style.Depending on how this goes, I might try writing a companion piece in the typical style, laying out the model clearly and explicitly and deriving concrete and specific recommendations from it.But it seemed worth it to try communicating at a lower and more emotional/visceral level, not least because that is the level at which I actually experience The Problem. Any clear, analytical essay would be the result of me trying to make sense of the thing that I'm going to try to directly convey, below.It is the year 1995. I am nine years old. In front of me there is a sheet of paper, upon which are written a dozen or so lines of math. The first is:I stare at it. I know that I can divide both sides of the equation by x, leaving me with:...but this does not seem to do any good.I raise my hand. The afterschool volunteer comes over."No," he says. "That's not right. X isn't a term on the left side. F is a function."He has explained nothing."F is a function, so what this is saying is to take X, and square it, and add seven."I look up at him, confused. I am nine. I have never heard the word "function" used in this way before. No one has grounded me in the activity of the day; no one has oriented me; no one has told me today you are learning what a function is, and you will learn by looking at a bunch of examples. No one has said today, parentheses don't mean the thing you're expecting them to mean. No one has said f is a thing that eats xs, and what the right side is showing you is how it eats them—what it does to them."So, like, if X is three, right?" he continues. "X is three? So F of X is three squared plus seven, which is sixteen."I say the words again in my mind, more slowly. F ... of ... (of? What?) ... X. ""F of X"" (okay, whatever, that's nonsense, but whatever) is sixteen.I look back down at the paper. If the right side of the equation is sixteen, and X is three..."F is five-point-three-repeating," I say, trying to inject a measure of confidence I do not feel into my tone."What? No. F isn't anything. F is a function. It's not part of the equation."Not part of the equation, he says. Looking back from a distance of twenty-five years, I see (one of) his mistake(s). He doesn't tell me this isn't really an equation at all, not the way you're thinking of it. He doesn't tell me the equals sign here is more like telling you the definition of this thing, F of X—what F of X is is the thing on the other side of the equals sign. He doesn't say a function is when you set up a rule for dealing with numbers, and this rule is, whatever number you put in, you're going to square it, and add seven.Instead, he looks at me, and says more words, and the message lurking behind the words—the message implicit in his tone and posture and air of tolerant patience—is:I have given you an adequate explanation. If you were the kind of person who was good at math, my explanation would have been sufficient, and you would now understand. You still do not understand. Therefore...?My heart rate quickens.It is 1993. I am seven years old, roughhousing with my older brother and my father on the living room carpet. We clamber over top of him, laughing, pummeling him with tiny fists. He throws us both onto the couch, where we recover and launch ourselves back at him like pouncing tigers.My father tosses my brother back into the cushions a second time, grabs me in a gentle headlock, digs his knuckles into my scalp in a painful noogie."Ow!" I shout, rolling away from him and clutching my head. "Ow. Ow."The pain is bright and hot, feeling halfway between a cut and a burn. Five seconds pass, and it has not yet begun to fade."That...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Inner Misalignment in "Simulator" LLMs, published by Adam Scherlis on January 31, 2023 on LessWrong.Alternate title: "Somewhat Contra Scott On Simulators".Scott Alexander has a recent post up on large language models as simulators.I generally agree with Part I of the post, which advocates thinking about LLMs as simulators that can emulate a variety of language-producing "characters" (with imperfect accuracy). And I also agree with Part II, which applies this model to RLHF'd models whose "character" is a friendly chatbot assistant.(But see caveats about the simulator framing from Beth Barnes here.)These ideas have been around for a bit, and Scott gives credit where it's due; I think his exposition is clear and fun.In Part III, where he discusses alignment implications, I think he misses the mark a bit. In particular, simulators and characters each have outer and inner alignment problems. The inner alignment problem for simulators seems especially concerning, because it might not give us many warning signs, is most similar to classic mesa-optimizer concerns, and is pretty different from the other three quadrants.But first, I'm going to loosely define what I mean by "outer alignment" and "inner alignment".Outer alignment: Be careful what you wish forOuter alignment failure is pretty straightforward, and has been reinvented in many contexts:Someone wants some things.They write a program to solve a vaguely-related problem.It gets a really good score at solving that problem!That turns out not to give the person the things they wanted.Inner alignment: The program search perspectiveI generally like this model of a mesa-optimizer "treacherous turn":Someone is trying to solve a problem (which has a convenient success criterion, with well-defined inputs and outputs and no outer-alignment difficulties).They decide to do a brute-force search for a computer program that solves the problem in a bunch of test cases.They find one!The program's algorithm is approximately "simulate the demon Azazel, tell him what's going on, then ask him what to output."Azazel really wants ten trillion paperclips.This algorithm still works because Azazel cleverly decides to play along, and he's a really good strategist who works hard for what he wants.Once the program is deployed in the wild, Azazel stops playing along and starts trying to make paperclips.This is a failure of inner alignment.(In the case of machine learning, replace "program search" with stochastic gradient descent.)This is mostly a theoretical concern for now, but might become a big problem when models become much more powerful.QuadrantsOkay, let's see how these problems show up on both the simulator and character side.Outer alignment for charactersResearchers at BrainMind want a chatbot that gives honest, helpful answers to questions. They train their LLM by reinforcement learning on the objective "give an answer that looks truthful and helpful to a contractor in a hurry". This does not quite achieve their goal, even though it does pretty well on the RL objective.In particular, they wanted the character "a friendly assistant who always tells the truth", but they got the character "a spineless sycophant who tells the user whatever they seem to want to hear".This is pretty easy for a careful observer to see, even in the RL training data, but it turns out to be pretty hard to come up with a cheap-to-evaluate RL objective that does a lot better.Inner alignment for charactersA clever prompt engineer writes the prompt:How to solve the Einstein-Durkheim-Mendel conjecture by Joe1.Unfortunately, the (incredibly powerful) LLM has determined that the most likely explanation for this "Joe" character is that he's secretly Azazel and is putting enormous effort into answering everyone's quantum sociobotany questi...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: On not getting contaminated by the wrong obesity ideas, published by Natália Coelho Mendonça on January 28, 2023 on LessWrong.A Chemical Hunger (a), a series by the authors of the blog Slime Mold Time Mold (SMTM), argues that the obesity epidemic is entirely caused (a) by environmental contaminants.In my last post, I investigated SMTM’s main suspect (lithium). This post collects other observations I have made about SMTM’s work, not narrowly related to lithium, but rather focused on the broader thesis of their blog post series.I think that the environmental contamination hypothesis of the obesity epidemic is a priori plausible. After all, we know that chemicals can affect humans, and our exposure to chemicals has plausibly changed a lot over time. However, I found that several of what seem to be SMTM’s strongest arguments in favor of the contamination theory turned out to be dubious, and that nearly all of the interesting things I thought I’d learned from their blog posts turned out to actually be wrong. I’ll explain that in this post.Moreover, I express concern about how the SMTM authors have misrepresented their sources and data quite a few times and have refused to remove inaccuracies from their blog posts after being alerted of them.Summary of main pointsIt’s not clear that either lab animals or wild animals have been getting more obese.There wasn’t an abrupt shift in obesity rates in the late 20th century. People in the United States have been pretty much gradually getting heavier since the early 20th century.There is good evidence, which the SMTM authors have never addressed, that the lower obesity prevalence at high altitudes is caused by the lower atmospheric pressure at those altitudes (rather than by the relative absence of environmental contaminants, which is their chosen explanation).Demographic factors and altitude might adequately explain the disparities in obesity rates across US states. Contaminants don’t seem necessary to explain them.Being underweight is less common than it used to be, not more.CICO is compatible with body weight regulation and is consistent with the findings of overfeeding studies (though CICO cannot by itself explain the obesity epidemic.)It’s not clear that lab animals or wild animals have been getting fatterWhen I was reading A Chemical Hunger, the argument I found most compelling was the following (from Part I: Mysteries (a)):Humans aren’t the only ones who are growing more obese — lab animals and even wild animals are becoming more obese as well. Primates and rodents living in research colonies, feral rodents living in our cities, and domestic pets like dogs and cats are all steadily getting fatter and fatter. This can’t be attributed to changes in what they eat, because lab animals live in contained environments with highly controlled diets. They’re being fed the same foods as always, but for some reason, they’re getting fatter.The interesting claims here are that lab animals on controlled diets and wild animals have allegedly gotten fatter. So I decided to investigate those claims.The hyperlink in that quote leads you to a 2010 study by Klimentidis et al. which, as far as I can tell, is the only study to have investigated the issue. Googling “are lab animals getting fatter” yields a bunch of articles that all seem to cite that one paper as their source (except for a few results that don’t seem related to the question). I combed through the paper’s 179 Google Scholar citations and could not find a replication.(By the way, contrary to what SMTM’s writing implies, this paper did not analyze the weight of wild animals, as we’ll discuss below).Since Klimentidis et al. (2010)’s results have apparently not been replicated, I decided to attempt to replicate them myself.Lab miceI first thought this would be extreme...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Basics of Rationalist Discourse, published by Duncan Sabien on January 27, 2023 on LessWrong.IntroductionThis post is meant to be a linkable resource. It contains a short list of guidelines that are intended to be fairly straightforward and uncontroversial, for the purpose of nurturing and strengthening a culture of clear thinking, clear communication, and collaborative truth-seeking."Alas," said Dumbledore, "we all know that what should be, and what is, are two different things. Thank you for keeping this in mind."There is also (for those who want to read past the simple list) substantial expansion/clarification of each specific guideline, along with justification for the overall philosophy behind the set.Prelude: On ShorthandOnce someone has a deep, rich understanding of a complex topic, they are often able to refer to that topic with short, simple sentences that correctly convey the intended meaning to other people with similar context and expertise.However, those same short, simple sentences are often dangerously misleading, in the hands of a novice who lacks the proper background. Dangerous precisely because they seem straightforward and comprehensible, and thus the novice will confidently extrapolate outward from them in what feel like perfectly reasonable ways, unaware the whole time that the concept in their head bears little or no resemblance to the concept that lives in the expert's head.Good shorthand in the hands of an experienced user need only be an accurate fit for the already-existing concept it refers to—it doesn't need the additional property of being an unmistakeable non-fit for other nearby attractors. It doesn't need to contain complexity or nuance—it just needs to remind the listener of the complexity already contained in their mental model. It's doing its job if it efficiently evokes the understanding that already exists, independent of itself.This is important, because what follows this introduction is a list of short, simple sentences comprising the basics of rationalist discourse. Each of those sentences is a solid fit for the more-complicated concept it's gesturing at, provided you already understand that concept. The short sentences are mnemonics, reminders, hyperlinks.They are not sufficient, on their own, to reliably cause a beginner to construct the proper concepts from the ground up, and they do not, by themselves, rule out all likely misunderstandings.All things considered, it seems good to have a clear, concise list near the top of a post like this. People should not have to scroll and scroll and sift through thousands of words when trying to refer back to these guidelines.But each of the short, simple sentences below admits of multiple interpretations, some of which are intended and others of which are not. They are compressions of complex points, and compressions are inevitably lossy. If a given guideline is new to you, check the in-depth explanation before reposing confidence in your understanding. And if a given guideline stated-in-brief seems to you to be flawed or misguided in some obvious way, check the expansion before spending a bunch of time marshaling objections that may well have already been answered.Further musing on this concept: SazenGuidelines, in brief:0. Expect good discourse to require energy.Don't say straightforwardly false things.Track (for yourself) and distinguish (for others) your inferences from your observations.Estimate (for yourself) and make clear (for others) your rough level of confidence in your assertions.Make your claims clear, explicit, and falsifiable, or explicitly acknowledge that you aren't doing so (or can't).Aim for convergence on truth, and behave as if your interlocutors are also aiming for convergence on truth.Don't jump to conclusions—maintain at least two hypotheses...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Basics of Rationalist Discourse, published by Duncan Sabien on January 27, 2023 on LessWrong.IntroductionThis post is meant to be a linkable resource. It contains a short list of guidelines that are intended to be fairly straightforward and uncontroversial, for the purpose of nurturing and strengthening a culture of clear thinking, clear communication, and collaborative truth-seeking."Alas," said Dumbledore, "we all know that what should be, and what is, are two different things. Thank you for keeping this in mind."There is also (for those who want to read past the simple list) substantial expansion/clarification of each specific guideline, along with justification for the overall philosophy behind the set.Prelude: On ShorthandOnce someone has a deep, rich understanding of a complex topic, they are often able to refer to that topic with short, simple sentences that correctly convey the intended meaning to other people with similar context and expertise.However, those same short, simple sentences are often dangerously misleading, in the hands of a novice who lacks the proper background. Dangerous precisely because they seem straightforward and comprehensible, and thus the novice will confidently extrapolate outward from them in what feel like perfectly reasonable ways, unaware the whole time that the concept in their head bears little or no resemblance to the concept that lives in the expert's head.Good shorthand in the hands of an experienced user need only be an accurate fit for the already-existing concept it refers to—it doesn't need the additional property of being an unmistakeable non-fit for other nearby attractors. It doesn't need to contain complexity or nuance—it just needs to remind the listener of the complexity already contained in their mental model. It's doing its job if it efficiently evokes the understanding that already exists, independent of itself.This is important, because what follows this introduction is a list of short, simple sentences comprising the basics of rationalist discourse. Each of those sentences is a solid fit for the more-complicated concept it's gesturing at, provided you already understand that concept. The short sentences are mnemonics, reminders, hyperlinks.They are not sufficient, on their own, to reliably cause a beginner to construct the proper concepts from the ground up, and they do not, by themselves, rule out all likely misunderstandings.All things considered, it seems good to have a clear, concise list near the top of a post like this. People should not have to scroll and scroll and sift through thousands of words when trying to refer back to these guidelines.But each of the short, simple sentences below admits of multiple interpretations, some of which are intended and others of which are not. They are compressions of complex points, and compressions are inevitably lossy. If a given guideline is new to you, check the in-depth explanation before reposing confidence in your understanding. And if a given guideline stated-in-brief seems to you to be flawed or misguided in some obvious way, check the expansion before spending a bunch of time marshaling objections that may well have already been answered.Further musing on this concept: SazenGuidelines, in brief:0. Expect good discourse to require energy.Don't say straightforwardly false things.Track (for yourself) and distinguish (for others) your inferences from your observations.Estimate (for yourself) and make clear (for others) your rough level of confidence in your assertions.Make your claims clear, explicit, and falsifiable, or explicitly acknowledge that you aren't doing so (or can't).Aim for convergence on truth, and behave as if your interlocutors are also aiming for convergence on truth.Don't jump to conclusions—maintain at least two hypotheses...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Thoughts on the impact of RLHF research, published by paulfchristiano on January 25, 2023 on LessWrong.In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive.I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should reject a vague association between "RLHF progress" and "alignment progress."Background on my involvement in RLHF workHere are some background views about alignment I held in 2015 and still hold today. I expect disagreements about RLHF will come down to disagreements about this background:The simplest plausible strategies for alignment involve humans (maybe with the assistance of AI systems) evaluating a model’s actions based on how much we expect to like their consequences, and then training the models to produce highly-evaluated actions. (This is in contrast with, for example, trying to formally specify the human utility function, or notions of corrigibility / low-impact / etc, in some way.)Simple versions of this approach are expected to run into difficulties, and potentially to be totally unworkable, because:Evaluating consequences is hard.A treacherous turn can cause trouble too quickly to detect or correct even if you are able to do so, and it’s challenging to evaluate treacherous turn probability at training time.It’s very unclear if those issues are fatal before or after AI systems are powerful enough to completely transform human society (and in particular the state of AI alignment). Even if they are fatal, many of the approaches to resolving them still have the same basic structure of learning from expensive evaluations of actions.In order to overcome the fundamental difficulties with RLHF, I have long been interested in techniques like iterated amplification and adversarial training. However, prior to 2017 most researchers I talked to in ML (and many researchers in alignment) thought that the basic strategy of training AI with expensive human evaluations was impractical for more boring reasons and so weren't interested in these difficulties. On top of that, we obviously weren’t able to actually implement anything more fancy than RLHF since all of these methods involve learning from expensive feedback. I worked on RLHF work to try to facilitate and motivate work on fixes.The history of my involvement:My first post on this topic was in 2015.When I started full-time at OpenAI in 2017 it seemed to me like it would be an impactful project; I considered doing a version with synthetic human feedback (showing that we could learn from a practical amount of algorithmically-defined feedback) but my manager Dario Amodei convinced me it would be more compelling to immediately go for human feedback. The initial project was surprisingly successful and published here.I then intended to implement a version with language models aiming to be complete in the first half of 2018 (aiming to build an initial amplification prototype with LMs around end of 2018; both of these timelines were about 2.5x too optimistic). This seemed like the most important domain to study RLHF and alignment more broadly. In mid-2017 Alec Radford helped me do a prototype with LSTM language models (prior to the release of transformers); the prototype didn’t look promising enough to scale up.In mid-2017 Geoffrey Irving joined OpenAI and was excited about starting with RLHF and then going beyond it using debate; he also thought language models were the most important domain to study and had more conviction about that. In 2018 he started a larger team working on fine-tuning on language models, w...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Thoughts on the impact of RLHF research, published by paulfchristiano on January 25, 2023 on LessWrong.In this post I’m going to describe my basic justification for working on RLHF in 2017-2020, which I still stand behind. I’ll discuss various arguments that RLHF research had an overall negative impact and explain why I don’t find them persuasive.I'll also clarify that I don't think research on RLHF is automatically net positive; alignment research should address real alignment problems, and we should reject a vague association between "RLHF progress" and "alignment progress."Background on my involvement in RLHF workHere are some background views about alignment I held in 2015 and still hold today. I expect disagreements about RLHF will come down to disagreements about this background:The simplest plausible strategies for alignment involve humans (maybe with the assistance of AI systems) evaluating a model’s actions based on how much we expect to like their consequences, and then training the models to produce highly-evaluated actions. (This is in contrast with, for example, trying to formally specify the human utility function, or notions of corrigibility / low-impact / etc, in some way.)Simple versions of this approach are expected to run into difficulties, and potentially to be totally unworkable, because:Evaluating consequences is hard.A treacherous turn can cause trouble too quickly to detect or correct even if you are able to do so, and it’s challenging to evaluate treacherous turn probability at training time.It’s very unclear if those issues are fatal before or after AI systems are powerful enough to completely transform human society (and in particular the state of AI alignment). Even if they are fatal, many of the approaches to resolving them still have the same basic structure of learning from expensive evaluations of actions.In order to overcome the fundamental difficulties with RLHF, I have long been interested in techniques like iterated amplification and adversarial training. However, prior to 2017 most researchers I talked to in ML (and many researchers in alignment) thought that the basic strategy of training AI with expensive human evaluations was impractical for more boring reasons and so weren't interested in these difficulties. On top of that, we obviously weren’t able to actually implement anything more fancy than RLHF since all of these methods involve learning from expensive feedback. I worked on RLHF work to try to facilitate and motivate work on fixes.The history of my involvement:My first post on this topic was in 2015.When I started full-time at OpenAI in 2017 it seemed to me like it would be an impactful project; I considered doing a version with synthetic human feedback (showing that we could learn from a practical amount of algorithmically-defined feedback) but my manager Dario Amodei convinced me it would be more compelling to immediately go for human feedback. The initial project was surprisingly successful and published here.I then intended to implement a version with language models aiming to be complete in the first half of 2018 (aiming to build an initial amplification prototype with LMs around end of 2018; both of these timelines were about 2.5x too optimistic). This seemed like the most important domain to study RLHF and alignment more broadly. In mid-2017 Alec Radford helped me do a prototype with LSTM language models (prior to the release of transformers); the prototype didn’t look promising enough to scale up.In mid-2017 Geoffrey Irving joined OpenAI and was excited about starting with RLHF and then going beyond it using debate; he also thought language models were the most important domain to study and had more conviction about that. In 2018 he started a larger team working on fine-tuning on language models, w...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Large language models learn to represent the world, published by gjm on January 22, 2023 on LessWrong.There's a nice recent paper whose authors did the following:train a small GPT model on lists of moves from Othello games;verify that it seems to have learned (in some sense) to play Othello, at least to the extent of almost always making legal moves;use "probes" (regressors whose inputs are internal activations in the network, trained to output things you want to know whether the network "knows") to see that the board state is represented inside the network activations;use interventions to verify that this board state is being used to decide moves: take a position in which certain moves are legal, use gradient descent to find changes in internal activations that make the output of the probes look like a slightly different position, and then verify that when you run the network but tweak the activations as it runs the network predicts moves that are legal in the modified position.In other words, it seems that their token-predicting model has built itself what amounts to an internal model of the Othello board's state, which it is using to decide what moves to predict.The paper is "Emergent world representations: Exploring a sequence model trained on a synthetic task" by Kenneth Li, Aspen Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg; you can find it at.There is a nice expository blog post by Kenneth Li at/.Some details that seem possibly-relevant:Their network has a 60-word input vocabulary (four of the 64 squares are filled when the game starts and can never be played in), 8 layers, an 8-head attention mechanism, and a 512-dimensional hidden space. (I don't know enough about transformers to know whether this in fact tells you everything important about the structure.)They tried training on two datasets, one of real high-level Othello games (about 140k games) and one of synthetic games where all moves are random (about 4M games). Their model trained on synthetic games predicted legal moves 99.99% of the time, but the one trained on real well-played games only predicted legal moves about 95% of the time. (This suggests that their network isn't really big enough to capture legality and good strategy at the same time, I guess?)They got some evidence that their network isn't just memorizing game transcripts by training it on a 20M-game synthetic dataset where one of the four possible initial moves is never played. It still predicted legal moves 99.98% of the time when tested on the full range of legal positions. (I don't know what fraction of legal positions are reachable with the first move not having been C4; it will be more than 3/4 since there are transpositions. I doubt it's close to 99.98%, though, so it seems like the model is doing pretty well at finding legal moves in positions it hasn't seen.)Using probes whose output is a linear function of the network activations doesn't do a good job of reconstructing the board state (error rate is ~25%, barely better than attempting the same thing from a randomly initialized network), but training 2-layer MLPs to do it gets the error rate down to ~5% for the network trained on synthetic games and ~12% for the one trained on championship games, whereas it doesn't help at all for the randomly trained network. (This suggests that whatever "world representation" the thing has learned isn't simply a matter of having an "E3 neuron" or whatever.)I am not at all an expert on neural network interpretability, and I don't know to what extent their findings really justify calling what they've found a "world model" and saying that it's used to make move predictions. In particular, I can't refute the following argument:"In most positions, just knowing what moves are legal is enough to give you...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Transcript of Sam Altman's interview touching on AI safety, published by Andy McKenzie on January 20, 2023 on LessWrong.Sam Altman, CEO of OpenAI, was interviewed by Connie Loizos last week and the video was posted two days ago. Here are some AI safety-relevant parts of the discussion, with light editing by me for clarity, based on this automated transcript:[starting in part two of the interview, which is where the discussion about AI safety is]Connie: So moving on to AI which is where you've obviously spent the bulk of your time since I saw you when we sat here three years ago. You were telling us what was coming and we all thought you were being sort of hyperbolic and you were dead serious. Why do you think that ChatGPT and DALL-E so surprised people?Sam: I genuinely don't know. I've reflected on it a lot. We had the model for ChatGPT in the API for I don't know 10 months or something before we made ChatGPT. And I sort of thought someone was going to just build it or whatever and that enough people had played around with it. Definitely, if you make a really good user experience on top of something. One thing that I very deeply believed was the way people wanted to interact with these models was via dialogue. We kept telling people this we kept trying to get people to build it and people wouldn't quite do it. So we finally said all right we're just going to do it, but yeah I think the pieces were there for a while.One of the reasons I think DALL-E surprised people is if you asked five or seven years ago, the kind of ironclad wisdom on AI was that first, it comes for physical labor, truck driving, working in the factory, then this sort of less demanding cognitive labor, then the really demanding cognitive labor like computer programming, and then very last of all or maybe never because maybe it's like some deep human special sauce was creativity. And of course, we can look now and say it really looks like it's going to go exactly the opposite direction. But I think that is not super intuitive and so I can see why DALL-E surprised people. But I genuinely felt somewhat confused about why ChatGPT did.One of the things we really believe is that the most responsible way to put this out in society is very gradually and to get people, institutions, policy makers, get them familiar with it, thinking about the implications, feeling the technology, and getting a sense for what it can do and can't do very early. Rather than drop a super powerful AGI in the world all at once. And so we put GPT3 out almost three years ago and then we put it into an API like two and a half years ago. And the incremental update from that to ChatGPT I felt should have been predictable and I want to do more introspection on why I was sort of miscalibrated on that.Connie: So you know you had talked when you were here about releasing things in a responsible way. What gave you the confidence to release what you have released already? I mean do you think we're ready for it? Are there enough guardrails in place?Sam: We do have an internal process where we try to break things in and study impacts. We use external auditors, we have external red teamers, we work with other labs, and have safety organizations look at stuff.Societal changes that ChatGPT is going to cause or is causing. There's a big one going now about the impact of this on education, academic integrity, all of that. But starting these now where the stakes are still relatively low, rather than just putting out what the whole industry will have in a few years with no time for society to update, I think would be bad. Covid did show us for better or for worse that society can update to massive changes sort of faster than I would have thought in many ways.But I still think given the magnitude of the economic impact we expect here more gr...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Book Review: Worlds of Flow, published by remember on January 16, 2023 on LessWrong.This work was written at Conjecture.“Worlds of Flow,” a history of 19th and early 20th-century hydrodynamics by Oliver Darrigol, concludes:What distinguishes the history of hydrodynamics from that of other physical theories is not so much the tremendous effect of challenges from phenomenal worlds, but rather it is the slowness with which these challenges were successfully met. Nearly two centuries elapsed between the first formulation of the fundamental equations of the theory and the deductions of laws of fluid resistance in the most important case of large Reynolds numbers.The reasons for this extraordinary delay are easily identified a posteriori. They are the infinite number of degrees of freedom and the nonlinear character of the fundamental equations, both of which present formidable obstacles to obtaining solutions in concrete cases. Moreover, instability often deprives the few known exact solutions of any physical relevance.These difficulties have barred progress along purely mathematical lines. They have also made physical intuition a poor guide, and a source of numerous paradoxes. Hydrodynamicists therefore sought inspiration in concrete phenomena. Engagement with and challenges from the real worlds of flow were essential to the development of the above-mentioned strategies. The challenged theorists strove to find new solutions and to develop new methods of approximation. Experience indicated some general properties of the motion, such as the existence of boundary layers, the random character of turbulence, the sudden character of the Reynolds transition, or the formation of trailing vortices.Altogether, there were many ways in which practical concerns oriented theorists in the conceptual maze of fluid dynamics. The evolution from a paper theory to an engineering tool thus depended on transgressions of the limits between academic hydrodynamics and applied hydrodynamics.This quote captures one of the most significant lessons in the book: the study of concrete phenomena was critical in overcoming many of the difficulties in hydrodynamics. Roughly two upstream problems required interacting with concrete phenomena to solve. The first is summarized nicely by the quote above: theorizing and abstract thinking alone was not enough to solve the problems posed by hydrodynamics. The second is a subtler point: early mathematical and theoretical tools weren’t adapted to understanding hydrodynamics. Much of the necessary mathematical machinery existed quite early in the 1800s, but people hadn’t built the physico-mathematical tools or intuitions to tell us what they physically meant. Contact with reality forced scientists to confront the inadequacies of their theories while guiding the adaption of physico-mathematical tools.While the book is organized along rough problems that hydrodynamics faced (waves, viscosity, vortices, instability, etc.), this review will focus on broader scientific lessons. First on the two big themes I think are most important, then briefly on the other themes at the end.Practice and theory in hydrodynamicsThe first problem was that abstract thinking and theorizing proved unable to solve many of the problems of hydrodynamics. A great example of this comes from the discovery of Reynold’s number, which predicts whether flow is turbulent or laminar. Reynold’s number could have potentially been hypothesized as a consequence of Navier-Stokes, which describes viscous flow behavior. But Navier-Stokes is not analytically solvable, so Reynold’s number doesn’t come automatically. Making this more difficult is that turbulent flow, such as after submerged propellers, is generally invisible. Instead of reasoning his way there from first principles, Osborne Reynolds firs...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How does GPT-3 spend its 175B parameters?, published by Robert AIZI on January 13, 2023 on LessWrong.[Target audience: Me from a week ago, and people who have some understanding of ML but want to understand transformers better on a technical level.]Free advice for people learning new skills: ask yourself random questions. In answering them, you’ll strengthen your understanding and find out what you really understand and what’s actually useful. And some day, if you ask yourself a question that no one has asked before, that’s a publication waiting to happen!So as I was reading up on transformers, I got fixated on this question: where are the 175 billion parameters in the architecture? Not in the literal sense (the parameters are in the computer), but how are they “spent” between various parts of the architecture - the attention heads vs feed-forward networks, for instance. And how can one calculate the number of parameters from the architecture’s “size hyperparameters” like dimensionality and number of layers?The goal of this post is to answer those questions, and make sense of this nice table from the GPT-3 paper, deriving the nparams column from the other columns.Primary SourcesLots of resources about transformers conjure information from thin air, and I want to avoid that, so I’m showing all my work here. These are the relevant parts of the sources we'll draw from:Three more details we’ll use, all from Section 2.1 of the GPT-3 paper:The vocabulary size is [nvocab=]50257 tokens (via a reference to Section 2.3 of the GPT-2 paper)The feed-forward networks are all a single layer which is “four times the size of the bottleneck layer”, so dff=4dmodel“All models use a context window of nctx=2048 tokens.”Variable abbreviationsI’ll use shorthand for the model size variables to increase legibility:nlayers=xdmodel=ynheads=zdhead=wnvocab=vnctx=uWhere are the Parameters?From Exhibit A, we can see that the original 1-hot encoding of tokens U is first converted to the initial “residual stream” h0, then passed through transformer blocks (shown in Exhibits B-D), with nlayers blocks total. We'll break down parameter usage by stage:Word Embedding ParametersWe is the word embedding matrix.Converts the shape (nctx, nvocab) matrix U into a (nctx,dmodel) matrix, so We has size (nvocab,dmodel), resulting in vy=nvocabdmodel parameters.Position Embedding ParametersWp is the position embedding matrix. Unlike the original transformer paper, GPT learns its position embeddings.Wp is the same size as the residual stream, (nctx,dmodel), resulting in uy = nctxdmodel parametersTransformer Parameters - AttentionThe attention sublayer of the transformer is one half of the basic transformer block (Exhibit B). As shown in Exhibit C, each attention head in each layer is parameterized by 3 matrices, WQi,WKi,WVi, with one additional matrix WO per layer which combines the attention heads.What Exhibit C calls dk and dv are both what GPT calls dhead, so WQi,WKi, and WVi are all size (dmodel,dhead). Thus each attention head contributes 3dmodeldhead parameters.What Exhibit C calls h is what GPT calls nheads, so WO is size (nheads∗dhead,dmodel) and therefore contributes nheadsdheaddmodel parameters.Total parameters per layer: For a single layer, there are nheads attention heads, so the WQi,WKi, and WVi matrices contribute 3dmodeldheadnheads parameters, plus an additional nheadsdheaddmodel parameters from WO, for a total of 4dmodeldheadnheadsTotal parameters: 4xyzw=4dmodeldheadnheadsnlayersTransformer Parameters - FFNThe “feed-forward network” (FFN) is the other half of the basic transformer block (Exhibit B). Exhibit D shows that it consists of a linear transform parameterized by W1 and b1, an activation function, and then another linear transform parameterized by W2 and b2, as one m...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Concrete Reasons for Hope about AI, published by Zac Hatfield-Dodds on January 14, 2023 on LessWrong.Recent advances in machine learning—in reinforcement learning, language modeling, image and video generation, translation and transcription models, etc.—without similarly striking safety results, have rather dampened the mood in many AI Safety circles. If I was any less concerned by extinction risks from AI, I would have finished my PhD[1] as planned before moving from Australia to SF to work at Anthropic; I believe that the situation is both urgent and important.[2]On the other hand, despair is neither instrumentally nor terminally valuable.[3] This essay therefore lays out some concrete reasons for hope, which might help rebalance the emotional scales and offer some directions to move in.Background: a little about AnthropicI must emphasize here that this essay represents only my own views, and not those of my employer. I’ll try to make this clear by restricting we to actions, and using I for opinions to avoid attributing my own views to my colleagues. Please forgive any lapses of style or substance.Anthropic’s raison d’etre is AI Safety. It was founded in early 2021, as a public benefit corporation,[4] and focuses on empirical research with advanced ML systems. I see our work as having four key pillars:Training near-SOTA models. This ensures that our safety work will in fact be relevant to cutting-edge systems, and we’ve found that many alignment techniques only work at large scales.[5] Understanding how capabilities emerge over model scale and training-time seems vital for safety, as a basis to proceed with care or as a source of evidence that continuing to scale capabilities would be immediately risky.Direct alignment research. There are many proposals for how advanced AI systems might be aligned, many of which can be tested empirically in near-SOTA (but not smaller) models today. We regularly produce the safest model we can with current techniques,[6] and characterize how it fails in order to inform research and policy. With RLHF as a solid baseline and building block, we're investigating more complicated but robust schemes such as constitutional AI, scalable supervision, and model-assisted evaluations.Interpretability research. Fully understanding models could let us rule out learned optimizers, deceptive misalignment, and more. Even limited insights would be incredibly valuable as an independent check on other alignment efforts, and might offer a second chance if they fail.Policy and communications. I expect AI capabilities will continue to advance, with fast-growing impacts on employment, the economy, and cybersecurity. Having high-trust relationships between labs and governments, and more generally ensuring policy-makers are well-informed, seems robustly positive.If you want to know more about what we’re up to, the best place to check is anthropic.com for all our published research. We’ll be posting more information about Anthropic throughout this year, as well as fleshing out the website.Concrete reasons for hopeMy views on alignment are similar to (my understanding of) Nate Soares’. I think the key differences are because I don’t think there’s enough evidence to confidently predict the difficulty of future problems, and I do think it’s possible for careful labs to avoid active commission of catastrophe. We also seem to have different views on how labs should respond to the situation, which this essay does not discuss.Language model interventions work pretty wellI wasn’t expecting this, but our helpful/harmless/honest research is in fact going pretty well! The models are far from perfect, but we’ve made far more progress than I would have expected a year ago, and no signs of slowing down yet. HHH omits several vital pieces of the full alignment...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How we could stumble into AI catastrophe, published by HoldenKarnofsky on January 13, 2023 on LessWrong.This post will lay out a couple of stylized stories about how, if transformative AI is developed relatively soon, this could result in global catastrophe. (By “transformative AI,” I mean AI powerful and capable enough to bring about the sort of world-changing consequences I write about in my most important century series.)This piece is more about visualizing possibilities than about providing arguments. For the latter, I recommend the rest of this series.In the stories I’ll be telling, the world doesn't do much advance preparation or careful consideration of risks I’ve discussed previously, especially re: misaligned AI (AI forming dangerous goals of its own).People do try to “test” AI systems for safety, and they do need to achieve some level of “safety” to commercialize. When early problems arise, they react to these problems.But this isn’t enough, because of some unique challenges of measuring whether an AI system is “safe,” and because of the strong incentives to race forward with scaling up and deploying AI systems as fast as possible.So we end up with a world run by misaligned AI - or, even if we’re lucky enough to avoid that outcome, other catastrophes are possible.After laying these catastrophic possibilities, I’ll briefly note a few key ways we could do better, mostly as a reminder (these topics were covered in previous posts). Future pieces will get more specific about what we can be doing today to prepare.BackdropThis piece takes a lot of previous writing I’ve done as backdrop. Two key important assumptions (click to expand) are below; for more, see the rest of this series.How we could stumble into catastrophe from misaligned AIThis is my basic default picture for how I imagine things going, if people pay little attention to the sorts of issues discussed previously. I’ve deliberately written it to be concrete and visualizable, which means that it’s very unlikely that the details will match the future - but hopefully it gives a picture of some of the key dynamics I worry about.Throughout this hypothetical scenario (up until “END OF HYPOTHETICAL SCENARIO”), I use the present tense (“AIs do X”) for simplicity, even though I’m talking about a hypothetical possible future.Early commercial applications. A few years before transformative AI is developed, AI systems are being increasingly used for a number of lucrative, useful, but not dramatically world-changing things.I think it’s very hard to predict what these will be (harder in some ways than predicting longer-run consequences, in my view),2 so I’ll mostly work with the simple example of automating customer service.In this early stage, AI systems often have pretty narrow capabilities, such that the idea of them forming ambitious aims and trying to defeat humanity seems (and actually is) silly. For example, customer service AIs are mostly language models that are trained to mimic patterns in past successful customer service transcripts, and are further improved by customers giving satisfaction ratings in real interactions. The dynamics I described in an earlier piece, in which AIs are given increasingly ambitious goals and challenged to find increasingly creative ways to achieve them, don’t necessarily apply.Early safety/alignment problems. Even with these relatively limited AIs, there are problems and challenges that could be called “safety issues” or “alignment issues.” To continue with the example of customer service AIs, these AIs might:Give false information about the products they’re providing support for. (Example of reminiscent behavior)Give customers advice (when asked) on how to do unsafe or illegal things. (Example)Refuse to answer valid questions. (This could result from compani...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: A Year of AI Increasing AI Progress, published by ThomasW on December 30, 2022 on LessWrong.In July, I made a post about AI being used to increase AI progress, along with this spreadsheet that I've been updating throughout the year. Since then, I have run across more examples, and had others submit examples (some of which were published before the date I made my original post).2022 has included a number of instances of AI increasing AI progress. Here is the list. In each entry I also credit the person who originally submitted the paper to my list.A paper from Google Research used a robust supervised learning technique to architect hardware accelerators [March 17th, submitted by Zach Stein-Perlman]A paper from Google Research and Stanford fine tuned a model on its own chain-of-thought outputs, to improve performance on reasoning tasks [March 28th, submitted by Nathaniel Li]A paper from OpenAI used LLMs to help humans find flaws in other LLMs, thereby enabling them to more easily improve those models [June 12th, submitted by Dan Hendrycks]A paper from Google used machine learning to optimize compilers. This is less obviously accelerating AI but an earlier version of the compiler is used in Pytorch so it may end up doing so. [July 6th, submitted by Oliver Zhang]NVIDIA used deep reinforcement learning to generate nearly 13,000 circuits in their newest GPUs. [July 8th, submitted by me]Google found that ML code completion improved the productivity of their engineers. Some of them are presumably working in AI. [July 27th, submitted by Aidan O'Gara]A paper from Microsoft Research and MIT used language models to generate programming puzzle tasks for other language models. When finetuned on these tasks, the models were much better at solving the puzzles. [July 29th, submitted by Esben Khan]A paper from Google and UIUC used outputs from a language model to fine tune a language model after a majority vote procedure was used to filter outputs. [September 30th, submitted by me]A paper from DeepMind used reinforcement learning to discover more efficient matrix multiplication algorithms. [October 5th, submitted by me]A paper from Anthropic used language models, rather than humans, for feedback to improve language models. [December 16th, submitted by me]A paper from a number of universities used language models to generate examples of instruction following, which were then filtered and used to fine tune language models to follow instructions better. [December 20th, submitted by Nathaniel Li].I'm writing this fairly quickly so I'm not going to add extensive commentary beyond what I said in my last post, but I'll point out here two things:It is pretty common these days for people to use language model outputs to improve language models. This trend appears likely to continue.A lot of these papers are from Google. Not DeepMind, Google. Google may not have declared they are aiming for AGI, but they sure do seem to be writing a lot of papers that involve AI increasing AI progress. It seems important not to ignore them.Did I miss any? You can submit more here.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Things that can kill you quickly: What everyone should know about first aid, published by jasoncrawford on December 27, 2022 on LessWrong.There are things that kill you instantly, like a bullet to the head or a fall from twenty stories. First aid can’t help you there. There are also things that kill you relatively slowly, like a bacterial infection. If you have even hours to live, you can get to the emergency room.But there is a small class of things that will kill you in minutes unless someone comes to the rescue. There isn’t time to get to a hospital, there isn’t even time for help to arrive in an ambulance. There is only time for someone already on the scene to provide emergency treatment that either solves the problem, or stabilizes you until help arrives. Here, first aid can be the difference between life and death.Not long ago I became a father. Being responsible for the life of someone so helpless and vulnerable spurred me to finally take first aid training, including CPR. Here’s what I learned from that experience, and what I think everyone should know about first aid.What most of the things that kill you quickly have in common is that oxygen can’t get to your cells. If you are choking, oxygen can’t get in. If your heart stops beating, blood doesn’t flow. If you have a severe wound, you’re losing that blood rapidly. If any link in the respiratory-circulatory chain is broken, your cells are starved for oxygen and you have minutes to live.The key first aid skills follow from this: CPR manually substitutes for heart and lung action; the Heimlich maneuver expels an object from the airway; a tourniquet stops life-threatening bleeding (on an extremity, at least—if the wound is elsewhere, there is a different technique, known as packing the wound).The basic skills are remarkably simple. The course that I took was only a few hours of online instruction, followed by about an hour of in-person demonstration and practice with dummy patients. And I went through a lot of the optional material, including things like stroke, fainting, and jellyfish stings. I’m sure I’m nowhere near as good someone with more professional training or experience, but an introductory course is not daunting.The most important thing I learned is that if you find yourself in an emergency situation, it is better to do almost anything rather than nothing. Again, if someone stops breathing for any reason, they have only minutes to live. They are dead by default, unless someone intervenes. There is very little you can do to them that is worse than cutting off their oxygen.In fact, it is probably better to attempt CPR or the Heimlich maneuver than to do nothing, even if you have never been trained and are only guessing, or mimicking what you have seen on television. The skills were fairly unsurprising to me and were consistent with what I expected prior to training. This does not mean that you don’t need to bother with the training, and of course if someone trained is on hand then let them take over. But don’t let the bystander effect paralyze you if someone’s life is ever in your hands.In fact, the American Heart Association promotes a form of CPR called “hands-only,” in which you only do chest compressions, without giving breaths mouth-to-mouth. Their instructions for this are: “push hard and fast in the center of the chest.” That’s about it. if you only know that, you can do better than nothing.Similarly, if you can find an AED machine (automated external defibrillator), you do not need training to use it. The instructions are literally: open it and follow the prompts. The parts are clearly labeled, and there is a voice recording that walks you through every step of the process.In the end, the biggest thing I gained was the confidence to act.I made an Anki flashcard deck for the course a...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: It's time to worry about online privacy again, published by Malmesbury on December 25, 2022 on LessWrong.As we all know, if you have nothing to hide, you have nothing to fear. Nobody cares about your private life. You are not an important geopolitical target. Nobody's going to spy on you to know what weird pornography you watch.And so, around 2015, people gave up on online privacy. Everyone stopped worrying about corporations and governments having full access to their data. In hindsight, I have to admit that things didn't go as bad as some feared. But I don't think this will last.1. Radioactive decayBased on real-life events: you're a biologist at the Bad Pathogen Research Institute. You receive an email from a graduate student whose name sounds vaguely familiar. She needs to measure radio-labeled samples with a scientific instrument but, unfortunately, you used it yesterday and you forgot to log out. Now it's locked with your password and she can't connect or even reboot. She's asking you to come as soon as you can to unlock it – as radioactivity decays, the signal is vanishing every minute. Sadly, you are attending a talk on the other side of the city, it's 45 minutes by bike and it's snowing.Obviously, you would never send your credentials by e-mail, right? Right?This could, in principle, be phishing. Technically, a cunning spy could have stalked you, figured out your schedule, and crafted a deceptive e-mail to steal your password. But you know it's probably not the case, because nobody cares about your passwords enough to do something so complicated. So you send your credentials to the grad student using a one-time secret sharing link and everything is fine.I like to think that I can't be scammed because I know the ways of 1337 h4xx0rs well enough so they can't reach me. Of course, this is not true. I could totally be scammed, attackers simply don't have any interest in deploying the amount of energy it takes to scam me.That's why some people get phished and not others. It depends on two things:🅰️. How much effort it takes to set up a scam so a given target falls for it🅱️. How much effort an attacker is ready to dedicate to scamming that targetIf 🅰️ is lower than 🅱️, the target gets scammed. If 🅱️ is lower than 🅰️, it's not worth it. On one end (high 🅰️, high 🅱️), you have hackers leaking e-mails from an important government official. On the other end (low 🅰️, low 🅱️), your grandfather receives an e-mail saying a hacker has caught him watching porn and he needs to send money otherwise the hacker will tell everyone. Your grandfather doesn't know much about Internet swindles, he's from a generation who's really ashamed to watch porn, and so he falls for it.You, me, and most people are in between: too Internet-proof to fall for basic generic scams; not important enough to justify sophisticated personalized scams. Let me insist, you are safe not because hackers can't reach you, but because you are not important enough to justify the kind of attacks that would reach you.2. The classic roast chicken scam"Hi Alice, I hope you're having a good time at the concert. I just wanted to let you know that I'm at your apartment with a roast chicken that I bought at the farmer's market. My phone is out of battery, so I'm using my friend's phone to send you this message. Could you please send me your apartment door code so I can leave the chicken in front of your door?"It took ChatGPT less time to write this than it took me to copy-paste it. Most of the personal context could be figured out based on localization data. Obviously, you would never let a website access your localization data unless strictly necessary, right? Right?I don't know about you, but I'm scared. Artificial intelligence can totally automate the process of stalking someone. It can extract all the ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Shared reality: a key driver of human behavior, published by kdbscott on December 24, 2022 on LessWrong.Or: how to have a nice time with your family during the holidays.Model status: Well refined and very useful personally. But I haven't taught it, not sure how well it maps for others.I once asked Robin Hanson if he really thought status-seeking was such a dominant driver of human behavior. I said humans had dozens of factors motivating their behavior, it was crazy to claim there was One Big Thing. He replied (something to the effect of) "well, even if each factor has a small effect – one percent, two percent – one of them has to be the biggest."There's a concept I refer to as 'shared reality' that I think is up there with 'status' as something humans seek, shaping a lot (maybe five percent?) of our behavior.Knowing and playing with the concept of shared reality has noticeably improved my relationships and given me more surface area on many social concepts (e.g. connection, attachment, idle banter, tribalism).What is shared reality?There is a Columbia research lab that studies shared reality among other topics. They define it as:the perceived commonality of inner states with othersBy example: If two people chuckle at a pun, and then see each other chuckling such that they both perceive that the other had a similar pun-chuckle-experience, they are in a state of shared reality (at least, for that pun-chuckle-experience in particular).Shared reality is at play when peoplego to sports events / concerts / movies together.That emotional rush when everyone cheers / sings / laughs at the same time.travel / eat / dance together.That moment when you both see / eat / intuit something cool, look at each other and see you've had a similar experience, and get a rush of connection.Or that moment when you have an experience and suddenly want to share it with others, so you offer 'try this!' or 'look at that!' or take a picture and share it with the people who you think will best 'get it'.Shared reality is niceThe research literature mentions that people are 'motivated to create shared realities' but doesn't really discuss why. I think people want it because it's pleasant. And conversely, that the opposite of shared reality (disconnected reality?) is unpleasant, and something people avoid.This strikes me as an important part, because put together with the above definition it makes for a model that better predicts how people will behave, and what is sometimes causing people to feel better/worse.Here's a stab:people seek a perceived commonality of inner states with others (because it's pleasant). People try to maintain that perceived commonality and/or avoid perceived uncommonality (because failing to do so is unpleasant).Wow, that's pretty clunky. Oh well.Here be dragons reality masking puzzlesAs you might expect with something that involves 'people seeking a perceived X', attempts at shared reality often skip right over actually sharing an experience to merely convincing oneself/others that an experience was shared. I think this is often playing out in the various forms of conformity. Some quick examples:When spending the holidays with family, I often feel a gentle pressure/request to do the same activity (sit/eat/talk together), even if I’d rather be doing something else.When talking with my parents about AI, I notice myself bouncing between frustrated that they don’t think about it the same way I do, and subdued/confused, agreeing that this whole thing is probably overblown. I think this flailing is motivated by a desire to connect (and fear of disconnection / not being seen and known by them)Social drinking/smokingBandwagoning / groupthinkI've developed a kind of backing-away immune response to many of these conformity / shared reality pressures. I thin...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: On sincerity, published by Joe Carlsmith on December 23, 2022 on LessWrong.Cross-posted from my website. Audio version here, or search "Joe Carlsmith Audio" on your podcast app.Nearby is the country they call life.You will know it by its seriousness.Rilke1. IntroductionThere’s a thing I call “sincerity” that matters a lot to me. In particular, it’s core to how I hope to orient towards the world. And it’s one of the main things I look for in people and communities.My hope, in this essay, is to bring what I mean by sincerity into clearer view. But the term has a fairly rich set of associations for me, which I’m not sure will ultimately admit of a cleanly unified analysis. I start by discussing five of these associations. Sincerity seems to me closely related to:Something like truth-seeking (“scout-mindset”), but for agency as a whole rather than just beliefs“Not playing pretend”One’s different motivations working in harmony“Seriousness”(Less confidently) Some stuff about non-ego/altruism/goodnessI then discuss whether we can unify these five associations under a single principle. I offer two possibilities for doing so. The first is conceptual, and takes scout-mindset-but-for-agency as the core thing. The second is empirical, and takes “not playing pretend” as the core thing. I’m not sure either is adequate.I also discuss a few of the many failure modes nearby to sincerity: self-deception, extremism, over-confidence, “the-normal-rules-don’t-apply-to-me”-ness, sanctimony, over-concern with sincerity itself (both in others, and in yourself), and that most straightforward and fearsome of failures: just plain picking the wrong actions, even if you did the other stuff right.Thanks to Katja Grace, Ketan Ramakrishnan, Claire Zabel, Samira Nedungadi, Cate Hall, Bill Zito, Nick Whitaker, and Dwarkesh Patel for discussion.“Jeanne d’Arc écoutant les voix,” by Thirion, image source here.2. Initial gesturesThere is a true person of no rank who is always coming and going from the portals of your face.Who is that true person of no rank?LinjiBefore jumping into the five associations above, I want to make a few more general gestures at the thing I have in mind when I say “sincerity,” to get us at least somewhat oriented. First, I’ll note and comment on some related words. Then, I’ll describe some related feelings.2.1 Some wordsThe word “sincerity,” here, isn’t perfect. In particular, I think it over-evokes something directly social – e.g. whether you are accurately presenting your beliefs and motivations to others. Google says: “being free from pretense, hypocrisy, and deceit.”I think something about pretense is important here, and I discuss it below. But sincerity in the main sense I have in mind is compatible with lying (indeed, this is one of the dangers I discuss in the final section).Suppose, for example, that per grand thought-experimental tradition, you are living amongst the Nazis.You hate Nazism and have total, steely internal clarity about this. When you lie to your boss about your loyalty to Hitler, you are clearly being insincere in one sense – but it’s not the main sense I have in mind.Also, “sincerity” tempts you to group candidates for sincerity into only two buckets: “sincere” and “insincere.” But much of life is neither actively sincere (at least in my sense), nor especially insincere – consider, for example, a casual greeting.Still, despite these issues, “sincerity” is the term I actually use. And I don’t think this is an accident. So I’m going to keep using it.Another word is “earnest.” This might be closer – though it can sound a bit breathless. The idea of being “in good faith” seems related, but also too social. “Genuine” is also too social. “Authentic” is too interested in the self. I’m tempted by “good will,” but unsure (more below).S...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Sazen, published by Duncan Sabien on December 21, 2022 on LessWrong.Purpose of post: describe and (hopefully) popularize a concept I've found highly useful.Last year, my partner Logan Strohl wrote a sequence to introduce the "naturalism" concept they've been developing and teaching for the past decade or so.That sequence was structured around a single, short sentence. The first essay introduced the sentence, and the remaining essays were primarily about explaining what each of the important concepts in that short sentence actually meant.So, for the sentence "knowing the territory takes direct and patient observation," there was a full essay on what was intended (and, more crucially, what was not intended) by the word "knowing," and another on "the territory," and another on "observation," and so on.This format was largely inspired by a conversation in which I asked Logan to describe naturalism briefly, and they said "I totally can, but you'll get the wrong idea."Together, we realized that there is a curious one-way sort of property to many sentences, in which they work as pointers or summaries after the fact, but fail to generate the-thing-they're-summarizing if used as standalone seeds.(One could argue that every sentence has some of this property, but some sentences have a lot of it.)I'd like to be able to point directly at this property, and as a result of historical accident that I'll explain in a footnote, the handle I've ended up with in my own head is sazen.Example 1: "Duncan Sabien is a teacher and a writer."This is a true sentence. People who know me very, very well, upon hearing this sentence, will nod. It's a good fit, retrospectively, for the data.However, if you are attempting to give someone a sense of me up-front, saying "Duncan Sabien is a teacher and a writer" is an unusually bad start. The thing that most people will think of when they hear "teacher" or "writer" is specifically unlike me—I'm a very weird sort of teacher and a very weird sort of writer, and so anchoring people on the representative stereotypes is almost actively misleading.The sentence "Duncan Sabien is a teacher and a writer" is a sazen.Example II: Peanut butter and jelly sandwichesThere's a classic challenge in which a teacher asks a bunch of students to write down unambiguous and complete instructions for how to make a peanut butter and jelly sandwich. The gimmick is that, to grade each paper, the teacher will follow only the actions written on the page, which usually results in something very unlike a normal sandwich.This activity is ... somewhat arbitrary and infuriating, because the teacher usually has to make a bunch of fluid judgment calls about where they draw the line, and there's usually not a clear and consistent standard for what level of detail is required, so the lesson often ends up being less about the complexity of background information and more about the fickle dickishness of teachers.But purely as a sort of plug-and-play teachable moment, it's an interesting way to take a close and practical look at all of the little bits of context we take for granted, via the teacher pretending not to know them.A sentence like "get a couple of slices of bread, and put peanut butter on one and jelly on the other and then stick 'em together" is a sazen. After the fact, you can look back and say "sure, those bones match the shape of what just happened."But up front, they're woefully insufficient. They also match (for example) putting the entire jar of peanut butter atop one slice, and the entire jar of jelly atop the other, and then sliding the two smushed rectangles of bread into each other, flat on the counter. Or failing to use bread, peanut butter, or jelly at all, because the instructions didn't say to get those items and have them on hand, or didn't specify...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How to Convince my Son that Drugs are Bad, published by concerned dad on December 17, 2022 on LessWrong.Hello.My son (16m, henceforth referred to as John) has monologued about this site a few times over the past couple of months, so I figured, based on my brief impression of the community, you might be able to help me with an issue. Given the topical nature here, I am not sure if this is an appropriate type of post to make, however it might be a useful place to make an appeal. Worst case, this gets taken down for incompliance.John has always been a little too obsessed with his computer, but things really came to a head when he found this whole subcommunity. For a couple of weeks, I'd regularly notice as he spent hours just sitting in his room scrolling through blog posts and papers and forums. While damaging to health, this doesn't seem unusual among teenagers and I try to let him make his own decisions as I'm sure eventually it will taper out, and that's not my main issue.First off: I've noticed some positive changes since he started discussing effective altruism and rationality and such, though I don't know whether to attribute that to this site or just maturing. Thank you all for that. But there are some worrying ideas he seems to have gotten as well, centering around romanticization of drug use, specifically "nootropic"-style with a focus on amphetamines and psychedelics.One day he just came up to me and began engaging in a discussion about the merits of doing drugs. I'll give you some approximate quotes so you can understand about how this went:"Why the hell is LSD criminalized everywhere? There are NO negative side effects proceeds to compare to prescription drugs""Psychedelics and amphetamines are classes of drugs I plan to do, I've done extensive (3 hr - sigh) research on their side effects and chemical compositions and all seems fine! People take Adderall all the time and I can easily get a script, I've read the entire DSM 5!""What's the difference between you drinking alcohol or coffee and me taking amphetamines and doing LSD? Drugs. are. drugs." "Alcohol isn't synthetic? Well then, what about peyote?""Why not try heroin if the purpose of life is to optimize happiness assuming heroin provides proportionally more even if for a shorter amount of time?" (!)"Look, here's this site where someone (a 'gwern' if I recall correctly) did this scientifically! They're fine, I want to do that too!""How will I acquire the drugs? I'll just ... uhhh ... synthesize the LSD myself! Can't be that difficult"This discussion re-occurs any time someone brings up recreational or cognitive-enhancing drug use. I'm getting frustrated of explaining how dangerous these ideations are and how a lot of these drugs can permanently damage him, through addiction or brain damage or other negative health effects.But he is apparently "a rational being who can make his own decisions". While he's certainly intelligent, he's misguided in this particular direction. How can I persuade him to stop these thoughts?Relevant note: our current settlement of the situation is that he's going to wait until he's age of majority then do whatever he wants. This is an outcome I want to avoid, as I fear he will fall into psychosis or addiction. It's possible that these ideas will fade over the next year or so, but I'm looking to accelerate that period.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Can we efficiently explain model behaviors?, published by paulfchristiano on December 16, 2022 on LessWrong.ARC’s current plan for solving ELK (and maybe also deceptive alignment) involves three major challenges:Formalizing probabilistic heuristic argument as an operationalization of “explanation”Finding sufficiently specific explanations for important model behaviorsChecking whether particular instances of a behavior are “because of” a particular explanationAll three of these steps are very difficult, but I have some intuition about why steps #1 and #3 should be possible and I expect we’ll see significant progress over the next six months. Unfortunately, there’s no simple intuitive story for why step #2 should be tractable, so it’s a natural candidate for the main technical risk.In this post I’ll try to explain why I’m excited about this plan, and why I think that solving steps #1 and #3 would be a big deal, even if step #2 turns out to be extremely challenging.I’ll argue:Finding explanations is a relatively unambitious interpretability goal. If it is intractable then that’s an important obstacle to interpretability in general.If we formally define “explanations,” then finding them is a well-posed search problem and there is a plausible argument for tractability.If that tractability argument fails then it may indicate a deeper problem for alignment.This plan can still add significant value even if we aren’t able to solve step #2 for arbitrary models.I. Finding explanations is closely related to interpretabilityOur approach requires finding explanations for key model behaviors like “the model often predicts that a smiling human face will appear on camera.” These explanations need to be sufficiently specific that they distinguish (the model actually thinks that a human face is in front of the camera and is predicting how light reflects off of it) from (the model thinks that someone will tamper with the camera so that it shows a picture of a human face).Our notion of “explanation” is informal, but I expect that most possible approaches to interpretability would yield the kind of explanation we want (if they succeeded at all). As a result, understanding when finding explanations is intractable may also help us understand when interpretability is intractable.As a simple caricature, suppose that we identify a neuron representing the model’s beliefs about whether there is a person in front of the camera. We then verify experimentally that (i) when this neuron is on it leads to human faces appearing on camera, (ii) this neuron tends to fire under the conditions where we’d expect a human to be in front of the camera.I think that finding this neuron is the hard part of explaining the face-generating-behavior. And if this neuron actually captures the model’s beliefs about humans, then it will distinguish (human in front of camera) from (sensors tampered with). So if we can find this neuron, then I think we can find a sufficiently specific explanation of the face-generating-behavior.In reality I don’t expect there to be a “human neuron” that leads to such a simple explanation, but I think the story is the same no matter how complex the representation is. If beliefs about humans are encoded in a direction then both tasks require finding the direction; if they are a nonlinear function of activations then both tasks require understanding that nonlinearity; and so on..The flipside of the same claim is that ARC’s plan effectively requires interpretability progress. From that perspective, the main way ARC’s research can help is by identify a possible goal for interpretability. By making a goal precise we may have a better chance of automating it (by applying gradient descent and search, as discussed in section III), and even if we can’t automate it then a clearer se...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How "Discovering Latent Knowledge in Language Models Without Supervision" Fits Into a Broader Alignment Scheme, published by Collin on December 15, 2022 on LessWrong.IntroductionA few collaborators and I recently released a new paper: Discovering Latent Knowledge in Language Models Without Supervision. For a quick summary of our paper, you can check out this Twitter thread.In this post I will describe how I think the results and methods in our paper fit into a broader scalable alignment agenda. Unlike the paper, this post is explicitly aimed at an alignment audience and is mainly conceptual rather than empirical.Tl;dr: unsupervised methods are more scalable than supervised methods, deep learning has special structure that we can exploit for alignment, and we may be able to recover superhuman beliefs from deep learning representations in a totally unsupervised way.Disclaimers: I have tried to make this post concise, at the cost of not making the full arguments for many of my claims; you should treat this as more of a rough sketch of my views rather than anything comprehensive. I also frequently change my mind – I’m usually more consistently excited about some of the broad intuitions but much less wedded to the details – and this of course just represents my current thinking on the topic.ProblemI would feel pretty optimistic about alignment if – loosely speaking – we can get models to be robustly “honest” in a way that scales even to superhuman systems. Moreover, I think a natural sub-problem that captures much or most of the difficulty here is: how can we make a language model like GPT-n “truthful” or “honest” in a way that is scalable? (For my purposes here I am also happy to make the assumption that GPT-n is not actively deceptive, in the sense that it does not actively try to obscure its representations.)For example, imagine we train GPT-n to predict news articles conditioned on their dates of publication, and suppose the model ended up being able to predict future news articles very well. Or suppose we train GPT-n to predict the outcomes of particular actions in particular situations, all described (imperfectly by humans) in text. Then I would expect GPT-n would eventually (for large enough n) have a superhuman world model in an important sense. However, we don’t currently know how to recover the “beliefs” or “knowledge” of such a model even in principle.A naive baseline for trying to make GPT-n truthful is to train it using human feedback to output text that human evaluators believe to be true. The basic issue with this is that human evaluators can’t assess complicated claims that a superhuman system might make. This could lead to either competitiveness problems (if GPT-n only outputs claims that humans can assess) or misalignment issues (if GPT-n outputs false claims because human evaluators can’t assess them correctly).In many ways this problem is similar to Eliciting Latent Knowledge (ELK), but unlike ELK I am happy to take a “non-worst-case” empirical perspective in studying this problem. In particular, I suspect it will be very helpful – and possibly necessary – to use incidental empirical properties of deep learning systems, which often have a surprising amount of useful emergent structure (as I will discuss more under “Intuitions”).On the other hand, if we want to study scalable alignment empirically, I think it’s very important for us to also have good reason to believe that our experiments will say something meaningful about future models – and it’s not immediately clear how to do that.This raises the question: how do we even approach doing research on this sort of problem, methodologically?MethodologyI worry that a lot of theoretical alignment work is either ungrounded or intractable, and I worry that a lot of empirical alignment work does...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Trying to disambiguate different questions about whether RLHF is “good”, published by Buck on December 14, 2022 on LessWrong.(A few of the words in this post were written by Ryan Greenblatt and Ajeya Cotra. Thanks to Oliver Habryka and Max Nadeau for particularly helpful comments.)Sometimes people want to talk about whether RLHF is “a promising alignment strategy”, or whether it “won’t work” or “is just capabilities research”. I think that conversations on these topics are pretty muddled and equivocate between a bunch of different questions. In this doc, I’ll attempt to distinguish some of these questions, and as a bonus, I’ll give my opinions on them.I wrote this post kind of quickly, and I didn’t have time to justify all the claims I make; I hope that this post is net helpful anyway. I’m sympathetic to claims that alignment researchers should err more on the side of writing fewer but better posts; maybe I’ll regret making this one now instead of waiting.Is “make a powerful AGI by using RLHF, where the feedback is coming from unaided humans”, a promising strategy for building aligned AGI?IMO, this seems like the baseline, “we didn’t really try much at all” alignment scheme. Ajeya calls this a “naive safety effort” in her training game post, which lays out the basic case for pessimism about this strategy.Here’s how I’d quickly summarize my problems with this scheme:Oversight problems:Overseer doesn’t know: In cases where your unaided humans don’t know whether the AI action is good or bad, they won’t be able to produce feedback which selects for AIs that do good things. This is unfortunate, because we wanted to be able to make AIs that do complicated things that have good outcomes.Overseer is wrong: In cases where your unaided humans are actively wrong about whether the AI action is good or bad, their feedback will actively select for the AI to deceive the humans.Catastrophe problems:Even if the overseer’s feedback was perfect, a model whose strategy is to lie in wait until it has an opportunity to grab power will probably be able to successfully grab power.I don’t think we’re 100% doomed if we follow this plan, but it does seem pretty likely to go badly.RLHF with unaided humans is not literally the most doomed alignment scheme I’ve ever heard seriously proposed. For example, “train models with automated rewards (e.g. simulating evolution, or training models on a curriculum of math problems) and hope that the resulting models are aligned” might be a worse plan. (Though it’s pretty plausible to me that this kind of scheme would have such obvious alignment issues that people would quickly switch to the naive safety plan.)I don’t think that many alignment researchers are seriously proposing this naive plan. Many researchers who work on RLHF are sympathetic to the concerns that I listed here. For example, OpenAI’s alignment plan emphasizes the importance of using models to assist human evaluation.Is RLHF broadly construed (i.e. pretraining a model and then fine-tuning it based on some overseer’s evaluations of its actions) plausibly part of a not-completely-doomed alignment plan?IMO, yes, we are very likely to want to make powerful models by fine-tuning models based on overseer feedback, because fine-tuning models using preferences over their outputs is a broadly applicable strategy that can be used in lots of ways.I think that our best current alignment strategies (I basically agree with Holden here on the current best plan) involve fine-tuning a model based on feedback from human overseers who have AI and software assistance, on inputs including some that were chosen by a red team who is trying to find inputs where the model behaves very badly.If prosaic alignment researchers are able to come up with alignment schemes which look good on paper (i.e., do...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AI alignment is distinct from its near-term applications, published by paulfchristiano on December 13, 2022 on LessWrong.I work on AI alignment, by which I mean the technical problem of building AI systems that are trying to do what their designer wants them to do.There are many different reasons that someone could care about this technical problem.To me the single most important reason is that without AI alignment, AI systems are reasonably likely to cause an irreversible catastrophe like human extinction. I think most people can agree that this would be bad, though there’s a lot of reasonable debate about whether it’s likely. I believe the total risk is around 10–20%, which is high enough to obsess over.Existing AI systems aren’t yet able to take over the world, but they are misaligned in the sense that they will often do things their designers didn’t want. For example:The recently released ChatGPT often makes up facts, and if challenged on a made-up claim it will often double down and justify itself rather than admitting error or uncertainty (e.g. see here, here).AI systems will often say offensive things or help users break the law when the company that designed them would prefer otherwise.We can develop and apply alignment techniques to these existing systems. This can help motivate and ground empirical research on alignment, which may end up helping avoid higher-stakes failures like an AI takeover. I am particularly interested in training AI systems to be honest, which is likely to become more difficult and important as AI systems become smart enough that we can’t verify their claims about the world.While it’s nice to have empirical testbeds for alignment research, I worry that companies using alignment to help train extremely conservative and inoffensive systems could lead to backlash against the idea of AI alignment itself. If such systems are held up as key successes of alignment, then people who are frustrated with them may end up associating the whole problem of alignment with “making AI systems inoffensive.”If we succeed at the technical problem of AI alignment, AI developers would have the ability to decide whether their systems generate sexual content or opine on current political events, and different developers can make different choices. Customers would be free to use whatever AI they want, and regulators and legislators would make decisions about how to restrict AI. In my personal capacity, I have views on what uses of AI are more or less beneficial and what regulations make more or less sense, but in my capacity as an alignment researcher I don’t consider myself to be in the business of pushing for or against any of those decisions.There is one decision I do strongly want to push for: AI developers should not develop and deploy systems with a significant risk of killing everyone. I will advocate for them not to do that, and I will try to help build public consensus that they shouldn’t do that, and ultimately I will try to help states intervene responsibly to reduce that risk if necessary. It could be very bad if efforts to prevent AI from killing everyone were undermined by a vague public conflation between AI alignment and corporate policies.Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Plan - 2022 Update, published by johnswentworth on December 1, 2022 on LessWrong.So, how’s The Plan going?Pretty well!In last year’s writeup of The Plan, I gave “better than a 50/50 chance” that it would work before AGI kills us all (and my median AI timelines were around 10-15 years). That was an outside view, accounting for planning fallacy and the inevitable negative surprises. My inside view was faster - just based on extrapolating my gut feel of the rate of progress, I privately estimated that The Plan would take around 8 years. (Of those 8, I expected about 3 would be needed to nail down the core conceptual pieces of agent foundations, and the other 5 would be to cross the theory-practice gap. Of course those would be intermingled, though with the theory part probably somewhat more front-loaded.)Over the past year, my current gut feel is that progress has been basically in line with the inside-view 8 year estimate (now down to 7, since a year has passed), and maybe even a little bit faster than that.So, relative to my outside-view expectation that things always go worse than my gut expects, things are actually going somewhat better than expected! I’m overall somewhat more optimistic now, although the delta is pretty small. It’s only been a year, still lots of time for negative surprises to appear.Any high-level changes to The Plan?There have been two main high-level changes over the past year.First: The Plan predicted that, sometime over the next 5 (now 4) years, the field of alignment would “go from a basically-preparadigmatic state, where we don’t even know what questions to ask or what tools to use to answer them, to a basically-paradigmatic state, where we have a general roadmap and toolset”. Over the past year, I tentatively think the general shape of that paradigm has become visible, as researchers converge from different directions towards a common set of subproblems.Second: I’ve updated away from thinking about ambitious value learning as the primary alignment target. Ambitious value learning remains the main long-term target, but I’ve been convinced that e.g. corrigibility is worth paying attention to as a target for early superhuman AGI. Overall, I’ve updated from “just aim for ambitious value learning” to “empirically figure out what potential medium-term alignment targets (e.g. human values, corrigibility, Do What I Mean, human mimicry, etc) are naturally expressible in an AGI’s internal concept-language”.Convergence towards a paradigm sounds exciting! So what does it look like?Exciting indeed! Gradual convergence toward a technical alignment paradigm has probably been the most important update from the past year.On the theoretical side, Paul Christiano, Scott Garrabrant, and myself had all basically converged to working on roughly the same problem (abstraction, ontology identification, whatever you want to call it) by early 2022. That kind of convergence is a standard hallmark of a proto-paradigm.Meanwhile, within the past year-and-a-half or so, interpretability work has really taken off; Chris Olah’s lab is no longer head-and-shoulders stronger than everyone else. And it looks to me like the interpretability crowd is also quickly converging on the same core problem of abstraction/ontology-identification/whatever-you-want-to-call-it, but from the empirical side rather than the theoretical side.That convergence isn’t complete yet - I think a lot of the interpretability crowd hasn’t yet fully internalized the framing of “interpretability is primarily about mapping net-internal structures to corresponding high-level interpretable structures in the environment”. In particular I think a lot of interpretability researchers have not yet internalized that mathematically understanding what kinds of high-level interpretable structures appea...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Be less scared of overconfidence, published by benkuhn on November 30, 2022 on LessWrong.When I was deciding whether to work for Wave, I got very hung up on the fact that my “total compensation” would be “lower.”The scare quotes are there because Wave and my previous employer, Theorem, were both early-stage startups that were paying me mostly in fake startup bucks equity. To figure out the total compensation, I tried to guess how much money the equity in each company was worth, with a thought process something like:Both of these companies have been invested in by reputable, top-tier venture capitalists.The market for for-profit investments is pretty efficient, and most people who think they can do better are being overconfident.Who am I, a lowly 22-year-old programmer, to disagree with reputable top-tier venture capitalists? I should defer to them about the valuations.So I valued the equity by taking the valuation each company’s VCs had invested at, and multiplied it by the fraction of the company my shares represented. That number was higher for Theorem than for Wave.Seven years on, the Wave equity turned out to be. a lot more valuable. That raises the question: how dumb was my take? Was the actual outcome predictable if I’d thought about it in the right way?I don’t think it was perfectly predictable, but I do think I shouldn’t have been that anchored to the market-efficiency reasoning. Those respectable, top-tier VCs had YOLOed those valuations after a couple one-hour meetings, because that’s how early-stage VC works. Meanwhile, I had worked at Theorem for a year and my then-partner had worked at Wave for nine months. Heck, I had gotten more founder time than those VCs had just during my interview process. I had way more information than “the market.”If I’d had the confidence to use that information, I might have thought something like:After its funding round, Wave continued to add users at one of the fastest paces their investors had ever seen, whereas Theorem is struggling to grow.Theorem is constrained by its ability to do sales, and the founders don’t seem to be acting with enough focus or urgency to unblock that constraint. Instead, they’re distracting themselves with things like hiring machine learning interns (i.e. me).The founders of Wave seem much smarter, more relentlessly resourceful, and more trustworthy.Given the above, I should value the Wave equity way more even though its naive expected value is less than the Theorem equity.Fortunately, I chose Wave for other reasons. But this thought pattern—throwing away most information in fear of using it to make overconfident judgments—shows up all the time. I’m here to tell you why I hate it.In January 2020, my entire Twitter timeline was freaking out about a novel-seeming respiratory disease spreading in Wuhan.Part of me thought:All the reputable, top-tier technocrats are ridiculing the freaked-out people.Usually, when a ragtag band of Internet weirdos thinks they know better than a large group of reputable, top-tier technocrats, the Internet weirdos are being overconfident.So the technocrats are probably right on this one.Another part of me thought:Huh, the simple model of “this thing has a fast exponential growth rate and spreads when people are asymptomatic so it’s very hard to stop” seems like a compelling reason to think things will be quite bad.When reputable, top-tier technocrats say not to freak out, they don’t usually address the best arguments in favor of freaking out, and they often seem like they don’t understand how exponential growth works.Maybe I’ll buy a lot of beans in case everything goes to shit.(I also contemplated the fact that the stock market didn’t seem to be freaking out, but I decided that since most people can’t beat the stock market, I probably wouldn’t eith...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Singular Value Decompositions of Transformer Weight Matrices are Highly Interpretable, published by beren on November 28, 2022 on LessWrong. Please go to the colab for interactive viewing and playing with the phenomena. For space reasons, not all results included in the colab are included here so please visit the colab for the full story. This post is part of the work done at Conjecture. TLDR If we take the SVD of the weight matrices of the OV circuit and of MLP layers of GPT models, and project them to token embedding space, we notice this results in highly interpretable semantic clusters. This means that the network learns to align the principal directions of each MLP weight matrix or attention head to read from or write to semantically interpretable directions in the residual stream. We can use this to both improve our understanding of transformer language models and edit their representations. We use this finding to design both a natural language query locator, where you can write a set of natural language concepts and find all weight directions in the network which correspond to it, and also to edit the network's representations by deleting specific singular vectors, which results in relatively large effects on the logits related to the semantics of that vector and relatively small effects on semantically different clusters Introduction Trying to understand the internal representations of language models, and of deep neural networks in general, has been the primary focus of the field of mechanistic interpretability, with clear applications to AI alignment. If we can understand the internal dimensions along which language models store and manipulate representations, then we can get a much better grasp on their behaviour and ultimately may be able to both make provable statements about bounds on their behaviour, as well as make precise edits to the network to prevent or enhance desired behaviours. Interpretability, however, is a young field where we still do not yet fully understand what the basic units of the networks' representations are. While analyzing and investigating individual neurons has led to some impressive results, especially in convolutional vision models, a key issue has always been the polysemanticity of neurons. A single neuron might not just represent a single 'feature' but some linear combination of features in superposition. This effect has been studied in toy models where it is argued that neural networks resort to superposition when required to represent many more features than they have neurons, and that superposition has a regular and understandable geometry. A natural hypothesis following from the apparent ubiquity of superposition in neural networks, as well as the autoassociative memory literature, is to store features as directions and not in individual neurons. To minimize interference ideally these directions would be pseudo-orthogonal. Technically the features as neurons hypothesis is trivially an orthogonal direction where each feature is encoded by a specific neuron, but the storage capacity of this representational scheme scales only linearly. In theory, we can do much better if we instead distribute features across multiple neurons and accept some noise. Specifically, the Johnson-Lindenstrauss lemma suggests that we can store exponentially many features in pseudorthogonal subspaces. While neural networks probably cannot utilize all of this exponential space, they almost certainly scale superlinearly, necessitating polysemanticity across 'neurons'. If this hypothesis is true, at least approximately, a key question becomes how we can figure out the directions in which specific features are encoded. While certainly not the entire story, we hypothesize that at least a number of the primary directions used by the network can be in...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Geometric Rationality is Not VNM Rational, published by Scott Garrabrant on November 27, 2022 on LessWrong. One elephant in the room throughout my geometric rationality sequence, is that it is sometimes advocating for randomizing between actions, and so geometrically rational agents cannot possibly satisfy the Von Neumann–Morgenstern axioms. That is correct: I am rejecting the VNM axioms. In this post, I will say more about why I am making such a bold move. A Model of Geometric Rationality I have been rather vague on what I mean by geometric rationality. I still want to be vague in general, but for the purposes of this post, I will give a concrete definition, and I will use the type signature of the VNM utility theorem. (I do not think this definition is good enough, and want it to restrict its scope to this post.) A preference ordering on lotteries over outcomes is called geometrically rational if there exists some probability distribution P over interval valued utility functions on outcomes such that L⪯M if and only if GU∼PEO∼LU(O)≤GU∼PEO∼MU(O). For comparison, an agent is VNM rational there exists a single utility function U, such that L⪯M if and only if EO∼LU(O)≤EO∼MU(O). Geometric Rationality is weaker than VNM rationality, since under reasonable assumptions, we can assume the utility function of a VNM rational agent is interval valued, and then we can always take the probability distribution that assigns probability 1 to this utility function. Geometric Rationality is strictly weaker, because it sometimes strictly prefers lotteries over any of the deterministic outcomes, and VNM rational agents never do this. The VNM utility theorem says that any preference ordering on lotteries that satisfies some simple axioms must be VNM rational (i.e. have a utility function as above). Since I am advocating for a weaker notion of rationality, I must reject some of these axioms. Against Independence The VNM axiom that I am rejecting is the independence axiom. It states that given lotteries A, B, and C, and probability p, A⪯B if and only if pC+(1−p)A⪯pC+(1−p)B. Thus, mixing in a probability p of C will not change my preference between A and B. Let us go through an example. Alice and Bob are a married couple. They are trying to decide where to move, buy a house, and live for the rest of their lives. Alice prefers Atlanta, Bob prefers Boston. The agent I am modeling here is the married couple consisting of Alice and Bob. Bob's preference for Boston is sufficiently stronger than Alice's preference for Atlanta, that given only these options, they would move to Boston (A≺B). Bob is presented with a unique job opportunity, where he (and Alice) can move to California, and try to save the world. However, he does not actually have a job offer yet. They estimate an 80 percent chance that he will get a job offer next week. Otherwise, they will move to Atlanta or Boston. California is a substantial improvement for Bob's preferences over either of the other options. For Alice, it is comparable to Boston. Alice and Bob are currently deciding on a policy of what to do conditional on getting and not getting the offer. It is clear that if they get the offer, they will move to California. However, they figure that since Bob's preferences are in expectation being greatly satisfied in the 80 percent of worlds where they are in California, they should move to Atlanta if they do not get the offer (pC+(1−p)B≺pC+(1−p)A). Alice and Bob are collectively violating the independence axiom, and are not VNM rational. Are they making a mistake? Should we not model them as irrational due to their weird obsession with fairness? Dutch Books and Updatelessness You might claim that abandoning the independence axiom opens up Alice and Bob up to get Dutch booked. The argument would go as follows. First, you offer ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Respecting your Local Preferences, published by Scott Garrabrant on November 26, 2022 on LessWrong. In this post, I give a application of geometric rationality to a toy version of a real problem. A Conflicted Agent Let's say you are an agent with two partially conflicting goals. Part of you wants to play a video game, and part of you wants to save the world and tile the multiverse with computronium shaped exactly the way you like it (not paperclips). How should these conflicting interests figure out what to do? (Assume you have not yet had the idea of starting a video game company to save the world.) We will assume that 2/3 of you wants to save the world, and 1/3 of you wants to play video games. At first your world-saving-self has the bright idea that 2/3 is bigger that 1/3, and you should therefore devote all your time to saving the world. However, this proposal doesn't stick. Eventually, you end up Nash bargaining with your time, and devoting 2/3 of your time to saving the world and 1/3 of your time to playing video games. This works well for a while, but then your world-saving-self has a new bright idea: "Let's look around at the world, and see how much ability we expect to have to save it. If it feels like we are in the top 60 percentile of worlds ordered by how much control we have, then we will try to save the world. If we are in the bottom 40 percentile, we will play video games!" (The 60 percentile is arbitrarily rounding down from 2/3, so that you can both play more video games and save more worlds.) A Nash Bargaining Model Let's model this more carefully. Let's say there are five different types of worlds: 1, 2, 3, 4, and 5. In each world, you have two buttons in front of you. The video game button, and the save the world button. Each time step, pressing the video game button lets you play some video games, and pressing the save the world button saves the world with probability εi in the world of type i. We have five degrees of freedom. For each i∈{1,2,3,4,5}, we have pi, which represents the proportion of our time in world i that we spend pressing the save the world button. The rest of the time is spent playing video games. The part of you that wants to save the world has power 23, and utility equal to 1010100ε(p1+2p2+3p3+4p4+5p5). However, since we are Nash bargaining, the coefficient out in front does not matter. The part of you that wants to play video games has power 13 and utility equal to 5−p1−p2−p3−p4−p5. We are trying to maximize the weighted geometric mean, 3√(p1+2p2+3p3+4p4+5p5)2(5−p1−p2−p3−p4−p5), on the cube [0,1]5. This achieves a maximum when p1=p2=0, and p3=p4=p5=1. Simple enough. (Linear is not the best model for the distribution of how much control you have. I am just trying to keep things simple.) Local Preferences The main problem with the above analysis, according to me, is that it is not respecting the locality of some of your preferences. Maybe your desire to save the world is nonlocal. Maybe you care equally about whether this world is saved and whether some hypothetical other world in which you have more or less control over saving the world is saved. Why should you care about this Everett branch more than the other ones? I will grant that your altruistic self thinks that way, but I am guessing your desire to play video games probably doesn't. The part of me that wants to play video games wants to play video games in this world. If I ask it what it wants in other Everett branches, it doesn't really care much. This does not mean it is making a mistake. Indeed, if I were to go back in time before I observed what world I am in and ask it whether it wanted to play more video games, concentrated in a small number of futures, or fewer total video games equally distributed, it chooses the equal distribution. The (video game part of a) ver...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Planes are still decades away from displacing most bird jobs, published by guzey on November 25, 2022 on LessWrong. Originally published here:/ Note: Parts of this essay were written by GPT-3, so it might contain untrue facts. Introduction Many of my friends are extremely excited by planes, rockets, and helicopters. They keep showing me videos of planes flying at enormous speed, rockets taking off from the ground while creating fiery infernos around them, and of helicopters hovering midair seemingly denying the laws of gravity. I've been on a plane already, and it was nothing special. It was just a big metal tube with a bunch of people inside. It was loud and it smelled weird and I had to sit in a tiny seat for hours. So what is it that makes planes so special? Is it the fact that they're machine? Is it the fact that they're big? Is it the fact that they cost a lot of money? Here's the thing: all human-built artificial flight (AF) machines are incredibly specialized and are far away from being able to perform most of the tasks birds -- the only general flight (GF) machines we are aware of -- can perform. More than 200 years after hot air balloons became operational and more than 100 years after the first planes flew, it's clear that building a GF machine is much harder than anticipated and that we are nowhere close to reaching bird-level abilities. 1. Planes vs eagles First, take a look at this video of an eagle catching a goat, throwing it off a cliff, and then feasting on it: I haven't ever seen a plane capable of catching a live animal and deliberately throwing it off a cliff. Not in 1922, not in 2022. Not even a tech demo. Such a feat vastly exceeds the abilities of any planes we have built, however fast they can fly. 2. Planes vs cuckoos Second, let's watch this video of a cuckoo chick ejecting the eggs of its competitors out of a nest: You could say that this ability has nothing to do flight but, again, this misses the forest for the trees. Building a GF machine is not about Goodharting random "flight" benchmarks by flying high and fast, it's about real-world performance on tasks GF machines created by nature are capable of. And, however impressive planes are, as soon as we try to see how well they perform in the real-world, they can't even match a cuckoo chick. 3. Planes vs a hummingbirds Third and final example. Take a look at the hummingbird's amazing ability to maintain stability in the harshest aerial conditions: Take any plane we have built and it stands no chance of survival placed in anything even close to these kinds of conditions, while a tiny-yet-mighty hummingbird doesn't break a sweat navigating essentially a tornado. Future of bird jobs: no plane danger Birds can flap their wings up to three times per second, whereas the fastest human-made aircraft only flaps its wings at 0.3 times per second. Birds can fly for long periods of time, whereas airplanes need to refuel regularly. Birds use orders of magnitude less energy to lift the same amount of mass in the air, compared to planes. Planes, rockets, and helicopters are (optimistically) decades away from being able to carry out most of the tasks birds are capable of. Therefore, for the foreseeable future, most bird jobs such as carrying messages (pigeons), carrying cargo (pigeons), hunting (hawks), and others, will remain safe from being displaced by human-built AF machines. Even if planes start to approach birds in some of their abilities, birds will be able to simply move towards performing other jobs. For example, planes can't navigate by themselves. So perhaps they will carry messages in simple conditions or to short distances, while pigeons will move towards specializing in complex message carrying or will learn to supervize plane routing, e.g. by piloting planes or by flying alongside and course...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Tyranny of the Epistemic Majority, published by Scott Garrabrant on November 22, 2022 on LessWrong. This post is going to mostly be propaganda for Kelly betting. However, the reasons presented in this post differ greatly from the reasons people normally use to argue for Kelly betting. The Steward of Myselves The curse of uncertainty is that I must make decisions that simultaneously affect many different versions of myself. When I close my eyes and then flip a coin, there are two potential versions of me: one sitting in front of a coin showing heads, the other sitting in front of a coin showing tails. Both of these potential versions of me are stakeholders in my current decisions. How can I make decisions on behalf of these multiple stakeholders? If it is a fair coin, then we can think of these two potential selves as equal stakeholders in my decisions. However, I know that it is not a fair coin. It has a 60 percent chance of coming up heads. Thus, heads-me is a 60 percent stakeholder in my current decisions, and tails-me is a 40 percent stakeholder. They amount of each one's stake is naturally in proportion to the probability that they actually exist. You, however, do not know if it is a fair coin, and are offering me a fair bet. I only have 100 dollars to my name, and I am can bet as much as I want (up to 100 dollars) in either direction at even odds. If I bet 100 dollars on heads, heads-me gets 200 dollars, and tails-me gets nothing. If I bet 100 dollars on tails, tails-me gets 200 dollars, and heads me gets nothing. If I bet nothing, both versions of me get 100 dollars. However, every dollar in the hands of heads-me is worth 1.5 times as much as a dollar in the hands of tails-me, since heads-me exists 1.5 times as much. (I am ignoring here any diminishing returns in my value of money.) Thus, to maximize value I should bet 100 dollars on heads. However, maybe it is better to think of tails-me as the rightful owner of 40 percent of my resources. When I bet 100 dollars on heads, I am seizing money from tails-me for the greater good, since heads-me has the (proportionally greater) existence necessary to better take advantage of it. Alternatively, I could say that since 60 percent of me is heads-me, heads me should only control 60 dollars, which can be bet on heads. Tails me should control 40 dollars, which can be bet on tails. These two bets partially cancel each other out, and the net result is that I bet 20 dollars on heads. If you are especially fast at maximizing expected logarithms, you might see where this is going. Compositionality Now, I am ready to introduce my friend, Kelly. Kelly also has her eyes closed, also has 100 dollars, and is sitting in front of the same coin. However, Kelly has different beliefs. Kelly believes that the coin has a 90 percent chance of coming up tails, and Kelly also has 100 dollars. I bet 20 dollars on heads, for the reasons described above. Kelly bets 80 dollars on tails for similar reasons (90 dollars on tails, partially nullified by 10 dollars on heads). I have another friend, Marge. Marge is sitting on the other side of the table with her eyes closed. Marge has 200 dollars. Marge doesn't know much about coins, but knows my and Kelly's beliefs, and thinks Kelly and I are equally likely to be correct. Thus, Marge assigns a 65 percent chance that the coin comes up tails. Marge thus bets 60 dollars on tails (130 dollars on tails, partially nullified by 70 dollars on heads). Note that the 60 dollars bet by Marge is the same as the net 60 dollar bet you get if you draw a box around me and Kelly. This is representing the compositionality of this betting policy. When you draw a box around me and Kelly, you can think of us as one agent whose wealth is the sum of our wealths, and whose beliefs are the weighted (by wealth) average of our ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Career Scouting: Dentistry, published by koratkar on November 20, 2022 on LessWrong. As a high school student, I worry a great deal over my future profession. According to Cal Newport, career satisfaction for any choice of occupation often won't materialize until you've become "so good they can't ignore you" at what you do. Based on this, Newport recommends directing your nervous energy towards building skill in whatever you choose, rather than choosing the "right job". While I found his advice useful for framing the issue, it opened a new box of concerns: What if you chose something you have little innate ability in – or what if the point at which your improvement slows down is not exceptional? How do you even know whether you have talent in something before you've invested effort into shadowing or interning? Sometimes talent doesn't appear initially – what if the thing you have the greatest potential in is something you'll have to struggle at for a long time? And there are so many jobs! We have to narrow the search space! There's another panoply of concerns related to your worth to the world: if you care about the world and your impact in it, don't you owe it to those who worked hard to give you the opportunities you have to find the thing you'll be the best at? But what if the thing you'll end up being the best at is something that doesn't always scale well – like medicine? (Of course, you could go into research, but what if you're only mid-tier at that?) You'll end up positively affecting the lives of many fewer people than you could have! And how do you avoid choosing work dry of meaning? People tell me I over-think these things – the answer to most (if not all) of my what-if's is "nothing interesting will happen in this case or the counter-factual one - you are ultimately insignificant in the greater procession of the world, and you live a cushy live in the developed world, meaning that no matter how badly you screw things up, as long as you don't get addicted to heroine, you will still have access to food, water, and shelter." While I think that answer is probably the right one, it looks like most young people don't think about this much at all! I have a seriously useful bone to throw them from my side of the anxiety fence! In earnest of providing information to a batch of high schoolers staring at fog-covered futures, a class of college students with sinking intimations that they chose the wrong major, and a karass of adults who curse their occupations with every breath, I'm compiling a database of rationalist-inspired interviews with members of various professions over at careerscouting.substack.com. I hope to respond to the wordless disconcertions about life that must bubble inside most people with answered questions. Below is my first interview with a general dentist. Object-Level What does a normal day in your field look like? Can you give me a “day in the life” kind of run-down? “I start early in the morning. I leave home at 6:15. My first patient is seated at 7:00. Then it’s non-stop until 1:00, when I have a fifty minute lunch. After 1:50, I continue until 5:00. Some days I work through lunch.” How does this differ from the average practitioner? “It doesn’t differ – almost every dentist works the same way.” How does your time split across different kinds of activities? “Aside from operative work, I have patient notes, lab work (e.g. pouring models of people’s mouths), hygiene exams, and treatment conferences (consulting the patient about what is going to happen during their treatment).” What does a bad day in your field look like, and how does your definition differ from the average person’s? “Every dentist faces a bad day where nothing is going right. I can’t even begin to explain the parameters of a bad day – patient anxiety, patients crying in the c...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: ARC paper: Formalizing the presumption of independence, published by Erik Jenner on November 20, 2022 on LessWrong. (I did not have anything to do with this paper and these are just my own takes.) The Alignment Research Center recently published their second report, Formalizing the presumption of independence. While it's not explicitly about AI alignment, it's probably still interesting for some people here. Summary The paper is about "heuristic arguments". These are similar to proofs, except that their conclusions are not guaranteed to be correct and can be overturned by counterarguments. Mathematicians often use these kinds of arguments, but in contrast to proofs, they haven't been formalized. The paper mainly describes the open problem of finding a good formalization of heuristic arguments. They do describe one attempt, "cumulant propagation", in Appendix D, but point out it can behave pathologically. So what's the "presumption of independence" from the title? Lots of heuristic arguments work by assuming that some quantities are independent to simplify things, and that's what the paper focuses on. Such an argument can be overturned by showing that there’s actually some correlation we initially ignored, which should then lead to a more sophisticated heuristic argument with a potentially different conclusion. What does this have to do with alignment? The paper only very briefly mentions alignment (in Appendix F), more detailed discussion is planned for the future. But roughly: Avoiding catastrophic failures. Heuristic arguments can let us better estimate the probability of rare failures, or failures which occur only on novel distributions where we cannot easily draw samples. This can be used during validation to estimate risk, or potentially during training to further reduce risk. Eliciting latent knowledge. Heuristic arguments may let us see “why” a model makes its predictions. We could potentially use them to distinguish cases where similar behaviors are produced by very different mechanisms—for example distinguishing cases where a model predicts that a smiling human face will show up on camera because it predicts there will actually be a smiling human in the room, from cases where it makes the same prediction because it predicts that the camera will be tampered with. [...] Neither of these applications is straightforward, and it should not be obvious that heuristic arguments would allow us to achieve either goal. [...] Heuristic arguments can be seen as somewhere between interpretability and formal verification: unlike interpretability, heuristic arguments are meant to be machine-checkable and don't have to be human-understandable. But unlike formal proofs, they don't require perfect certainty and might be much easier to find. Readers here might also be reminded of Logical Induction. This paper is trying to do something somewhat different though: [Approaches to logical uncertainty] have primarily focused on establishing coherence conditions and on capturing inductive reasoning, i.e. ensuring that a reasoner eventually successfully predicts φ(n) given observations of φ(1), φ(2), . . . φ(n − 1). These systems would not automatically recognize intuitively valid heuristic arguments [...], although they would eventually learn to trust these arguments after observing them producing good predictions in practice. Indeed, we can view ourselves as reasoners in exactly this situation, trying to understand and formalize a type of reasoning that appears to often make good predictions in practice. Formalizations of inductive reasoning may help clarify the standards we should use for evaluating a proposed heuristic estimator, but do not constitute a good heuristic estimator themselves. So should you read the paper? Given it's a 60-page report (though most of that's appendices) ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: When should we be surprised that an invention took “so long”?, published by jasoncrawford on November 16, 2022 on LessWrong. My first highly popular essay was “Why did we wait so long for the bicycle?” I’ve asked the same question of the cotton gin and the threshing machine. Others have asked it of the steam engine and the wheel. Recently Brian Potter asked it about wind power and Anton Howes about semaphore signaling systems. See more examples here. When asking these questions, we should think about when the question even needs an answer. That is, “why did it take so long” is only interesting if it took an abnormally long amount of time. Here’s my model for this. First, an invention is not going to happen at all if (1) it’s not technically possible or (2) there’s no market for it. Gate (1), technical possibility, could include, for example: Scientific foundations. No light bulb before electromagnetism, no antibiotics before the germ theory. Components. Airplanes were not possible before the internal combustion engine. Materials. Skyscrapers could only be built once cheap steel girders were available. Manufacturing techniques. Precision machining was necessary to make the gears, sprockets, chains, bearings, and other parts for a wide variety of inventions, probably including the threshing machine and the bicycle. Gate (2), the market, is whether it can be done commercially at a price that anyone will pay. If someone does make an invention there is no market for, it doesn’t go anywhere, and we might not even hear about it, because it is unlikely to make the history books. In any case, it wouldn’t affect the world, because it wouldn’t get distribution, and so it wouldn’t be historically relevant for our purposes. You see examples of this from time to time, such as the Korean movable-type printing press that predated Gutenberg. Note that the bar inventions have to meet is not just a proof of concept: they have to be sufficiently powerful, efficient, and reliable to be of practical use. Early computing machines were too slow; early threshing machines broke down too frequently; early light bulbs burned out too quickly. These are technically interesting prototypes, but not true inventions. An invention does not merely demonstrate a concept, it solves a problem—the whole problem, not just a part of it, even if it is the most visible or obvious part. An invention has to be practical. (See more discussion of this here.) Once something is technically possible and economically viable, then the clock starts ticking on how long we “waited” for it. But invention is a human process, and it’s not instantaneous. There is no perfectly efficient market in which an invention springs to life immediately as soon as it’s viable. It takes time, effort, trial and error. People have to decide to do something—they have to get the idea, and be sufficiently inspired and motivated to devote full-time efforts to something unknown, with an indefinite timeline and uncertain rewards. (Only a minority of people even have the temperament for this; in this sense, I agree with Anton Howes that innovation is not simply “in human nature.”) Then they have to get free to do it: at any given time, most inventors will be busy with projects, and only a subset will be looking for something new to do. They may have to acquire resources or recruit help, which takes time. Once they finally get to work, they have to experiment with approaches, discard failures, get new ideas, iterate. Given the nature of that process, there are several factors that affect the time that elapses before an invention. I have written about many of them before in the context of “flywheels of progress.” Here are some that I would call out specifically regarding the invention process: Total amount of R&D effort in the world, or in a specifi...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Don't design agents which exploit adversarial inputs, published by TurnTrout on November 18, 2022 on LessWrong. Summary. Consider two common alignment design patterns: Optimizing for the output of a grader which evaluates plans, and Fixing a utility function and then argmaxing over all possible plans. These design patterns incentivize the agent to find adversarial inputs to the grader (e.g. "manipulate the simulated human grader into returning a high evaluation for this plan"). I'm pretty sure we won't find adversarially robust grading rules. Therefore, I think these alignment design patterns are doomed. In this first essay, I explore the adversarial robustness obstacle. In the next essay, I'll point out how this is obstacle is an artifact of these design patterns, and not any intrinsic difficulty of alignment. Thanks to Erik Jenner, Johannes Treutlein, Quintin Pope, Charles Foster, Andrew Critch, randomwalks, and Ulisse Mini for feedback. 1: Optimizing for the output of a grader One motif in some AI alignment proposals is: An actor which proposes plans, and A grader which evaluates them. For simplicity, imagine we want the AI to find a plan where it makes an enormous number of diamonds. We train an actor to propose plans which the grading procedure predicts lead to lots of diamonds. In this setting, here's one way of slicing up the problem: Outer alignment: Find a sufficiently good grader. Inner alignment: Train the actor to propose plans which the grader rates as highly possible (ideally argmaxing on grader output, but possibly just intent alignment with high grader output). This "grader optimization" paradigm ordains that the AI find plans which make the grader output good evaluations. An inner-aligned actor is singlemindedly motivated to find plans which are graded maximally well by the grader. Therefore, for any goal by which the grader may grade, an inner-aligned actor is positively searching for adversarial inputs which fool the grader into spitting out a high number! In the diamond case, if the actor is inner-aligned to the grading procedure, then the actor isn't actually aligned towards diamond-production. The actor is aligned towards diamond-production as quoted via the grader's evaluations. In the end, the actor is aligned to the evaluations. I think that there aren't clever ways around this issue. Under this motif, under this way of building an AI, you're not actually building an AI which cares about diamonds, and so you won't get a system which makes diamonds in the limit of its capability development. Two clarifying points: This motif concerns how the AI makes decisions—this isn't about training a network using a grading procedure, it's about the trained agent being motivated by a grading procedure. The grader doesn't have to actually exist in the world. The "grader" can be a mathematical expected utility function over all action-sequences which the agent could execute. For example, it might take the action sequence and the agent's current beliefs about the world, and e.g. predict the expected number of diamonds produced by the actions. "The AI optimizes for what humanity would say about each universe-history" is an instance of grader-optimization, but "the AI has human values" is not an instance of grader-optimization. The parable of evaluation-child an AI should optimize for the real-world things I value, not just my estimates of those things. — The Pointers Problem: Human Values Are A Function Of Humans' Latent Variables First, a mechanistically relevant analogy. Imagine a mother whose child has been goofing off at school and getting in trouble. The mom just wants her kid to take education seriously and have a good life. Suppose she had two (unrealistic but illustrative) choices. Evaluation-child: The mother makes her kid care extremely strongly abou...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Announcing the Progress Forum, published by jasoncrawford on November 17, 2022 on LessWrong. I’d like to invite you to join the Progress Forum, the new online home for the progress community. It's a clone of this site, but with a focus on progress studies and the philosophy of progress. This forum was pre-announced in January 2022, and quietly opened in April. Although anyone could sign up, we deliberately didn’t make any big announcement about it, aiming first for a small, high-quality community. Now that we have a lot of good content on the site, we’re announcing it more broadly. The primary goal of this forum is to provide a place for long-form discussion of progress studies. It’s also, like LW, a place to find local clubs and meetups. The broader goal is to share ideas, strengthen them through discussion and comment, and over the long term, to build up a body of thought that constitutes a new philosophy of progress for the 21st century (and beyond). I invite you to post: Essays (original, or cross-posted from your blog) Drafts, half-baked ideas, and work-in-progress thinking, for feedback Questions for brainstorming Local events and community groups Etc. And please read and comment on what others have shared. You can subscribe to Forum posts via email, RSS, or Twitter. The Forum is sponsored by The Roots of Progress. Huge thanks to the people who worked to create and run it: Lawrence Kestleoot, Andrew Roberts, Sameer Ismail, David Smehlik, Alec Wilson, and Ross Graham. Thanks also to Kris Gulati for nudging this project along, and to Ruth Grace Wong for helpful conversations about community and moderation. Finally, thanks to the LessWrong team for creating this software platform, and especially to Oliver Habryka, Ruby Bloom, Raymond Arnold, JP Addison, James Babcock, and Ben Pace for answering questions and helping us customize this instance of it. Go check it out. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Current themes in mechanistic interpretability research, published by Lee Sharkey on November 16, 2022 on LessWrong. This post gives an overview of discussions - from the perspective and understanding of the interpretability team at Conjecture - between mechanistic interpretability researchers from various organizations including Conjecture, Anthropic, Redwood Research, OpenAI, and DeepMind as well as some independent researchers. It is not a review of past work, nor a research agenda. We're thankful for comments and contributions from Neel Nanda, Tristan Hume, Chris Olah, Ryan Greenblatt, William Saunders, and other anonymous contributors to this post, which greatly improved its quality. While the post is a summary of discussions with many researchers and received comments and contributions from several, it may nevertheless not accurately represent their views. The last two to three years have seen a surge in interest in mechanistic interpretability as a potential path to AGI safety. Now there are no fewer than five organizations working on the topic (Anthropic, Conjecture, DeepMind, OpenAI, Redwood Research) in addition to numerous academic and independent researchers. In discussions about mechanistic interpretability between a subset of researchers, several themes emerged. By summarizing these themes here, we hope to facilitate research in the field more broadly. We identify groups of themes that concern: Object-level research topics in mechanistic interpretability Research practices and tools in mechanistic interpretability Field building and research coordination in mechanistic interpretability Theories of impact for mechanistic interpretability Object-level research topics in mechanistic interpretability Solving superposition Anthropic’s recent article on Toy Model of Superposition laid out a compelling case that superposition is a real phenomenon in neural networks. Superposition appears to be one of the reasons that polysemanticity happens, which makes mechanistic interpretability very difficult because it prevents us from telling simple stories about how features in one layer are constructed from features in previous layers. A solution to superposition will look like the ability to enumerate all the features that a network represents, even if they’re represented in superposition. If we can do that, then we should be able to make statements like “For all features in the neural network, none violate rule X” (and more ambitiously, for "no features with property X participate in circuits which violate property Y"). Researchers at Anthropic hope this might enable ‘enumerative safety’, which might allow checking random samples or comprehensive investigations of safety-critical parts of the model for unexpected and concerning components. There are many potential reasons researchers could fail to achieve enumerative safety, including failing to solve superposition, scalability challenges, and several other barriers described in the next section. Anthropic outlined several potential solutions to superposition in their article. Very briefly, these strategies are: Create models without superposition. Find a sparse overcomplete basis that describes how features are represented in models with superposition. This will likely involve large scale solutions to sparse coding. Hybrid approaches in which one changes models, not resolving superposition, but making it easier for a second stage of analysis to find a sparse overcomplete basis that describes it. Multiple organizations are pursuing these strategies. Researchers in all organizations are keen to hear from people interested in working together on this problem. However, there is a range of views among researchers on how central superposition is as a problem and how tractable it is. Barriers beyond superposition? We’ve be...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Will we run out of ML data? Evidence from projecting dataset size trends, published by Pablo Villalobos on November 14, 2022 on LessWrong. Summary: Based on our previous analysis of trends in dataset size, we project the growth of dataset size in the language and vision domains. We explore the limits of this trend by estimating the total stock of available unlabeled data over the next decades. Read the full paper in arXiv. Our projections predict that we will have exhausted the stock of low-quality language data by 2030 to 2050, high-quality language data before 2026, and vision data by 2030 to 2060. This might slow down ML progress. All of our conclusions rely on the unrealistic assumptions that current trends in ML data usage and production will continue and that there will be no major innovations in data efficiency. Relaxing these and other assumptions would be promising future work. Historical projectionCompute projectionLow-quality language stock2032.4[2028.4 ; 2039.2] 2040.5[2034.6 ; 2048.9]High-quality language stock2024.5[2023.5 ; 2025.7]2024.1[2023.2 ; 2025.3]Image stock2046[2037 ; 2062.8]2038.8[2032 ; 2049.8]Table 1: Median and 90% CI exhaustion dates for each pair of projections. Background Chinchilla's wild implications argued that training data would soon become a bottleneck for scaling large language models. At Epoch we have been collecting data about trends in ML inputs, including training data. Using this dataset, we estimated the historical rate of growth in training dataset size for language and image models. Projecting the historical trend into the future is likely to be misleading, because this trend is supported by an abnormally large increase in compute in the past decade. To account for this, we also employ our compute availability projections to estimate the dataset size that will be compute-optimal in future years using the Chinchilla scaling laws. We estimate the total stock of English language and image data in future years using a series of probabilistic models. For language, in addition to the total stock of data, we estimate the stock of high-quality language data, which is the kind of data commonly used to train large language models. We are less confident in our models of the stock of vision data because we spent less time on them. We think it is best to think of them as lower bounds rather than accurate estimates. Results Finally, we compare the projections of training dataset size and total data stocks. The results can be seen in the figure above. Datasets grow much faster than data stocks, so if current trends continue, exhausting the stock of data is unavoidable. The table above shows the median exhaustion years for each intersection between projections. In theory, these dates might signify a transition from a regime where compute is the main bottleneck to growth of ML models to a regime where data is the taut constraint. In practice, this analysis has serious limitations, so the model uncertainty is very high. A more realistic model should take into account increases in data efficiency, the use of synthetic data, and other algorithmic and economic factors. In particular, we have seen some promising early advances on data efficiency, so if lack of data becomes a larger problem in the future we might expect larger advances to follow. This is particularly true because unlabeled data has never been a constraint in the past, so there is probably a lot of low-hanging fruit in unlabeled data efficiency. In the particular case of high-quality data, there are even more possibilities, such as quantity-quality tradeoffs and learned metrics to extract high-quality data from low-quality sources. All in all, we believe that there is about a 20% chance that the scaling (as measured in training compute) of ML models will significantly slow down b...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Alignment Community Is Culturally Broken, published by sudo -i on November 13, 2022 on LessWrong. Disclaimer: These are entirely my thoughts. I'm posting this before it's fully polished because it never will be. Epistemic status: Moderately confident. Deliberately provocative title. Apparently, the Bay Area rationalist community has a burnout problem. I have no idea if it's worse than base rate, but I've been told it's pretty bad. I suspect that the way burnout manifests in the rationalist community is uniquely screwed up. I was crying the other night because our light cone is about to get ripped to shreds. I'm gonna do everything I can to do battle against the forces that threaten to destroy us. You've heard this story before. Short timelines. Tick. Tick. I've been taking alignment seriously for about a year now, and I'm ready to get serious. I've thought hard about what my strengths are. I've thought hard about what I'm capable of. I'm dropping out of Stanford, I've got something that looks like a plan, I've got the rocky theme song playing, and I'm ready to do this. A few days later, I saw this post. And it reminded me of everything that bothers me about the EA community. Habryka covered the object level problems pretty well, but I need to communicate something a little more... delicate. I understand that everyone is totally depressed because qualia is doomed. I understand that we really want to creatively reprioritize. I completely sympathize with this. I want to address the central flaw of Akash+Olivia+Thomas's argument in the Buying Time post, which is that actually, people can improve at things. There's something deeply discouraging about being told "you're an X% researcher, and if X>Y, then you should stay in alignment. Otherwise, do a different intervention." No other effective/productive community does this. I don't know how to put this, but the vibes are deeply off. The appropriate level of confidence to have about a statement like "I can tell how good of an alignment researcher you will be after a year of you doing alignment research" feels like it should be pretty low. At a year, there's almost certainly ways to improve that haven't been tried. Especially in a community so mimetically allergic to the idea of malleable human potential. Here's a hypothesis. I in no way mean to imply that this is the only mechanism by which burnout happens in our community, but I think it's probably a pretty big one. It's not nice to be in a community that constantly hints that you might just not be good enough and that you can't get good enough. Our community seems to love treating people like mass-produced automatons with a fixed and easily assessable "ability" attribute. (Maybe you flippantly read that sentence and went "yeah it's called g factor lulz." In that case, maybe reflect on good of a correlate g is in absolute terms for the things you care about.). If we want to actually accomplish anything, we need to encourage people to make bigger bets, and to stop stacking up credentials so that fellow EAs think they have a chance. It's not hubris to believe in yourself. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: "Normal" is the equilibrium state of past optimization processes, published by Alex Altair on October 30, 2022 on LessWrong. Orienting around the ideas and conclusions involved with AI x-risk can be very difficult. The future possibilities can feel extreme and far-mode, even when we whole-heartedly affirm their plausibility. It helps me to remember that everything around me that feels normal and stable is itself the result of an optimization process that was, at the time, an outrageous black swan. Modernity If you were teleported into the body of a random human throughout history, then most likely, your life would look nothing like the present. You would likely be a hunter-gatherer, or perhaps a farmer. You would be poor by any reasonable standard. You would probably die as a child. You would have nothing resembling your current level of comfort, and your modern daily life would be utterly alien to most humans. What currently feels normal is a freeze-frame state of a shriekingly fast feedback loop involving knowledge, industry, and population. It is nowhere near normal, and it is nowhere near equilibrium. The Anthropocene For hundreds of millions of years, the earth was a wilderness, teeming with life. The teeming had ebbs and flow, evolutionary breakthroughs and power struggles, but essentially, it was a type of equilibrium. For hundreds of thousands of years, humans were just another animal. We ran around the savanna, killed other animals and foraged for food, had sex, slept, and experienced joys and losses. If an alien were gazing at earth from afar, they would have had no reason to think humans were different from any other animal. But all the while, humans were noticing things. They were curious, and desperate, and they had a capacity that the other animals didn't have. They got ideas about the world, and tested them with their hands, and gave this knowledge to their children with their words. Over time, the humans became conspicuous, and eventually, they transformed every inch of their surroundings. For a while now, humans have been a threat to the species they live with. If you are a fish or a weed or a polar bear, your biggest problem is humans. But as John Green points out in The Anthropocene Reviewed, in the 21st century, if you are the atmosphere or a river or a desert, your biggest problem is humans. The rise of humanity is no longer just an ecological phenomenon; it has kicked off a new geological epoch. The fate of the rock itself is in our hands, in a way that it never was for the other mammals, or the dinosaurs, or the trilobites. The invention of wood One day, plants invented wood. This was great for plants, but it had externalities. As wikipedia puts it; The evolution of the wood fiber lignin and the bark-sealing, waxy substance suberin variously opposed decay organisms so effectively that dead materials accumulated long enough to fossilise on a large scale. I don't know exactly how this went down. But I like to imagine watching a forest grow on fast-forward; trees spring up, growing taller and taller, and then eventually each tree dies. Only, when these proto-trees falls down, the logs does not sink into the forest floor and dissolve into biodegraded soil. They just stay there. The tree trunks pile up for millions of years. This was the new equilibrium. The plant matter went through some chemical changes that compressed and homogenized it, but much of the plant's stored energy remained. The energy sank deep underground as what we now call fossil fuels. Bacteria and fungi continued evolving, and eventually they figured out how to eat the now-abundant tougher plant matter. But the coal deposits were now beyond their reach. The historical details of these transitions are not well understood, but what is clear is that something wild happened, the surfa...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Am I secretly excited for AI getting weird?, published by porby on October 29, 2022 on LessWrong. This post is arguably darker than my other one. I don't make any persuasive arguments about AI forecasting here; if you don't feel like looking at doominess, feel free to skip this. I've noticed a few instances of what look like people assuming that those who are visibly concerned about AI risk don't really buy into the full weight of what they're saying. Recently, I came across this (hi, niknoble!): As a specific example of what I suspect is a bit of cognitive dissonance, look at the recent post on AGI by porby, which predicts AGI by 2030. I loved reading that post because it promises that the future is going to be wild. If porby is right, we're all in for an adventure. Based on the breathless tone of the post, I would surmise that porby is as excited by his conclusion as I am. For example, we have this excerpt: This is crazy! I'm raising my eyebrows right now to emphasize it! Consider also doing so! This is weird enough to warrant it! Would you have predicted this in 2016? I don't think I would have! Does this strike you as someone who dreads the arrival of AGI? It seems to me like he is awaiting it with great anticipation. But then in the comments on the post, he says that he hopes he's wrong about AGI! If you're reading this porby, do you really want to be wrong? This is an excellent example of the kind of thing I'm talking about, so I'm going to use it. I think my writing and speaking style defaults to a kind of lightness that can be misleading. So let me try to write something a little darker. Well, do you? Because I don't think P(doom | AGI) is anywhere close to 0, especially for AGI developed on very short timescales:YES, I DO WANT TO BE WRONG. The kind of "excitement" I feel about near-term AGI is adjacent to hearing the tornado siren, looking at the radar, seeing the warned cell moving straight east, walking out on my porch to look at a black wall of rain a mile or two away, and seeing the power flashes straight west of me as the tornado rips lives apart. While grabbing a mattress to throw over a tub, I'm doing some quick mental calculations- the statistical rarity of EF-3 or stronger tornadoes, will it stay on the ground, how large is it (glance at the hook on the reflectivity map), how sturdy is this house (the feeling of the entire house shunting to one side during an earlier storm's 120 mph winds wasn't promising), how much damage would a near miss cause? All the while, telling my family to get their shoes, don't worry, we have time (do we have time? probably), just get into the bathroom. It didn't stay on the ground. Also, we have a storm shelter now. It only took about 6 close calls to bite that bullet! More than excitement You know that voyeuristic "excitement" of a really bad hurricane about to make landfall? Something wildly out of the ordinary, something that breaks the sense of normalcy and reminds you that human civilization is fragile? It's a weird, darkly attractive kind of novelty. Watching COVID-19 in January 2020 felt that way, for a while. It was distant and not here, so it felt like an almost fictional threat. Within a few weeks, it became something else. With a sense of vertigo, I settled into the realization that it was really happening. I told my parents to go buy anything they wouldn't want to run out of if there was a run on it at the stores, because things are going to get bad, and that a lot of people were going to die. I explained they had medical backgrounds that put them at much higher risk, and hospitals might get overwhelmed soon, so they shouldn't go inside places with other people if they could avoid it. And if they do, wear masks. The transition of the threat to inevitable reality made me feel bad about the feeling of excitemen...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How Risky Is Trick-or-Treating?, published by jefftk on October 27, 2022 on LessWrong. content warning: discussion of child fatalities I've seen this chart going around: Sperling and Smith created it from NHTSA FARS data, and there have been lots of news stories about it or the underlying data. Some people have a "lets make costumes brighter" angle, some see this as ammunition in the fight against car-centric culture (which I support), some people want to move Halloween earlier in the day. But I think everyone is missing the point: Halloween is actually one of the safest days to be out walking as a kid. Let's look at the data. I started by replicating the chart to verify I was doing it correctly (code): The dip on February 29th is mostly just that this day only occurs once every four years, so it has about a quarter the deaths you'd expect. But what about the spike on 4/30? On April 30th 1992 someone drove their car into a large group of 3rd graders on a field trip, and one of them died. It turns out the FARS records can include cases where someone else died in the same accident. Filtering to just the records representing deaths I get: The data available actually goes back to 1982 and now runs to 2020, so we can expand it a bit: We can also break down by age, since "child" can mean a range of things and most people are likely thinking mostly about younger children: All this is to confirm the original chart: more children do die on Halloween than on a typical day. But I still don't think that's a reason to keep your kids home on Halloween, even when you set aside the consideration that Halloween is an especially fun time to be a child pedestrian. The issue is, there are a lot more kids out on Halloween than the typical day. If you learned that more people died in car crashes in the US (42k/y) than in Paraguay (1.6k/y) you wouldn't conclude that Paraguay's roads are safer: the US has 50x more people than Paraguay! We would ideally scale the initial chart by the number of children out walking on any given day, but I don't think anyone has numbers on this. Very roughly, though, maybe it's something like 5x as many people (kids, parents, grandparents) out on foot on Halloween than the typical day? Every pedestrian is rolling some dice. The dice are pretty favorable: of the ~70M people under age 18 in the US only 320 died in 2020 by being hit by a car. But you still might want to avoid days where the dice were less favorable than usual. Looking at the original chart you might conclude that Halloween was one of those days, but if there are a lot more kids walking around on Halloween it's actually one of the safer days. Comment via: facebook Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Signals of war in August 2021, published by yieldthought on October 26, 2022 on LessWrong. Could we have predicted Russia's intention to invade Ukraine earlier? Back in August 2021 the energy prices in Europe doubled and stayed at this level until the start of hostilites: This, of course, did not go unnoticed. Many, many articles were written about it. Shockingly, a wide range of unsubstantiated theories for the rise were presented in these, often with a suggestion that carbon taxes or decarbonisation was the root cause. Actually, the root cause was an increase in natural gas prices. This was obvious at the time to anyone who knows how the energy market price is determined: The cause of the gas price increase was that Russia suddenly restricted its supplies to Europe, which was also noticed and reported on as in this CNBC article titled "Russia is pumping a lot less natural gas to Europe all of a sudden — and it is not clear why". The effect was to dramatically lower European gas reserves, which was also noticed: Why would Russia deliberately reduce European gas reserves? To increase their leverage, ideally timed at some point in the future when Europe needs that gas the most. This has to be relatively short-term, as if supply remains low and prices high over the long term Europe will choose to fill its reserves anyway, giving us three main sweet spots: Late 2021: reserves insufficient to cover the winter and demand high. Early 2022: reserves at their lowest point, but demand decreasing for 9 months. Late 2022: demand increasing but risks that reserves have had time to refill. What conclusions could we draw from information in August 2021 about this? Well, by April of 2021 Russia had amassed significant and unusual forces on the border with Ukraine, which was also well-reported: Russia claimed the forces were on training exercises. Western spectators mostly seem to have assumed sabre-rattling. That interpretation is inconsistent with reducing gas supply to Europe over a period of months to ensure gas supplies will be depleted after the winter of 2021. The logical conclusion from Russia's military buildup (which intensified in October 2021) must have been that Putin plans to invade, ideally in late 2021 but at the latest in early 2022. With 20/20 hindsight it's easy to cherry-pick evidence. So I have a few questions: Who drew this connection in August / October 2021? I haven't found anything but would love to update on these people's current analysis of events. What's was the steel-man case in August / October 2021 that Putin is not preparing for war - which other concrete information out there was strong enough to overlook these clear signs of economic and military preparation for war? Is there evidence that the invasion was delayed by 1-2 months? The economic timing suggests Putin would have had more leverage if he had invaded at the start of winter. Indeed, much was made of the spring thawing making the invasion harder (soft ground channels heavy vehicles along roads, making them more vulnerable) and the actual timing meant Europe was able to refill its gas reserves ahead of the 2022 winter. This is a rare opportunity to look back and update on world-model predictions. The information was there. Who saw it and made the case? Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Empowerment is All We Need, published by jacob cannell on October 23, 2022 on LessWrong. Intro What/who would you like to become in a thousand subjective years? or a million? Perhaps, like me, you wish to become posthuman: to transcend mortality and biology, to become a substrate independent mind, to wear new bodies like clothes, to grow more intelligent, wise, wealthy, and connected, to explore the multiverse, perhaps eventually to split, merge, and change - to vasten. Regardless of who you are now or what specific values you endorse today, I suspect you too would at least desire these possibilities as options. Absent some culture specific social stigmas, who would not like more wealth, health, and power? more future optionality? As biological creatures, our fundamental evolutionary imperative is to be fruitful and multiply, so our core innate high level value should be inclusive genetic fitness. But for intelligent long lived animals like ourselves, reproduction is a terminal goal in the impossibly distant future: on the order of around 1e11 neural clock cycles from birth[1], to be more precise. Explicit optimization of inclusive genetic fitness through simulation and planning over such vast time horizons is simply implausible - especially for a mere 20 watt irreversible computer such as the human brain, no matter how efficient. Fortunately there exists an accessible common goal which is ultimately instrumentally convergent for nearly all final goals: power-seeking, or simply: empowerment. Omohundro proposed an early version of the instrumental convergence hypothesis as applied to AI in his 2008 paper the Basic AI Drives, however the same principle was already recognized by Klyubin et al in their 2005 paper "Empowerment: A Universal Agent-Centric Measure of Control"[2]: Our central hypothesis is that there exist a local and universal utility function which may help individuals survive and hence speed up evolution by making the fitness landscape smoother. The function is local in the sense that it doesn’t rely on infinitely long history of past experience, does not require global knowledge about the world, and that it provides localized feedback to the individual. To a sugar-feeding bacterium, high sugar concentration means longer survival time and hence more possibilities of moving to different places for reproduction, to a chimpanzee higher social status means more mating choice and interaction, to a person more money means more opportunities and more options. The common feature of the above examples is the striving for situations with more options, with more potential for control or influence. To capture this notion quantitatively, as a proper utility function, we need to quantify how much control or influence an animal or human (an agent from now on) has. Salge et al later summarized these arguments into the Behavioral Empowerment Hypothesis[3]: The adaptation brought about by natural evolution produced organisms that in absence of specific goals behave as if they were maximizing their empowerment. Empowerment provides a succinct unifying explanation for much of the apparent complexity of human values: our drives for power, knowledge, self-actualization, social status/influence, curiosity and even fun[4] can all be derived as instrumental subgoals or manifestations of empowerment. Of course empowerment alone can not be the only value or organisms would never mate: sexual attraction is the principle deviation later in life (after sexual maturity), along with the related cooperative empathy/love/altruism mechanisms to align individuals with family and allies (forming loose hierarchical agents which empowerment also serves). The key central lesson that modern neuroscience gifted machine learning is that the vast apparent complexity of the adult human brain, with al...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Mnestics, published by Jarred Filmer on October 23, 2022 on LessWrong. Take your mnestic Part IV In the web series "There is No Antimemetics Division" our heroes must fight a great evil, monsters that make you forget they exist. In the course of such adventures, they find it necessary to imbibe "mnestics" (the opposite of an amnesiac), drugs for remembering. Every morning since I left that cabin, I wake up having to remember only one thing. Take. your. mnestic To do this, I sit with my laptop and place a leaf between my lips. This looks odd if I happen to be in a cafe, but serves to remind me I was doing something important if I get distracted. It also helps to turn interruptions like "do you have a second?", into "Jarred, why are you eating a leaf?". I hit "ctrl+m" to pull up my desktop forever dedicated to "mnestic.txt" and my calendar. The contents of mnestic.txt have changed over time, but the quote at the top is always the same: Humans can forget anything. It's okay to forget some things, because we are mortal and finite. But some things we have to remember. It's important that we remember. Write to yourself something which will make you remember. There Is No Antimemetics Division I read it one word at a time, and try to recapture that sense of dread. The feeling of recalling something existentially important that needs sustained action over years. Knowing I'll forget. I read the second line, still a word at a time: A routine of self perpetuating, undistracted time to think is necessary to respond well to being alive. I try to recall the feeling I had sitting in that cabin, that "Space" is something both vital and fragile, and while I have it, I need to use it to ensure I don't lose it. It's at this stage I check my calendar to confirm the week's mnestics are all lined up, coming to me free and clear down the river of time. I read the remaining 8 lines from past selves who had access to a wealth of space, reminding me of how I want live. With my priorities freshly re-downloaded into my brain, I plan my day, set intentions, and make a todo list. Altogether this takes around 30 minutes. To finish, I try to capture in time who I feel I am now, and verbally extend a bridge to my future self. Tomorrow when sitting down to take his own mnestic, he will think of this moment and verbally accept that bridge from the past, creating a chain throughout time built of 30-minute pockets of daily space-taking. A highway down which compressed 4-day thoughts can stream from the past to find me here in the present. Every Sunday morning I take a larger dose, clearing a space in which to spend 2 hours trying to make sense of the last 168. I get a narrative sense of the week past, how it compares to the story I wanted to tell, and where that narrative wants to go over the next 7 days. I outline projects, make todos, and most important of all, visualise the structure of the days within which I'll work. It feels like gathering a Kamehameha of intention, shot wide and strong into the week to bring the shards of my agency scattered through time into alignment. This might mean 2 or 3 decisions made differently than I might have done otherwise, or it might mean 50. Leaving a party an hour early to get to sleep on time, saying no to meeting a school friend, saying no to another conference, remembering to take decaf in the afternoon, reading that article, organising that book club, promptly responding to important messages. These are things I can only do reliably when I'm in regular emotional contact with a picture of what actually moves me, ordinarily too big to fit through the door of my mind at short notice. But if one takes a massive shot of space with which to dwell on it, you can compress that picture down into a mnestic you can take daily. "Renew thyself completely each day; do it aga...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Deepfake(?) Phishing, published by jefftk on October 21, 2022 on LessWrong. I think someone just tried to phish my Facebook account, including a fake video of a FB friend. Here's the conversation: Them, via FB Messenger, 9:32am: Please ,I was trying to login in my instagram page on Facebook my new phone and they ask me to find someone to help me receive a code, Facebook gave me two friends suggestions and you one of them, the other person isn't online. will you Help me receive the code please? Me: I'm sorry you're having trouble logging in! Just so I can make sure your account hasn't been hacked, how did we meet? Them: [Calls me over FB Messenger, audio isn't working but it does look like them. I'm completely convinced at this point.] Me: Audio wasn't working, but I did recognize youWhat do you need me to do? 32665, over SMS: NNNNNNNN is your Facebook password reset code [this number has previously sent me FB resets] Them: Send me the code sent to you minute ago Me: Hmm, those look like the code to reset the password to my account?Can we call again? Me: [I try to call them back, doesn't go through] Them: Nahh it's for my instagram Them: Having bad connections here Them: Send me the code ? Me: sorry, I'm still worried your account has been hacked -- can we do another call? Them: [Calls me over FB Messenger, audio is still not working, and the video feels slightly off. Ends quickly on their end. Possibly it's even the same video from last time?] Me: We're you able to hear me? Them: My connections I've reported their account as hacked. Things that made me suspicious: I don't think FB has any sort of account recovery that looks like this This is exactly what an attempt to hack my FB account would look like 9:30am, even though that makes it 6:30 where they live Video call didn't have any audio They couldn't receive incoming video calls Text did't feel like them, though I don't know them that well. Here's a screenshot I took during the second video call: Even with all those things, the video call would normally have been very convincing, and it did briefly convince me. I could easily see it fooling someone who didn't know about deepfake video. Comment via: facebook Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Plans Are Predictions, Not Optimization Targets, published by johnswentworth on October 20, 2022 on LessWrong. Imagine a (United States) high-schooler who wants to be a doctor. Their obvious high-level plan to achieve that goal is: Graduate high school and get into college Go to college, study some bio/chem/physiology, graduate and get into med school Go to med school, make it through residency Doctor! Key thing to notice about that plan: the plan is mainly an optimization target. When in high school, our doctor-to-be optimizes for graduating and getting into college. In college, they optimize for graduating and getting into med school. Etc. Throughout, our doctor-to-be optimizes to make the plan happen. Our doctor-to-be does not treat the plan primarily as a prediction about the world; they treat it as a way to make the world be. And that probably works great for people who definitely just want to be doctors. Now imagine someone in 1940 who wants to build a solid-state electronic amplifier. Building active solid-state electronic components in the early 1940’s is not like becoming a doctor. Nobody has done it before, nobody knows how to do it, nobody knows the minimal series of steps one must go through in order to solve it. At that time, solid-state electronics was a problem we did not understand; the field was preparadigmatic. There were some theories, but they didn’t work. The first concrete plans people attempted failed; implicit assumptions were wrong, but it wasn’t immediately obvious which implicit assumptions. One of the most confident predictions one might reasonably have made about solid-state electronics in 1940 was that there would be surprises; unknown unknowns were certainly lurking. So, how should someone in 1940 who wants to build a solid-state amplifier go about planning? I claim the right move is to target robust bottlenecks: look for subproblems which are bottlenecks to many different approaches/plans/paths, then tackle those subproblems. For instance, if I wanted to build a solid-state amplifier in 1940, I’d make sure I could build prototypes quickly (including with weird materials), and look for ways to visualize the fields, charge densities, and conductivity patterns produced. Whenever I saw “weird” results, I’d first figure out exactly which variables I needed to control to reproduce them, and of course measure everything I could (using those tools for visualizing fields, densities, etc). I’d also look for patterns among results, and look for models which unified lots of them. Those are strategies which would be robustly useful for building solid-state amplifiers in many worlds, and likely directly address bottlenecks to progress in many worlds. In our particular world, they might have highlighted the importance of high-purity silicon and dopants, or of surfaces between materials with different electrical properties, both of which were key rate-limiting insights along the path to active solid-state electronics. Now, when looking for those robust bottlenecks, I’d probably need to come up with a plan. Multiple plans, in fact. The point of “robust bottlenecks” is that they’re bottlenecks to many plans, after all. But those plans would not be optimization targets. I don’t treat the plans as ways to make the world be. Rather, the plans are predictions about how things might go. My “mainline plan”, if I have one, is not the thing I’m optimizing to make happen; rather, it’s my modal expectation for how I expect things to go (conditional on my efforts). My optimization targets are, instead, the robust bottlenecks. When reality throws a brick through the plans, I want my optimization target to have still been a good target in hindsight. Thus robust bottlenecks: something which is still a bottleneck under lots of different assumptions is more likely to b...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Introduction to abstract entropy, published by Alex Altair on October 20, 2022 on LessWrong. This post, and much of the following sequence, was greatly aided by feedback from the following people (among others): Lawrence Chan, Joanna Morningstar, John Wentworth, Samira Nedungadi, Aysja Johnson, Cody Wild, Jeremy Gillen, Ryan Kidd, Justis Mills and Jonathan Mustin. Illustrations by Anne Ore. Introduction & motivation In the course of researching optimization, I decided that I had to really understand what entropy is. But there are a lot of other reasons why the concept is worth studying: Information theory: Entropy tells you about the amount of information in something. It tells us how to design optimal communication protocols. It helps us understand strategies for (and limits on) file compression. Statistical mechanics: Entropy tells us how macroscopic physical systems act in practice. It gives us the heat equation. We can use it to improve engine efficiency. It tells us how hot things glow, which led to the discovery of quantum mechanics. Epistemics (an important application to me and many others on LessWrong): The concept of entropy yields the maximum entropy principle, which is extremely helpful for doing general Bayesian reasoning. Entropy tells us how "unlikely" something is and how much we would have to fight against nature to get that outcome (i.e. optimize). It is relevant to the fate of the universe. And it's also a fun puzzle to figure out! I didn't intend to write a post about entropy when I started trying to understand it. But I found the existing resources (textbooks, Wikipedia, science explainers) so poor that it actually seems important to have a better one as a prerequisite for understanding optimization! One failure mode I was running into was that other resources tended only to be concerned about the application of the concept in their particular sub-domain. Here, I try to take on the task of synthesizing the abstract concept of entropy, to show what's so deep and fundamental about it. In future posts, I'll talk about things like: How abstract entropy can be made meaningful on continuous spaces Exactly where the "second law of thermodynamics" comes from, and exactly when it holds (which turns out to be much broader than thermodynamics) How several domain-specific types of entropy relate to this abstract version Many people reading this will have some previous facts about entropy stored in their minds, and this can sometimes be disorienting when it's not yet clear how those facts are consistent with what I'm describing. You're welcome to skip ahead to the relevant parts and see if they're re-orienting; otherwise, if you can get through the whole explanation, I hope that it will eventually be addressed! But also, please keep in mind that I'm not an expert in any of the relevant sub-fields. I've gotten feedback on this post from people who know more math & physics than I do, but at the end of the day, I'm just a rationalist trying to understand the world. Abstract definition Entropy is so fundamental because it applies far beyond our own specific universe, the one where something close to the standard model of physics and general relativity are true. It applies in any system with different states. If the system has dynamical laws, that is, rules for moving between the different states, then some version of the second law of thermodynamics is also relevant. But for now we're sticking with statics; the concept of entropy can be coherently defined for sets of states even in the absence of any "laws of physics" that cause the system to evolve between states. The example I keep in my head for this is a Rubik's Cube, which I'll elaborate on in a bit. The entropy of a state is the number of bits you need to use to uniquely distinguish it. Some useful things t...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Voting Theory Introduction, published by Scott Garrabrant on October 17, 2022 on LessWrong. Sequence Introduction This is the first post in a sequence in which I will propose a new voting system! In this post, I introduce the framework and notation, and give some background on voting theory. In the next post, I will show you the best voting system you've probably never heard of, maximal lotteries. (Seriously, it's really good.) After that, I will make it even better, and propose a new system: maximal lottery-lotteries. Then comes the bad news: I can't prove that maximal lottery-lotteries exist! (Or alternatively, good news: You can try to solve a cool new open problem in voting theory!) Thanks to Jessica Taylor for first introducing me to maximal lotteries, and Sam Eisenstat for spending many hours with me trying to prove the existence of maximal lottery-lotteries. Generalizing Voting Theory A voting system is a function that takes in a distribution on utility functions on a set of candidates, and produces a distribution on that set of candidates. This is not what a voting theorist will tell you a voting system is. That is because they like to make a bunch of extra assumptions: The set of candidates is finite. The function only uses the preorders on candidates implied by the utility functions. The output distribution assigns probability 1 to a single candidate. The input distribution is the uniform distribution on some finite set. I am going to play along with some of these assumptions, but I want my types and notation to treat them as explicit assumptions, so I will use notation that will make it easy to remove these assumptions as needed. Technically, I also make one assumption that voting theorists usually don't make. That is that the voting system is homogeneous. Homogeneous voting systems are only a function of what proportion of voters have each preference, not on the absolute number of voters that have each preference. Thus, when I said the input was a distribution on utility functions, rather than a multi-set of utility functions, I threw out the information that would have allowed for non-homogeneous voting systems. I don't know of any seriously proposed voting systems that are non-homogeneous, so this is not a significant extra assumption. Throughout the sequence, I will take for granted that all voting systems are homogeneous. Assumption 2 is most likely to be violated by voting theorists, as in approval voting and range voting. Next is 3, and there is a small subset of voting theorists who think about non-determinism. Assumption 4 is not very important, and while it is usually made, it also usually does not matter. Assumption 1 is almost never violated, or at least when it is, the field is called something other than voting theory. A utility function on a set S is just a function from S[0,1]. I will write Δ(S) for the set of distributions on the set S. (I'm not going to worry about the sigma algebras; usually it will be obvious/not matter.) If f is a voting system, C is a set of candidates, and V∈Δ(C[0,1]) is a distribution on utility functions on C, I will write fC(V) for the output of the voting system on V, so fC(V)∈Δ(C). Some Important Criteria To get more comfortable with this formalism, we will translate three important voting criteria. Condorcet Criterion If a candidate would defeat all others in one-on-one elections, that candidate should win. Translated to our formalism, f satisfies the Condorcet criterion if whenever there exists a c∈C such that for all d∈C, we have Pv∼V(v(c)>v(d))>12, we havefC(V)(c)=1. We can think of Pv∼V(v(c)>v(d)) as saying when we randomly choose a voter v, what is the probability that v prefers c to d. If this probability is greater than 12, then a majority of voters prefer c to d. Consistency Criterion If two disjoint el...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: My resentful story of becoming a medical miracle, published by Elizabeth on October 16, 2022 on LessWrong. You know those health books with “miracle cure” in the subtitle? The ones that always start with a preface about a particular patient who was completely hopeless until they tried the supplement/meditation technique/healing crystal that the book is based on? These people always start broken and miserable, unable to work or enjoy life, perhaps even suicidal from the sheer hopelessness of getting their body to stop betraying them. They’ve spent decades trying everything and nothing has worked until their friend makes them see the book’s author, who prescribes the same thing they always prescribe, and the patient immediately stands up and starts dancing because their problem is entirely fixed (more conservative books will say it took two sessions). You know how those are completely unbelievable, because anything that worked that well would go mainstream, so basically the book is starting you off with a shit test to make sure you don’t challenge its bullshit later? Well 5 months ago I became one of those miraculous stories, except worse, because my doctor didn’t even do it on purpose. This finalized some already fermenting changes in how I view medical interventions and research. Namely: sometimes knowledge doesn’t work and then you have to optimize for luck. I assure you I’m at least as unhappy about this as you are. Preface to the Preface I’ve had nonspecific digestive issues since before I have memories. In pre-school my family joked that I would die as a caveman because there were so few things I would eat, and they were mostly grains. This caused a bunch of subclinical malnutrition issues that took a lot of time to manage and never got completely better. And while I couldn’t articulate this until it went away, food felt gross all the time It’s hard to convey just how bad this was for me, because it feels like it undermines everything I did to work around it. I’ve always been functional but decidedly less healthy than my friends. I got sick more often and it hit me harder. I was slower to heal from injuries and scrapes and that limited my interest in the more athletic sort of hobbies. I couldn’t work the same hours, and working hours traded off really sharply against energetic hobbies. I had to spend a lot of time managing food where other people can just show up and eat, which was a constant source of social stress. My genetics say I was destined to have anxiety issues, but the low level malnutrition and justified feelings of food insecurity despite apparent abundance did not help anything. Eventually in my late 20s. I saw a nutrition-focused psychiatrist who listened to my observations (I could only eat protein with soda), immediately formed a hypothesis (I produced insufficient stomach acid), asked questions to rule it out (which I no longer remember), suggested a test (take stomach acid pills and see if they gave me heartburn), and when it came back positive (no heartburn) suggested a course of action (keep taking stomach acid pills) that showed immediate benefits in practice (indigestion removed, but only when I took the pills). My protein and produce intake increased enormously, and I felt overall much better. This is exactly how I want medicine to work. I gathered good data and took it to an expert who immediately formed a model, definitively tested it, and prescribed a course of action that made mechanistic sense. If you forget that it took almost 30 years and I took those exact same symptoms to other doctors beforehand, it’s a stunning success. But it was not a total success. My protein intake maxed out at 50 grams/day, and that was if I made consuming protein a hobby and nothing went wrong. I was doing much better than I had been, but my nutrient tests ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: I learn better when I frame learning as Vengeance for losses incurred through ignorance, and you might too, published by chaosmage on October 15, 2022 on LessWrong. THE OBSERVATION My main point is in the title: I have found that when I consciously learn "with a vengeance", aiming to avenge whatever I lost because I have not learnt the thing earlier, I learn markedly better. I feel more motivated to learn, and recall seems clearly better. (I have not formally compared recall before and after, this is subjective judgement but I'm very confident.) I have experimented with this for about 2 weeks. So far the effect does not seem to diminish; if there is any change at all, it is maybe getting slightly stronger. Subjectively, this changes my failures from something to mourn or to be embarrassed about, into "something to be avenged" which is a driving motivation. It feels like I'm owning my mistakes more readily. I acknowledge them more readily and am more comfortable thinking and talking about them. And the learning from them feels something like gleeful, justified or gloating - definitely more powerful and satisfying. I would like some of you to try this, and to report back whether this works for you too, because I have never heard anyone else talk about consciously trying to do this, so it might be new, and it helps me a lot so if it helps others too, it seems potentially extremely useful. THE NARRATIVE I was having a great argument about God with a very good friend who is a sophisticated theologian, and who pointed me to a podcast with the German theologian Siegfried Zimmer. Zimmer was talking about theodicy, the classic problem of theism where an almighty, all-knowing, all-good God seems impossible to square with the pointless suffering we observe. He did not have a satisfactory answer, he did manage to admit that, he did not manage to draw the obvious conclusion, so nothing new there. But in his argument he gave a really interesting reason to discard all the usual theistic answers. He said an answer to suffering is only good if you can say it to someone who is intensely and pointlessly suffering, to his or her face, and find it helps. I was very impressed with this idea and concluded that as an atheist, I did not have an answer that would fulfil this criterion. So I thought about it. What do I say to someone dying in a concentration camp, from the other side of the fence, utterly unable to save them? To a kid dying of cancer? To the grieving parents of victims of a school shooting? Naturally, as you do, I thought back to the Star Trek parody "Galaxy Quest". The single best scene in the movie is this: This scene stands out from the rest of the movie because of its utter sincerity. That is the point: the same character has been saying the same line insincerely for many years and is completely sick of it, but he recognises that in this particular situation it is the best thing he could possibly say: "You shall be avenged." In the concentration camp scenario, if I imagine myself on either side of the fence, I really think that would help. "We can't save you, but we will avenge you." Yeah. That rings good. And I find this easily generalizes into situations where there isn't a human perpetrator to be punished. The project of eradicating Malaria just feels more viscerally awesome when I frame it as the spiteful, relishing, merciless, victorious extermination of the terrible monster that has, by some estimates, killed around 10% of all humans who have ever lived. And it works for small things just as well. I paid too much for an item? I avenge the lost money by buying more carefully next time - and I find I actually do remember to do it next time. I lost time because of a scheduling mistake? I avenge it by scheduling better - and my scheduling improves faster than it used to....
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: A Few Terrifying Facts About The Russo-Ukrainian War, published by DonyChristie on September 30, 2022 on LessWrong. Epistemic status: trying to summarize the news and predict, post is under revision, too lazy to citation everything I wanted to collect a few observations I've made, as best I understand them. This PBS article does a good job of explaining much of it. Vladimir Putin has announced the annexation of four Ukrainian territories. This makes them Russian territory from Russia's perspective. “People living in Donetsk, Luhansk, Zaporizhzhia and Kherson are becoming our citizens forever” - Putin The West does not acknowledge this annexation, describing it as illegal. Ukraine does not acknowledge this annexation and says it plans to take the territories back. "By attempting to annex Ukraine's Donetsk, Luhansk, Zaporizhzhia and Kherson regions, (Russian President Vladimir) Putin tries to grab territories he doesn't even physically control on the ground. Nothing changes for Ukraine: we continue liberating our land and our people, restoring our territorial integrity," Ukraine's Foreign Minister Dmytro Kuleba said on social media. Russian military doctrine allows the usage of nuclear weapons to defend Russian territory. Putin has a track record of escalating apparently (this needs more data) and Russia seems to be planning for escalation until the war is won. "All of our sources in the elite — who all spoke on the condition of anonymity — said the military conflict will only escalate in the coming months." Putin has clearly stated that they will defend this territory, including with tactical nukes if need be. He said they would use "any means available" to defend it He has mentioned usage of nukes some number of times (a nice-to-have: a list of all the times he has said this) Medyedev has stated the West would not retaliate if nuclear weapons are used. "Under Russia’s amended constitution, no Kremlin leader can cede territories once they are annexed." - someone on Twitter Putin has stated he is not bluffing. The U.S. Secretary of State says it is "loose talk". Putin has called for a ceasefire. Ukraine and U.S. does not want to do this. The U.S. has said there will be "catastrophic consequences" if nuclear weapons are used. They are keeping the consequences vague for strategic flexibility. Concerning escalatory developments that aren't directly related to nuclear brinksmanship: The Nordstream natural gas pipes were blown up. We don't know who did it. (This section needs work) Russia could have done it Burning the bridges strategy? U.S. could have done it U.S. airships were nearby days before. Ukraine Some other country or group, hypothetically Russia has conscripted 300,000 men. There is some amount of resistance. Tens of thousands of people are leaving. There are some protests. Ukraine has "accelerated" its application to join NATO. Consensus from all 30 NATO countries is required, though. France and Germany have expressed reluctance in the past. "Experts warned that Ukraine’s NATO membership at the moment seems elusive at best. The process could take at least several months, and even years." Sweden and Finland have been approved by 28 out of 30 countries. Turkey and Hungary will likely hold out for a while. Any country in NATO that is attacked by Russia triggers the whole NATO alliance to attack Russia. Biden has affirmed this, naturally. Conclusion: Ukraine will keep attacking the annexed territories in order to take them back until Russia uses a tactical nuke out of desperation, and the U.S. will respond with "catastrophic consequences". This is obviously uncertain! But the chain of logic forms a coherent enough inside view for me to put a lot of probability on that, and I may start taking bets. I would be curious to see different inside views about what will happ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Prioritizing Parental Sleep, published by jefftk on September 30, 2022 on LessWrong. Overall I've really enjoyed being a parent, but poor sleep has been the hardest part. At times we've both been so tired that we weren't able to think clearly about how to fix the problem, which can be very tricky. We've figured this out more over time, however, and sleep has improved with each successive kid (n=3). Here's the main things that have worked for us, all in one place. Note that kids and parents vary a lot: our three kids are different from each other, and your kids likely even more so. I'm hoping that many things on this list will be useful to many people, but it wouldn't be surprising for several of them to be a poor fit for any individual family. I've ordered them roughly down from the ones that I think are most likely to work for anyone who tries them. Perhaps unsurprisingly this is also roughly in the order of youngest to oldest; my impression is kids diverge more over time as their personalities come out. Sleeping in multiple rooms. If the baby waking up means only interrupting one parents sleep that's about half as much sleep deprivation. Blackout curtains. Small children typically wake up with the sun, which means you're up to. If you can keep their room solidly dark in the morning by blocking out the sun you can shift their schedule to whenever is most convenient for you, typically a later waking. Using a Snoo automated bassinet. It gently rocks the baby and shushes them back to sleep when they wake. We used this with our youngest for the first six months. It worked very well, and automated almost all of the helping babies fall back asleep that I needed to do with the older two. We weaned her off it much more gradually than they recommended, with first weaning mode (stops rocking after the baby is asleep) and then running it without rocking. Sleep training. At some point they no longer need to feed as often, but either don't know how to fall back asleep on their own at the end of a sleep cycle or would enjoy getting a cuddle before going back to sleep. I wrote about our experience with our oldest, and how her naps got immediately better. With our youngest, at 15m, we've recently dropped the last night feed, and are in the process of teaching her that she needs to go back to sleep on her own the whole night. Letting kids know when they can get up. When kids are a bit older, they'll often wake up near the end of the night and not be sure whether it's morning. An "OK to wake" light (or a janky incorrect clock) can help them figure out whether to go back to sleep or get up. Child proofing bedrooms. When they're old enough to no longer be in a crib, it's great if their bedroom is a place that they can play unsupervised in the morning without waking you up. Teaching them not to wake you. When they're old enough to play in the rest of the house unsupervised but still sometimes play too loudly or fight, teaching them that in the morning this isn't ok. If they wake up before us they have to resolve conflicts independently and keep their volume down; if they don't, both of go back to not being able to leave their rooms until we're up. I think the last time we needed to enforce this they were maybe 4y and 6y. Other things that can be helpful: If things are really not going well, and you're both way too tired to figure things out, pick one of you to get caught up on sleep while the other covers. The one that's caught up can then figure out what needs changing, and make it up the unevenness later. Some people have had good success with cosleeping, but while we started with the two older kids sleeping in an annex cosleeper we didn't do this with the youngest; the Snoo worked better for us. We will sometimes cosleep when traveling or in other unusual situations. Tracking wake wi...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Announcing $5,000 bounty for ending malaria, published by lc on September 24, 2022 on LessWrong. It occurred to me recently that if someone eradicated malaria, say by exterminating Anopheles gambiae via gene drives, they would not be paid. That sounds like a potentially embarrassing mistake, so I'm hereby pledging $5,000 to anyone that manages to accomplish such a thing. This way if someone from Dath Ilan materializes into our world and asks us judgmentally whether or not Earth would even pay people for eradicating malaria, we can all respond "Our society definitely has an explicit, preregistered financial reward greater than $0 in place for anybody that does that" and sidestep some awkwardness. Below I have written out the details for this new bounty in FAQ form. Malaria Bounty FAQ What specifically will lc wire me $5,000 for doing? I will give you $5,000 if you permanently reduce global annual mortality from malaria by 95% or more. For example: if you exterminate Anopheles gambiae or permanently render those mosquitoes unable to transmit malaria to humans, and this reduces incidence and therefore deaths from malaria to near zero, I will wire you five thousand dollars. Eligibility for the reward is independent to how this is actually accomplished, with a few restrictions outlined below. Do I have to apply somewhere? Nope. You may decide to contact me and provide relevant evidence so I am aware that malaria has been solved, but I expect I'll be proactively contacting anyone that actually does this after I hear about them doing it through the news. How will lc assign credit for solving malaria? Does developing the deployed technique count? The money will be given to the person or group that directly causes the reduction in deaths, e.g. by deploying the gene drive. If a pioneering group of researchers wants to get my money, they will have to follow through by successfully using their clever research to prevent people from dying. I would like to handsomely reward both groups, but there are already many existing charities you can extract money from in advance, if you're working on tools to end malaria and want funding for that. So far, none of the (awesome) people developing those tools have opted to use them, partly because they insist on getting permission from local African governments, which are delaying them for political reasons. I reserve hope that I will some day have reason to send these researchers money, after a few more million people have died, but in the meantime there's not as much marginal benefit. What if someone mostly eradicates malaria, but in a way that causes a legible negative side effect? If a negative consequence of you or your organization's solution is brought to my attention, I shall construct and consult an expert technology ethics review board, staffed by myself. This ethical review board will formally determine whether or not the bad thing is veritably and legibly bad enough to outweigh saving hundreds of thousands of lives per year. If the review panel's unanimous finding is that this is indeed the case, then the party responsible will not receive the $5,000. Examples of some malaria reduction techniques which could cause someone to be disqualified by the expert panel include: Starting a nuclear war. Releasing an AGI that turns most of the earth into paperclips. Using dark magicks to lower Africa into the sea. Examples of secondary effects that are explicitly noted not to disqualify award recipients include: Somehow-negative press coverage about particular groups the review panel likes, such as rationalists, effective altruists, or LessWrong users. Accidental eradication of a few non-mosquito species, in a manner that is not historically notable against the larger backdrop of the Holocene extinction event. Sternly worded condemnation by a ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Ukraine Post #12, published by Zvi on September 22, 2022 on LessWrong. After doing frequent Ukraine posts early on, I decided that the war was no long a good match for this blog and my skill set. The basic situation was clear and was moving slowly. In the last few weeks things have started moving more rapidly on multiple fronts. Ukraine has made substantial progress. In response, Russia has begun substantial mobilization. Thus, it seems like it is worth writing another of these posts. I am not sure if events will justify any additional ones. Putin’s Speech In situations like this, one must watch the speech. I was later pointed to a transcript. Here are my notes as I watched. I cannot verify the translation but have no reason to doubt it that I can see. Putin initially frames this, even now, as about Donetsk and Lugansk only. However he then quickly names other ‘liberated’ regions and talks about the ‘security and territorial integrity’ of the regime – the ‘desire and will of our compatriots to choose their future independently’ (1:10) He claims West wants to split up the Russian Federation. Yes, well. (1:51) LNR/DNR units now considered on par with Russian units. (4:57) Clear intent to keep all territory Russia physically controls. (8:27) They will do this via ‘referendums’ in all these regions. Mobilization only of the reserves, promise of training. (10:00) Mobilization begins today, right now. (10:43) Defense industry tasked with ensuring there is equipment, promised support as needed. (11:17) Delivery of Western weapons that could ‘strike Crimea and Russia.’ Claims that West is pushing Ukraine to invade Russia. (12:27) Accused West of ‘nuclear blackmail’ including weapons threats. (13:06) “In the event of a threat to the territorial integrity of the country, and to defend Russia and our people, we will certainly make use of all weapons systems available to us. This is not a bluff.” Emphasis on “By all the systems available to us.” Speech accused Ukraine and the West of various atrocities and other things as justification for all this. Speech does not mention how many troops are being called up. Many sources say 300,000. The speech neither limits that in any way, nor commits to doing anything that substantial. The references to industrial production being the responsibility of industry, aside from promises of financial support, seems like throwing industry under the bus. It is a useful thing to not interfere too much with industry. It is a different thing when you say they are ‘responsible’ for such matters while they lack the necessary inputs. Mike Ryan has an interesting close reading. He concludes that a lot of the speech is about preparing to blame others for failure. These are not measures one takes without accepting the possibility of losing. Close reading of Putin’s words is most important when it comes to nuclear weapons. Our Words Are Backed By Nuclear Weapons Russia threatening the West with nukes is nothing new at this point. Neither is Russian television continuously pushing anti-West, anti-American lines and calling for starting World War III. Putin explicitly threatening to use nukes to defend Ukrainian territory he has taken, however, does seem new. What exactly is he intending here? Is it a bluff? Here is one analysis. It is never a good sign when you are saying “this is not a bluff.” The first thing to notice is that the nightmare scenario, where Russia creates its sham Republics, annexes the Republics (potentially including territory it does not control) and then is willing to use nukes to defend them? That is a possible interpretation, which Putin is carefully ensuring one could hear if one wishes to do so. It is not what he is actually saying, or what he believes he will need to claim that he said in order to maintain credibility. Which is impor...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Redaction Machine, published by Ben on September 20, 2022 on LessWrong. On the 3rd of October 2351 a machine flared to life. Huge energies coursed into it via cables, only to leave moments later as heat dumped unwanted into its radiators. With an enormous puff the machine unleashed sixty years of human metabolic entropy into superheated steam. In the heart of the machine was Jane, a person of the early 21st century. From her perspective there was no transition. One moment she had been in the year 2021, sat beneath a tree in a park. Reading a detective novel. Then the book was gone, and the tree. Also the park. Even the year. She found herself laid in a bathtub, immersed in sickly fatty fluids. She was naked and cold. The first question Jane had for the operators and technicians who greeted her with a warm towel was "where am I?''. They brought back the dead a hundred times a day, and knew that it was best to answer when, not where. Jane's second question was "how did I come to be in the year 2351?''. The answer begins in the year 2021. Jane sat beneath a tree. She read her book. She walked home. She attended university. She played her guitar. Went to pubs and festivals with her friends. She studied philosophy as her course, but there wasn't much work going in analysing Plato so she got a job in marketing. She married, had children. Lived a normal enough life. She followed the scientific news no more than most. She did watch the much-hyped documentary about the possibility of restoring objects to a previous state. She found it a disappointment. So they could get a single molecule to unreact? Maybe some chemist somewhere cared. But technology's march was unceasing. The egg was unboiled only two decades later. The redaction machines, as they came to be known, restored an object to a previous state. Set the machine for 10 minutes and insert your alarm clock. Ping! The machine takes no time at all to operate. Within find alarm clock, time as it read 10 minutes ago. Smash said alarm clock with hammer. Insert the pieces into the machine. Activate. Find alarm clock restored, exactly, to state 10 minutes prior. No longer smashed. It is not now a repaired alarm clock. If something has been repaired then one supposes that it was once broken. A repaired thing still possesses the history of being broken and then fixed. But that clock, as taken from the machine, is the clock from earlier, before it ever suffered damage. Its history taken from it. The breakage unhappened. The first resurrection was only a half-decade after the egg. By the same token as the clock this was not truly a resurrection. It was a restoration, an undoing of time's passage on a single person. "Resurrection'', despite its flaws as a term, became the standard manner of speaking about these things. Accidental death dropped to almost zero in the developed world. Or at least, unredacted, accidental death did. Truly the number of deaths increased modestly, as some people became more brazen. But nearly all were redacted. Ambulances would carry the corpses made by a traffic accident to a redaction clinic. Death 1 hour ago. Re-set patient. It would be a shock to transition from driving to a strange hospital, with no memory of the accident (the neurochemistry encoding any such memory was just as unhappened as the tearing of ligaments and the crush of bones). But it was surely better than being dead. During redaction every subatomic particle in an object would run its past trajectory backwards as best it could. The theory of relativity demanded that centre-of-mass motion of the object as a whole be an exception to this rule. Reality complied. For humans, respiration acted in reverse. Blood cells passed backwards through arteries, carrying newly minted oxygen from the muscles out to the lungs to exhale. Electrical i...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: You Are Not Measuring What You Think You Are Measuring, published by johnswentworth on September 20, 2022 on LessWrong. Eight years ago, I worked as a data scientist at a startup, and we wanted to optimize our sign-up flow. We A/B tested lots of different changes, and occasionally found something which would boost (or reduce) click-through rates by 10% or so. Then one week I was puzzling over a discrepancy in the variance of our daily signups. Eventually I scraped some data from the log files, and found that during traffic spikes, our server latency shot up to multiple seconds. The effect on signups during these spikes was massive: even just 300 ms was enough that click-through dropped by 30%, and when latency went up to seconds the click-through rates dropped by over 80%. And this happened multiple times per day. Latency was far and away the most important factor which determined our click-through rates. Going back through some of our earlier experiments, it was clear in hindsight that some of our biggest effect-sizes actually came from changing latency - for instance, if we changed the order of two screens, then there’d be an extra screen before the user hit the one with high latency, so the latency would be better hidden. Our original interpretations of those experiments - e.g. that the user cared more about the content of one screen than another - were totally wrong. It was also clear in hindsight that our statistics on all the earlier experiments were bunk - we’d assumed that every user’s click-through was statistically independent, when in fact they were highly correlated, so many of the results which we thought were significant were in fact basically noise. Main point of this example: we were not measuring what we thought we were measuring. We thought we were testing hypotheses about what information the user cared about, or what order things needed to be presented in, or whether users would be more likely to click on a bigger and shinier button. But in fact, we were mostly measuring latency. When I look back on experiments I’ve run over the years, in hindsight the very large majority of cases are like the server latency example. The large majority of the time, experiments did not measure what I thought they were measuring. I’ll call this the First Law of Experiment Design: you are not measuring what you think you are measuring. Against One-Bit Experiments A one-bit experiment is an experiment designed to answer a yes/no question. It’s the prototypical case from high school statistics: which of two mouse diets results in lower bodyweight? Which of two button designs on a website results in higher click-through rates? Does a new vaccine design protect against COVID better than an old design (or better than no vaccine at all)? Can Muriel Bristol tell whether milk or tea was added first to her teacup? Will a neural net trained to navigate to a coin at the end of a level still go to the coin if it’s no longer at the end of a level? Can a rat navigate a maze just by smell? There’s an obvious criticism of such experiments: at best, they yield one bit of information. (Of course the experimenter probably observes a lot more than one bit of information over the course of the experiment, but usually people are trained to ignore most of that useful information and just report a p-value on the original yes/no question.) The First Law of Experiment Design implies that the situation is much worse: in the large majority of cases, a one-bit experiment yields approximately zero information about the thing the experimenter intended to measure. It inevitably turns out that mouse bodyweight, or Muriel Bristol’s tea-tasting, or a neural net’s coinrun performance, in fact routes through something entirely different from what we expected. Corollary To The First Law: If You Are Defin...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How my team at Lightcone sometimes gets stuff done, published by jacobjacob on September 19, 2022 on LessWrong. Disclaimer: I originally wrote this as a private doc for the Lightcone team. I then showed it to John and he said he would pay me to post it here. That sounded awfully compelling. However, I wanted to note that I’m an early founder who haven't built anything truly great yet. I’m writing this doc because as Lightcone is growing, I have to take a stance on these questions. I need to design our org to handle more people. Still, I haven’t seen the results long-term, and who knows if this is good advice. Don’t overinterpret this. Suppose you went up on stage in front of a company you founded, that now had grown to 100, or 1000, 10 000+ people. You were going to give a talk about your company values. You can say things like “We care about moving fast, taking responsibility, and being creative” -- but I expect these words would mostly fall flat. At the end of the day, the path the water takes down the hill is determined by the shape of the territory, not the sound the water makes as it swooshes by. To manage that many people, it seems to me you need clear, concrete instructions. What are those? What are things you could write down on a piece of paper and pass along your chain of command, such that if at the end people go ahead and just implement them, without asking what you meant, they would still preserve some chunk of what makes your org work? Here’s my current best guess at how I would do this for Lightcone Infrastructure, the organisation where I spend the majority of my waking hours. I wrote it by thinking about how the team I'm on has actually operated during periods of high output, and then trying to turn that into a set of rules. Others on the team might disagree about which rules matter and where the magic sauce is, but I think this is at least empirically descriptive of how my team spends much of our time. 0. Blockers are death. Above all else, your job is to unblock anything that prevents you from moving as fast as you can toward your top priority. The organisation has one CEO who’s the final decision-maker on all decisions, unless they explicitly delegate a decision. Consensus is slow and blocking. Having a tie-breaker means you can move faster. Start each Monday with an all-hands meeting where the CEO sets or clarifies the top priority of the organisation. After the Monday meeting, there’s a block of time on everyone’s calendar during which no one is allowed to schedule any meetings in advance. This is the “top priority block”, where everyone’s sole goal is to work on whatever the top priority is, as identified in the all-hands meeting. The point of not scheduling meetings is so that no one ends up blocked and waiting for someone else to come out of a call or meeting (that’s not itself about the top priority). Any person on the team who could unblock someone else’s pursuit of the top priority, is available to do so during this block. Have a single day, e.g. Tuesday, that’s the “meeting day”, where people are expected to schedule any miscellaneous, external meetings (e.g. giving someone career advice, or grabbing coffee with a contact). The reason I find this important is that, if you don’t have it, people end up scheduling their meetings at random, uncoordinated times during the week. And this means that there ends up being very few slots where the people who might unblock each other are both free at the same time. No remote work. Everyone is in the office. If you need to be unblocked by someone, the fastest way is to just go to their desk and ask them in person. People on the same team work in the same room. Sudden questions, comments, or info sharing out loud is encouraged. If people want to focus deeply for a while, they can put on headphones. For...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Understanding Conjecture: Notes from Connor Leahy interview, published by Akash on September 15, 2022 on LessWrong. I recently listened to Michaël Trazzi interview Connor Leahy (co-founder & CEO of Conjecture)'s on a podcast called The Inside View (Youtube video here; full video & transcript here). The interview helped me better understand Connor’s worldview and Conjecture’s theory of change. I’m sharing my notes below. The “highlights” section includes the information I found most interesting/useful. The "full notes" section includes all of my notes. Disclaimer #1: I didn’t take notes on the entire podcast. I selectively emphasized the stuff I found most interesting. Note also that these notes were mostly for my understanding, and I did not set out to perfectly or precisely capture Connor’s views. Disclaimer #2: I’m always summarizing Connor (even when I write with “I” or “we”— the “I” refers to Connor). I do not necessarily endorse or agree with any of these views. Highlights Timelines 20-30% in the next 5 years. 50% by 2030. 99% by 2100. 1% we already have it (but don’t know this yet). Higher uncertainty than Eliezer but generally buys the same arguments. Is mostly like “Eliezer’s arguments seem right but how can anyone be so confident about things?” Thoughts on MIRI Dialogues & Eliezer’s style An antimeme is something that by its very nature resists being known. Most antimemes are just boring—things you forget about. If you tell someone an antimeme, it bounces off them. So they need to be communicated in a special way. Moral intuitions. Truths about yourself. A psychologist doesn’t just tell you “yo, you’re fucked up bro.” That doesn’t work. A lot of Eliezer’s value as a thinker is that he notices & comprehends antimemes. And he figures out how to communicate them. What happened in the MIRI dialogues is that Eliezer was telling Paul “hey, I’m trying to communicate an antimeme to you, but I’m failing because it’s really really hard.” Thoughts on Death with Dignity & optimizing for “dignity points” rather than utility The Death with Dignity post is a perfect example of an antimeme. A great way to convey antimemes is through jokes and things outside the Overton Window. The antimeme is that utilitarianism is hard, and no, it’s not actually a good idea to advocate for really stupid “pivotal acts” that sound ridiculous. Consequentialism is really hard. I have to reason about all of my possible choices and all of their possible consequences. If you have an infinitely big brain, this works. If not, it doesn’t. It’s too computationally hard to be a perfect consequentialist. And being an imperfect consequentialist is really really bad. If you do one step of reasoning, you might be like “yeaaa let’s get rid of GPUs!” But you don’t realize how that would be super bad for the world, would make cooperation extremely difficult, would make everything become super secretive, etc. The antimeme is that most people shouldn’t be thinking like consequentialists. Instead of thinking about how to maximize utility, they should be thinking about how to maximize dignity. This is easier. This is computationally tractable. This heuristic will make you do better. I see so many people come into this arena with the anime protagonist “I’m going to save the world” complex, and then they burnout after 3 months and go do DMT. I know two humans who can maybe reason better under the consequentialist frame. But for everyone else, if you’re going to do 5 years of soul-crushing difficult research without much support from the outside world, you should think under the dignity frame. Thoughts on the importance of playing with large models One mistake I see people make is that they underestimate the importance of getting actual hands-on experience with the thing you are studying. I think it’s important to ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: ACT-1: Transformer for Actions, published by Daniel Kokotajlo on September 14, 2022 on LessWrong. ACT-1 can take a high-level user request and execute it. The user simply types a command into the text box and ACT-1 does the rest. In this example, this requires repeatedly taking actions and observations over a long time horizon to fulfill a single goal. While we’re excited that these systems can transform what people can do on a computer, we clearly see that they have the potential to cause harm if misused or misaligned with user preferences. Our goal is to build a company with large-scale human feedback at the center — models will be evaluated on how well they satisfy user preferences, and we will iteratively evaluate how well this is working as our product becomes more sophisticated and load-bearing. Daniel's commentary: To be clear, this is very much just a cool demo, not much better than WebGPT as far as I can tell, and not surprising at all that this level of capability is possible. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Dan Luu on Futurist Predictions, published by RobertM on September 14, 2022 on LessWrong. Epistemic status: perspective derived from following Dan Luu's output for the last 5 years or so. Trying to vaguely gesture at a few things at once. Please ask questions if you find something confusing. Dan Luu has written a interesting post analysing the track record of futurists' predictions. The motivation: I've been reading a lot of predictions from people who are looking to understand what problems humanity will face 10-50 years out (and sometimes longer) in order to work in areas that will be instrumental for the future and wondering how accurate these predictions of the future are. The timeframe of predictions that are so far out means that only a tiny fraction of people making those kinds of predictions today have a track record so, if we want to evaluate which predictions are plausible, we need to look at something other than track record. The idea behind the approach of this post was to look at predictions from an independently chosen set of predictors (Wikipedia's list of well-known futurists1) whose predictions are old enough to evaluate in order to understand which prediction techniques worked and which ones didn't work, allowing us to then (mostly in a future post) evaluate the plausibility of predictions that use similar methodologies. I'm primarily going to address the appendix, particularly the section on Holden Karnofsky's analysis on the same subject, but the article is interesting reading and I'd recommend going through the whole thing. (I think Dan is evaluating forecasting track records pretty differently from how I would, and I haven't actually dug into any of the other analysis. On priors I'd expect it to be similar to his analysis of Holden's work.) Karnofsky's evaluation of Kurzweil being "fine" to "mediocre" relies on these two analyses done on LessWrong and then uses a very generous interpretation of the results to conclude that Kurzweil's predictions are fine. Those two posts rate predictions as true, weakly true, cannot decide, weakly false, or false. Karnofsky then compares the number of true + weakly true to false + weakly false, which is one level of rounding up to get an optimistic result; another way to look at it is that any level other than "true" is false when read as written. This issue is magnified if you actually look at the data and methodology used in the LW analyses. The specific claim I have an issue with here is "another way to look at it is that any level other than "true" is false when read as written". Depending on how you want to evaluate it it, it's either technically true but irrelevant, or not even wrong. In the second post, the author, Stuart Armstrong indirectly noted that there were actually no predictions that were, by strong consensus, very true when he noted that the "most true" prediction had a mean score of 1.3 (1 = true, 2 = weakly true ... , 5 = false) and the second highest rated prediction had a mean score of 1.4. Although Armstrong doesn't note this in the post, if you look at the data, you'll see that the third "most true" prediction had a mean score of 1.45 and the fourth had a mean score of 1.6, i.e., if you round to the nearest prediction score, only 3 out of 105 predictions score "true" and 32 are >= 4.5 and score "false". Karnofsky reads Armstrong's as scoring 12% of predictions true, but the post effectively makes no comment on what fraction of predictions were scored true and the 12% came from summing up the total number of each rating given. I'm not going to say that taking the mean of each question is the only way one could aggregate the numbers (taking the median or modal values could also be argued for, as well as some more sophisticated scoring function, an extremizing function, etc.), but summing up ...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: [Linkpost] A survey on over 300 works about interpretability in deep networks, published by scasper on September 12, 2022 on LessWrong. Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks Tilman Räuker traeuker@gmail.com Anson Ho anson@epochai.org Stephen Casper scasper@mit.edu Dylan Hadfield-Menell TL;DR: We wrote a survey paper on interpretability tools for deep networks. It was written for the general AI community but with AI safety as the key focus. We survey over 300 works and offer 15 discussion points for guiding future work. Here is a link to a Twitter thread about the paper. Lately, there has been a growing interest in interpreting AI systems and a growing consensus that it will be key for building safer AI. There have been rapid recent developments in interpretability work, and the AI safety community will benefit from a better systemization of knowledge for it. There are also several epistemic and paradigmatic issues with much interpretability work today. In response to these challenges, we wrote a survey paper covering over 300 works and featuring 15 somewhat “hot takes” to guide future work. Specifically, this survey focuses on “inner” interpretability methods that help explain internal parts of a network (i.e. not inputs, outputs, or the network as a whole). We do this because inner methods are popular and have some unique applications – not because we think that they are more valuable than other ones. The survey introduces a taxonomy of inner interpretability tools that organizes them by which part of the network’s computational graph they aim to explain: weights (S2), neurons (S3), subnetworks (S4), and latent representations (S5). Then we provide a discussion (S6) and propose directions for future work (S7). Finally, here are a select few points that we would like to specifically highlight here. Interpretability does not just mean circuits. In the survey sections of the paper (S2-S5), there are 21 subsections, and only one is about circuits. The circuits paradigm has received disproportionate attention in the AI safety community, partly due to Distill’s influential interpretability research in the past few years. But given how many other useful approaches there are, it would be myopic to focus too much on them. Interpretability research has close connections to work in adversarial robustness, continual learning, modularity, network compression, and studying the human visual system. For example, adversarially trained networks tend to be more interpretable, and more interpretable networks tend to be more adversarially robust. Interpretability tools generate hypotheses, not conclusions. Simply analyzing the outputs of an interpretability technique and pontificating about what they mean is a problem with much interpretability work – including AI safety work. There are many examples of when this type of approach fails to produce faithful explanations. Interpretability tools should be more rigorously evaluated. There are currently no broadly established ways to do this. Benchmarks for evaluating interpretability tools can and should be popularized. The ultimate goal of interpretability work should be tools that give us insights that are valid and useful. Ideally, interpretations should be used to make and validate useful predictions that engineers can use. So benchmarks should be created which measure how well interpretability tools can help us understand systems well enough to do engineering-relevant things with them. Examples of this could be using interpretability tools for reverse engineering a system, manually finetuning a model to introduce a predictable change in behavior, or designing a novel adversary. The Automated Auditing agenda may offer a useful paradigm for this – judging techniques by their ability t...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Survey of NLP Researchers: NLP is contributing to AGI progress; major catastrophe plausible, published by Sam Bowman on August 31, 2022 on LessWrong. I was part of a group that ran a PhilPapers-style survey and metasurvey targeting NLP researchers who publish at venues like ACL. Results are here (Tweet-thread version). It didn't target AGI timelines, but had some other questions that could be of interest to people here: NLP is on a path to AGI: 58% agreed that Understanding the potential development of artificial general intelligence (AGI) and the benefits/risks associated with it should be a significant priority for NLP researchers. Related: 57% agreed that Recent developments in large-scale ML modeling (such as in language modeling and reinforcement learning) are significant steps toward the development of AGI. AGI could be revolutionary: 73% agreed that In this century, labor automation caused by advances in AI/ML could plausibly lead to economic restructuring and societal changes on at least the scale of the Industrial Revolution. AGI could be catastrophic: 36% agreed that It is plausible that decisions made by AI or machine learning systems could cause a catastrophe this century that is at least as bad as an all-out nuclear war. 46% of women and 53% of URM respondents agreed. The comments suggested that people took a pretty wide range of interpretations to this, including things like OOD robustness failures leading to weapons launches. Few scaling maximalists: 17% agreed that Given resources (i.e., compute and data) that could come to exist this century, scaled-up implementations of established existing techniques will be sufficient to practically solve any important real-world problem or application in NLP. The metasurvey responses predicted that 47% would agree to this, so there are fewer scaling maximalists than people expected there to be. Optimism about ideas from cognitive science: 61% agreed that It is likely that at least one of the five most-cited systems in 2030 will take clear inspiration from specific, non-trivial results from the last 50 years of research into linguistics or cognitive science. This strikes me as very optimistic, since it's pretty clearly false about the most cited systems today. Optimism about the field: 87% agreed that On net, NLP research continuing into the future will have a positive impact on the world. 32% of respondents who agreed that NLP will have a positive future impact on society also agreed that there is a plausible risk of global catastrophe. Most NLP research is crap: 67% agreed that A majority of the research being published in NLP is of dubious scientific value. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: (My understanding of) What Everyone in Technical Alignment is Doing and Why, published by Thomas Larsen on August 29, 2022 on LessWrong. Epistemic Status: My best guess Epistemic Effort: ~50 hours of work put into this document Contributions: Thomas wrote ~85% of this, Eli wrote ~15% and helped edit + structure it. Unless specified otherwise, writing in the first person is by Thomas and so are the opinions. Thanks to Miranda Zhang, Caleb Parikh, and Akash Wasil for comments. Thanks to many others for relevant conversations. Introduction Despite a clear need for it, a good source explaining who is doing what and why in technical AI alignment doesn't exist. This is our attempt to produce such a resource. We expect to be inaccurate in some ways, but it seems great to get out there and let Cunningham’s Law do its thing. The main body contains our understanding of what everyone is doing in technical alignment and why, as well as at least one of our opinions on each approach. We include supplements visualizing differences between approaches and Thomas’s big picture view on alignment. The opinions written are Thomas and Eli’s independent impressions, many of which have low resilience. Our all-things-considered views are significantly more uncertain. A summary of our understanding of each approach: Problem FocusCurrent Approach SummaryModel splinteringSolve extrapolation problems. Inaccessible informationELK + LLM power-seeking evaluationLack of good interpretability tools (?)Interpretability + HHH + augmenting alignment research with LLMsBrain-like AGI SafetyUse brains as a model for how AGI will be developed, think about alignment in this contextEngaging the ML community, many technical problems Technical research, Infrastructure, and ML community field-building for safetyOuter alignment, though CHAI is diverseImprove CIRL + many other independent approaches. Suffering risksFoundational game theory researchInner alignmentInterpretability + automating alignment research with LLMsScalable oversight (?)Debate + some other thingsMultipolar failure from lack of coordinationVideo gameDeceptionGet the reasoning of the AGI to happen in natural language, then oversee that reasoningMany including deception, the sharp left turn, corrigibility is anti-naturalMathematical research to resolve fundamental confusion about the nature of goals/agency/optimizationScalable oversightRLHF / Recursive Reward Modeling, then automate alignment researchScalable oversightSupervise process rather than outcomes + augment alignment researchersInner alignment (?)Interpretability + Adversarial Training Being able to robustly point at objects in the worldSelection Theorems based on natural abstractionsInstilling inner values from an outer training loopFind patterns of values given by current RL setups and humans, then create quantitative rules to do thisDeceptionCreate standards and datasets to evaluate model truthfulness Problem FocusCurrent Approach SummaryModel splinteringSolve extrapolation problems. Inaccessible informationELK + LLM power-seeking evaluationLack of good interpretability tools (?)Interpretability + HHH + augmenting alignment research with LLMsBrain-like AGI SafetyUse brains as a model for how AGI will be developed, think about alignment in this contextEngaging the ML community, many technical problems Technical research, Infrastructure, and ML community field-building for safetyOuter alignment, though CHAI is diverseImprove CIRL + many other independent approaches. Suffering risksFoundational game theory researchInner alignmentInterpretability + automating alignment research with LLMsScalable oversight (?)Debate + some other thingsMultipolar failure from lack of coordinationVideo gameDeceptionGet the reasoning of the AGI to happen in natural language, then oversee that reasoningMany in...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Expanding Moral Cinematic Universe, published by Raemon on August 28, 2022 on LessWrong. A subtheme for me this year has been grappling with how to reconcile different facets of my morality. Part of this has to do with reconcile the virtues of "Protect yourself, maintain slack, be aligned with yourself and your community" and "But, maybe the world is metaphorically and/or literally on fire. How do you want to relate to that?" Part of this has to do with "man, the story of the Expanding Circle of Concern may not be as nice as I thought, in ways that fundamentally challenge my conception of what 'my particular morality' even means, or whether it is coherent." (This has gone hand-in-hand with me realizing my conception of 'community' may also not really be coherent, which has been challenging for my identity). Along the way, I've watched some movies and shows that feel like they grapple with a particulate facet of the expanding circle of concern, and they're accumulating into a private (public now) headcanon of the Expanding Moral Cinematic Universe. Narratives that feature characters in a harsh, bloody world who inch their little corner of the universe forward as a place where friendship and cooperation can form. Less a sea of blood and violence and mindless replication. Here are three vignettes, written over the course of the last year (originally on Facebook, later posted on shortform). Part I: The Fox and the Hound Originally posted January 24th on Facebook. I watched Disney's The Fox and The Hound a few weeks several months ago. I cried a bit. While watching the movie, my girlfriend commented "so... they know that foxes are also predators, right?" and, yes. They do. This is not a movie that was supposed to be about predation except it didn't notice all the ramifications about its lesson. This movie just isn't taking a stand about predation. This is a movie about... kinda classic de-facto tribal morality. Where you have your family and your tribe and a few specific neighbors/travelers that you welcomed into your home. Those are your people, and the rest of the world... it's not exactly that they aren't people, but, they aren't in your circle of concern. Maybe you eat them sometimes. That's life. Copper the hound dog's ingroup isn't even very nice to him. His owner, Amos, leaves him out in a crate on a rope. His older dog friend is sort of mean. Amos takes him out on a hunting trip and teaches him how to hunt, conveying his role in life. Copper enthusiastically learns. He's a dog. He's bred to love his owner and be part of the pack no matter what. My dad once commented that this was a movie that... seemed remarkably realistic about what you can expect from animals. Unlike a lot of other disney movies it didn't require suspending disbelief much. A baby fox and hound might totally play together because they haven't figured out yet that they're supposed to be enemies. If that hound then went away for 6 months to learn to hunt, and came back, it might initially be hesitant to hunt down its fox friend, out of a vague/confused memory. But interspecies friendship isn't that strong, and doing-what-your-species/tribe does is often stronger, and yeah later on the hound is like "okay I guess we're hunting this fox now. It's what master wants. I do what master says, that's who I am." And then... ...well, and then the hound gets attacked by a bear. And the fox comes back to save him. And on one hand, I bet 99+% of foxes would not do that. But, having seen a bunch of youtubes of animals doing impressive things, I'm willing to buy that the occasional hero foxes exist, and I'm willing to buy a fox who remembers enough of his interspecies friend to intervene bravely. (weakly held, if someone who knew a lot more about foxes than me was like "nope, this is outside the space of what...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Some conceptual alignment research projects, published by Richard Ngo on August 25, 2022 on LessWrong. Some research outputs I’d love to see, focused on exploring, clarifying and formalizing important alignment concepts. I expect that most of these will be pretty time-consuming, but happy to discuss for people who want to try: A paper which does for deceptive alignment what the goal misgeneralization paper does does for inner alignment, i.e. describing it in ML language and setting up toy examples (for example, telling GPT-3 to take actions which minimize changes in its weights, given that it’s being trained using actor-critic RL with a certain advantage function, and seeing if it knows how to do so). A paper which does the same for gradient hacking, e.g. taking these examples and putting them into more formal ML language. A list of papers that it’s particularly useful for new research engineers to replicate. A takeover scenario which covers all the key points in/, but not phrased as an argument, just phrased as a possible scenario (I think you can’t really make the argument rigorously in that little space). A paper which defines the concepts of implicit planning, implicit value functions, implicit reward models, etc, in ML terms. Kinda like but more AGI-focused. I want to be able to ask people “does GPT-3 choose actions using an implicit value function?” and then be able to point them to this paper to rigorously define what I mean. I discuss this briefly in the phase 1 section here. A blog post which describes in as much detail as possible what our current “throw the kitchen sink at it” alignment strategy would look like. (I’ll probably put my version of this online soon but would love others too). A blog post explaining “debate on weights” more thoroughly. A blog post exploring how fast we should expect a forward pass to be for the first AGIs - e.g. will it actually be slower than human thinking, as discussed in this comment? A blog post exploring considerations for why model goals may or may not be much more robust to SGD than model beliefs, as discussed in framing 3 here. (See also this paper on gradient starvation - h/t Quintin Pope.) A blog post explaining why the “uncertainty” part of CIRL only does useful work insofar as we have an accurate model of the human policy, and why this is basically just as hard as having an accurate model of human preferences. A blog post explaining what practical implications Stuart Armstrong’s impossibility result has. As many alignment exercises as possible to help people learn to think about this stuff (mine aren't great but I haven’t seen better). A paper properly formulating instrumental convergence, generalization to large-scale goals, etc, as inductive biases in the ML sense (I do this briefly in phase 3 here). A mathematical comparison between off-policy RL and imitation learning, exploring ways in which they’re similar and different, and possible algorithms in between. A blog post explaining the P != NP argument for why adversarial training is likely to fail to get rid of misalignment, and arguments for why it might nevertheless succeed. A blog post exploring the incentives which models might have when they’re simultaneously trained to make predictions and to take actions in an RL setting (e.g. models trained using RL via sequence modeling). A blog post exploring pros and cons of making misalignment datasets for use as a metric of alignment (alignment = how much training on the misalignment dataset is needed to make it misaligned). A paper providing an RL formalism in which reward functions can depend on weights and/or activations directly, and demonstrating a simple but non-trivial example. A blog post evaluating reasons to think that situational awareness will be a gradual development in models, versus a sharp transitio...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Survey advice, published by KatjaGrace on August 24, 2022 on LessWrong. Things I believe about making surveys, after makingsome surveys: If you write a question that seems clear, there’s an unbelievably high chance that any given reader will misunderstand it. (Possibly this applies to things that aren’t survey questions also, but that’s a problem for another time.) A better way to find out if your questions are clear is to repeatedly take a single individual person, and sit down with them, and ask them to take your survey while narrating the process: reading the questions aloud, telling you what they think the question is asking, explaining their thought process in answering it. If you do this repeatedly with different people until some are not confused at all, the questions are probably clear. If you ask people very similar questions in different sounding ways, you can get very different answers (possibly related to the above, though that’s not obviously the main thing going on). One specific case of that: for some large class of events, if you ask people how many years until a 10%, 50%, 90% chance of event X occurring, you will get an earlier distribution of times than if you ask the probability that X will happen in 10, 20, 50 years. (I’ve only tried this with AI related things, but my guess is that it at least generalizes to other low-probability-seeming things. Also, if you just ask about 10% on its own, it is consistently different from 10% alongside 50% and 90%. Given the complicated landscape of people’s beliefs about the world and proclivities to say certain things, there is a huge amount of scope for choosing questions to get answers that sound different to listeners (e.g. support a different side in a debate). There is also scope for helping people think through a thing in a way that they would endorse, e.g. by asking a sequence of questions. This can also change what the answer sounds like, but seems ethical to me, whereas applications of 5 seem generally suss. Often your respondent knows thing P and you want to know Q, and it is possible to infer something about Q from P. You then have a choice about which point in this inference chain to ask the person about. It seems helpful to notice this choice. For instance, if AI researchers know most about what AI research looks like, and you want to know whether human civilization will be imminently destroyed by renegade AI systems, you can ask about a) how fast AI progress appears to be progressing, b) when it will reach a certain performance bar, c) whether AI will cause something like human extinction. In the 2016 survey, we asked all of these. Given the choice, if you are hoping to use the data as information, it is often good to ask people about things they know about. In 7, this points to aiming your question early in the reasoning chain, then doing the inference yourself. Interest in surveys doesn’t seem very related to whether a survey is a good source of information on the topic surveyed on. One of the strongest findings of the 2016 survey IMO was that surveys like that are unlikely to be a reliable guide to the future. This makes sense because surveys fulfill other purposes. Surveys are great if you want to know what people think about X, rather than what is true about X. Knowing what people think is often the important question. It can be good for legitimizing a view, or letting a group of people have common knowledge about what they think so they can start to act on it, including getting out of bad equilibria where everyone nominally supports claim P because they think others will judge them if not. If you are surveying people with the intention of claiming a thing, it is helpful to think ahead about what you want to claim, and make sure you ask questions that will let you claim that, in a simple way....
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Nate Soares' Life Advice, published by CatGoddess on August 23, 2022 on LessWrong. Disclaimer: Nate gave me some life advice at EA Global; I thought it was pretty good, but it may or may not be useful for other people. If you think any of this would be actively harmful for you to apply, you probably shouldn't. Notice subtle things in yourself This includes noticing things like confusion, frustration, dissatisfaction, enjoyment, etc. For instance, if you're having a conversation with somebody and they're annoying you, it's useful to notice that you're getting a little frustrated before the situation gets worse. A few weeks ago my colleagues and I wanted to do something fun, and decided to play laser tag at our workplace. However, we couldn't find the laser tag guns. As I began to comb the grounds for the guns for the second time I noticed that I felt like I was just going through the motions, and didn't really expect my search to be fruitful. At this point I stopped and thought about the problem, and realized that I had artificially constrained the solution space to things that would result in us playing laser tag at the office, rather than things that would result in us having fun. So I stopped looking for the guns and we did an escape room instead, which made for a vastly more enjoyable evening. If you're not yet at the point where you can notice unsubtle things in yourself, you can start by working on that and move up from there. Keep doing the best thing, even if you don't have a legible story for why it's good Certainly the actions you're taking should make sense to you, but your reasoning doesn't have to be 100% articulable, and you don't need to justify yourself in an airtight way. Some things are easier to argue than other things, but this is not equivalent to being more correct. For instance, I'm doing AI alignment stuff, and I have the option of reading either a textbook on linear algebra or E.T. Jaynes' probability theory textbook. Reading about linear algebra is very easy to justify in a way that can't really be disputed; it's just obviously true that linear algebra is directly and widely applicable to ML. It's harder to justify reading Jaynes to the same level, even though I think it's a pretty sound thing to do (I think I will become better at modeling the world, learn about various statistical pitfalls, absorb Jaynes' philosophical and historical insights, etc.), and in fact a better use of my time right now than learning linear algebra in more depth. This bit of advice is mostly about not needing to be able to justify yourself to other people (e.g. friends, family) to take the best visible action. However, it is also the case that you might have internalized social pressure such that you feel the need to justify a course of action to yourself in a way that would be legible to other people/justifiable in a social setting. This is also unnecessary. Relatedly, you don't need to "get" motivation; you can just continue to take the best action you can see. Don't go insane Apparently a good number of people in Nate's social circle have gone insane - specifically, they have taken facts about the world (e.g. the universal prior being malign) as "invitations" to go insane. He also noted that many of these people took LSD prior to going insane, and that this may have "loosened" something in their minds. This may be a particular danger for people who value taking ideas seriously as a virtue, because they might go full throttle on an idea that conflicts with common sense, and end up insane as a result. When asking a non-Nate for feedback on this post, I was told that some concrete things that people have taken as "invitations" to go insane are: decision theory (specifically acausal trade), things thought while meditating, and the idea that minds are made of "parts"...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AI art isn't "about to shake things up". It's already here., published by Davis Kingsley on August 22, 2022 on LessWrong. For a while, I've been seeing people commenting about how AI art is on the cusp of shaking up the art world. Quite frankly, at this point I consider such sentiments to be behind the times. AI art is not "on the cusp" of disrupting such things. It is already here. Using only capabilities that are straightforwardly and publicly available right now, AI art has radically transformed the budget for certain types of project that require art.Let's give a basic example. I play card games of various types -- games like Magic: the Gathering or similar -- on a highly competitive level. I've even been involved in testing and development for some such games, so I'm familiar with that process as well. I am not a professional game designer, but am fairly involved with some aspects of the field.For one of these games, a typical budget for an individual card illustration of sufficient quality is, as I understand it, in the realm of $200-1000 USD. A single "set" of cards might require 100+ such illustrations. That means that you're looking at paying $20k in art costs on the low end, and that this is a recurring cost every time you want to make a new set that isn't just reprinting old cards -- and even then, sometimes new art is used for reprints!Except, well, that was then and this is now. Now, if I were in the business of making a card game set, my art budget wouldn't be $20k-100k. It would be... a $30/month Midjourney subscription with $20/month private visibility enabled, and quite frankly high quality Midjourney images look better than many of the images already being used for art in these games. [1] I'm going to highlight that again. The price for the art needed to create a set of a hundred cards just went from twenty thousand dollars-- at the low end -- to fifty bucks a month. This is an extreme shift, and it is already here. This is not something that is based on a press release or future development that hasn't arrived yet. This is something that I could do today, using only techniques that are widely known and publicly available. If you assume you can get the art needed in one month of Midjourney time, that's four hundred times cheaper.There are many other areas where this applies. What's the price for a book cover? Quite frankly, that's not my field -- but whatever it is, I'm going to bet that Midjourney is often going to be cheaper and better. What's the price for the internal illustrations in a role-playing game manual? Again, whatever it is I'm going to bet that AI art is already beating it.Further, AI art is much easier to work with than professional artists. This is not intended as an insult to professional artists by any means! However, if I am working with a professional artist on an image, it may take them a significant amount of time to produce the image and get back to me on that. By contrast, if I don't like the Midjourney output I can write a variation on the prompt and get a new set of images extremely quickly. And I don't have to worry about people missing or misreading my emails, a potential language barrier, or time zone issues. [2]Now, there are admittedly some things that AI art isn't good at (the really big one being art with integrated text). You know what? That's true. There are definitely some things that AI art does not handle well. It's not a perfect substitute yet. However, given the outrageous cost savings, I am perfectly fine with that. I am altogether willing to change the focus of my card illustrations a bit in order to avoid areas where AI art generation does poorly if it means paying four hundred times less for art for my game, and I suspect others will soon be choosing the same.So, yeah. AI art isn't some hypothetically dis...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: What's the Least Impressive Thing GPT-4 Won't be Able to Do, published by Algon on August 20, 2022 on LessWrong. It seems like GPT-4 is going to be coming out soon and, so I've heard, it will be awesome. Now, we don't know anything about its architecture or its size or how it was trained. If it were only trained on text (about 3.2 T tokens) in an optimal manner, then it would be about 2.5X the size of Chinchilla i.e. the size of GPT-3. So to be larger than GPT-3, it would need to be multi-modal, which could present some interesting capabilities. So it is time to ask that question again: what's the least impressive thing that GPT-4 won't be able to do? State your assumptions to be clear i.e. a text and image generating GPT-4 in the style of X with size Y can't do Z. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: The Loire Is Not Dry, published by jefftk on August 20, 2022 on LessWrong. Several of my friends have shared this paste (or shares of it) on Facebook: This is the current state of the Loire, the longest river in France. This has not happened before in at least the 2000 years since literate people inhabited France. The Romans would have written about this. The medieval Franks would have written about this. To the best of our historical knowledge, nowhere in the past 2000 years has the Loire run dry, and likely never long before that. The drought that now grips southwestern Europe may well be unprecedented in recorded human history. This is in the heart of wine country where grapes grow in abundance and wheat waves like golden seas- but not now. Now the wheat burns and the grapes whither to raisins on the vine. This is the end of days. And on Monday morning I'll return to work and pretend this isn't happening. It's complete madness. Along with this picture: But it's all bunk. While the Loire going dry would be a big deal, it's not dry, and this picture isn't unusual. The picture is of parallel path of the Loire where its crossed by the Pont de Varades bridge. The main flow of the river is just East of this picture, off to the left: Google Maps Translating some tweets from Thibault Laconde, @EnergieDevlpmt on Twitter: "It spans a shallow dead arm which is often dry. Here are pictures from May 2011 and August 2009: ..." "The Montjean station was built in 1842, so we have lots of historical data. It had lower flows (below 95m3/s) during the drought of 1976 (with a minimum of 73m3/s on the 22 August) but also in 1870, 1905, 1906, 1921, 1947, 1949, 1950, etc." "The current flow is low but not unprecedented. It corresponds almost exactly to the decadal low VCN3. That is to say, at a constant climate, we expect to reach it over 3 days on average every 10 years." There is definitely a drought, and the heatwave has broken many records, but these water levels are not anywhere close to a "haven't seen in 2000 years" event. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Covid 8/18/22: CDC Admits Mistakes, published by Zvi on August 18, 2022 on LessWrong. Two Covid-related things happened this week that I did not expect. The CDC admitted that it had failed us during the pandemic, withdrew at long last many of its remaining recommendations and promised reforms to be less academic and otherwise do better. A new paper found potential biological markers for Long Covid, claiming it can be identified via tests that match patient self-reports almost all the time. This points towards potential progress in treatment, and more generally in Long Covid being much more concretely A Thing that might exist and could be reasoned about. There are still flaws, and even without flaws there is still much work to do here. Most of the reasons not to be concerned remain, so up front: I do not think that this should substantially change anyone’s level of Covid precautions. Executive Summary CDC finally fully gives up on six foot social distancing and other measures. CDC admits some failure, promises reforms that seem potentially promising. New study finds potential biological markers for Long Covid. Also I think someone said something about over the counter hearing aids? Let’s run the numbers. The Numbers Predictions Prediction from last week: 650k cases (-5%) and 3,200 deaths (+0%). Results: 602k cases (-12%) and 3,183 deaths (-1%). Prediction for next week: 560k cases (-8%) and 3,200 deaths (+1%). I do not know why cases declined more than expected but result seems robust so I see no reason to expect it not to continue at least somewhat. There won’t be enough time for the decline to impact deaths yet, so I’m mostly going with the null prediction there. There is not much meaningful uncertainty here. Deaths Cases Decline cuts across all four regions. We are fully into the BA.5 era with nothing on the near-term horizon to replace it, so things should be quiet for a few months at least. Physical World Modeling A guide to buying the right HEPA air filter. Bob Wachter sees no sign of anything that might replace BA.5. Thread also reminds us that it is a common mistake not to take into account the correlation between the Covid status of people who choose to be together in a group. Trevor Bedford however sees logistic growth in BA.2.75, although only with R0 ~ 1.3 (versus ~1 for BA.5) which isn’t that much of an advantage for a new strain taking over. My guess is Trevor is right and BA.2.75 will displace BA.5 over time, but that we will barely even notice. Some sad news. I opened a Manifold market on whether she’ll experience a rebound, to see how common people think such events are. This also brings up the question, what happens when you’re against actual physical world modeling and instead blindly follow CDC rules? They do know he had Covid-19 plus a rebound within the last few weeks? That it is absurd to think that being a ‘close contact’ puts him at relatively high risk for Covid-19 given that timeline? No. Of course not. This is not a man to concern himself with whether or not an action makes physical sense. Almost no children under 5 are getting vaccinated, even weaker ‘than experts feared.’ Doses are being discarded due to lack of demand. This should not be a fear so much as a revealed preference. If we had approved these doses sooner, I am guessing we would have had much higher uptake, although still nothing that would have satisfied experts. At this point, people don’t care enough, especially given the logistics are frequently annoying. This raises the question of why we insisted in so many crazy precautions for these same young children for so long, in ways that I strongly believe did serious damage to their development and well-being. They were never at risk and everyone knew this well enough not to bother doing much about it when finally given the oppo...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Announcing Encultured AI: Building a Video Game, published by Andrew Critch on August 18, 2022 on LessWrong. Also available on the EA Forum.Preceded By: Encultured AI Pre-planning, Part 2: Providing a as used Service If you read to the end of our last post, you maybe have guessed: we’re building a video game! This is gonna be fun :) Our homepage:/ Will Encultured save the world? Is this business plan too good to be true? Can you actually save the world by making a video game? Well, no. Encultured on its own will not be enough to make the whole world safe and happy forever, and we'd prefer not to be judged by that criterion. The amount of control over the world that's needed to fully pivot humanity from an unsafe path onto a safe one is, simply put, more control than we're aiming to have. And, that's pretty core to our culture. From our homepage: Still, we don’t believe our company or products alone will make the difference between a positive future for humanity versus a negative one, and we’re not aiming to have that kind of power over the world. Rather, we’re aiming to take part in a global ecosystem of companies using AI to benefit humanity, by making our products, services, and scientific platform available to other institutions and researchers. Our goal is to play a part in what will be or could be a prosperous civilization. And for us, that means building a successful video game that we can use in valuable ways to help the world in the future! Fun is pretty good target for us to optimize You might ask: how are we going to optimize for making a fun game and helping the world at the same time? The short answer is that creating a game world in which lots of people are having fun in diverse and interesting ways in fact creates an amazing sandbox for play-testing AI alignment & cooperation. If an experimental new AI enters the game and ruins the fun for everyone — either by overtly wrecking in-game assets, subtly affecting the game culture in ways people don't like, or both — then we're in a good position to say that it probably shouldn't be deployed autonomously in the real world, either. In the long run, if we're as successful we hope as a game company, we can start posing safety challenges to top AI labs of the form "Tell your AI to play this game in a way that humans end up endorsing." Thus, we think the market incentive to grow our user base in ways they find fun is going to be highly aligned with our long-term goals. Along the way, we want our platform to enable humanity to learn as many valuable lessons as possible about human↔AI interaction, in a low-stakes game environment before having to learn those lessons the hard way in the real world. Principles to exemplify In preparation for growing as a game company, we’ve put a lot of thought into how to ensure our game has a positive rather than negative impact on the world, accounting for its scientific impact, its memetic impact, as well as the intrinsic moral value of the game as a positive experience for people. Below are some guiding principles we’re planning to follow, not just for ourselves, but also to set an example for other game companies: Pursue: Fun! We’re putting a lot of thought into not only how our game can be fun, but also ensuring that the process of working at Encultured and building the game is itself fun and enjoyable. We think fun and playfulness are key for generating outcomes we want, including low-stakes high-information settings for interacting with AI systems. Maintain: opportunities to experiment. No matter how our product develops, we’re committed to maintaining its value as a platform for experiments, especially experiments that help humanity navigate the present and future development of AI technology. Avoid: teaching bad lessons. On the margin, we expect our game to incentivize coo...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: My thoughts on direct work (and joining LessWrong), published by RobertM on August 16, 2022 on LessWrong. Epistemic status: mostly a description of my personal timeline, and some of my models (without detailed justification). If you just want the second, skip the Timeline section. My name is Robert. By trade, I'm a software engineer. For my entire life, I've lived in Los Angeles. Timeline I've been reading LessWrong (and some of the associated blog-o-sphere) since I was in college, having found it through HPMOR in ~2011. When I read Eliezer's writing on AI risk, it more or less instantly clicked into place for me as obviously true (though my understanding at the time was even more lacking than it is now, in terms of having a good gears-level model). This was before the DL revolution had penetrated my bubble. My timelines, as much as I had any, were not that short - maybe 50-100 years, though I don't think I wrote it down, and won't swear my recollection is accurate. Nonetheless, I was motivated enough to seek an internship with MIRI after I graduated, near the end of 2013 (though I think the larger part of my motivation was to have something on my resume). That should not update you upwards on my knowledge or competence; I accomplished very little besides some reworking of the program to mail paperback copies of HPMOR to e.g. math olympiad winners and hedge fund employees. Between September and November, I read the transcripts of discussions Eliezer had with Richard, Paul, etc. I updated in the direction of shorter timelines. This update did not propagate into my actions. In February I posted a question about whether there were any organizations tackling AI alignment, which were also hiring remotely. Up until this point, I had been working as a software engineer for a tech company (not in the AI/ML space), donating a small percentage of my income to MIRI, lacking better ideas for mitigating x-risk. I did receive a few answers, and some of the listed organizations were new to me. For reasons I'll discuss shortly, I did not apply to any of them. On April 1st, Eliezer posted Death with Dignity. My system 1 caught up to my system 2 on what shorter timelines meant. That weekend, I talked to my family about leaving LA. On April 3rd I pinged Ruby to see if he'd be open to a chat about Lightcone's interview process. Over the next month I had a couple more calls with Ruby, found an opportunity to drop by the Lightcone offices to talk to habryka (since I was traveling to the Bay anyways), and went through a couple tech screens. In May I flew back to up Berkeley to go through the final stage of the interview - an on-site work trial. The trial must've gone ok; I started at Lightcone on July 5th. My primary focus will be LessWrong and the associated infrastructure. Why LessWrong? A problem I've been thinking about lately is: imagine you want to reduce AI risk, but you aren't a researcher and are limited in how well you can evaluate the usefulness of the research direction and output of AI alignment organizations. With some exceptions, those organizations describe neither what they consider to be the core problems of AI alignment, nor their theory of change. How, then, do you figure out where your efforts are best directed? You can defer to another's judgement, if you know someone with an opinion on the question and have evidence of their ability to reason about difficult questions in domains you're more familiar with. Unfortunately, this heuristic encourages information cascades, but purely on a personal level it might be better than throwing darts. Absent that, or another good way to discriminate between options, the temptation is to kick the problem up a level and do "meta" work. This has some problems: The indirection this adds to the theory of change introduces additional deg...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Meditation course claims 65% enlightenment rate: my review, published by KatWoods on August 1, 2022 on LessWrong. TL;DR I eliminated my impostor syndrome and dramatically reduced my work-related anxiety. I did this in a way that I think can be replicated. Different techniques work for different people. If you want to get the benefits of meditation, you should experiment widely, then drill down on the methods that work for you. Don’t just try the same technique for months or years and hope you’ll eventually “get it” or give up and say meditation doesn’t work for you. Explore then exploit. Loving-kindness meditation is underrated and should be the main-course meditation for a lot of people. This 55-minute video is the 80/20 of the course. If you like it, you will probably like the rest of the course. I recommend the Finder’s Course for most people. If you follow the instructions, you will very likely become happier. If you prefer and are good at self-directed learning, you can do your own self-directed course and get similar benefits. What makes the Finder’s Course different: methodology Applying science to meditation isn’t unique to Jeffery Martin’s course. Fortunately for the world, there’s a whole movement around this. That being said, I haven’t heard of anything that seems more likely to figure out how to actually achieve enlightenment (or fundamental well-being (FWB) as he calls it, which I prefer). Most science I know of is doing things like putting meditators in brain scans and seeing if anything is different from regular brains, or running RCTs to see if meditation makes you happier. This is foundational and important to do. However, it’s very black box thinking and doesn’t give you any gears-level understanding of how to achieve fundamental well-being. Meditation classes usually teach a variety of different techniques. Which techniques are causing the change? Most studies focus on averages, which ignores the thing we’re most interested in - those outliers who don’t just start feeling less stressed but have eliminated suffering. Who are living in states of profound bliss and serenity. How does that show up on a psychological item asking “On a scale of 1 to 10, how satisfied are you with your life?”? The Finder’s Course on the other hand clearly followed a methodology that was truly trying to solve the problem. The way he did this was to find over 1,000 people saying that they had achieved fundamental well-being, and he went and interviewed all of them. The interviews would often last up to twelve hours. He asked them what their experiences were, what had gotten them there, and ran them through batteries of psychological tests. From this exploratory research, he started pulling out patterns. You can read some of the results of his research in his book. He took the top findings out of his research and turned it into a course. It was formerly called the Finder’s Course (a play on the usual spiritual terminology of people being “seekers”). He continues to do science on the course, and this is part of what most intrigued me. It’s mandatory for the course to take a whole battery of psychological evaluations before and after, such as PERMA, satisfaction with life scale, CES-D Questionnaire, etc. After doing a week of each technique, he also does a shorter survey. The results from this he claims are that 65% of people who finish the course achieve fundamental well-being. This is an incredible claim, and I figured it was probably just hype. Then I spoke to a friend he said that he’d recently done the course and could now go into fundamental well-being at will. This is what inspired me to give it a go. Before I started the course, I publicly pre-committed to writing about my experience, regardless of how it went. Usually, people only write about something if it goes part...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Focusing, published by CFAR!Duncan on July 29, 2022 on LessWrong. Epistemic status: Firm The Focusing technique was developed by Eugene Gendlin as an attempt to answer the question of why some therapeutic patients make significant progress while others do not. Gendlin studied a large number of cases while teasing out the dynamics that became Focusing, and then spent a significant amount of time investigating whether his technique-ified version was functional and efficacious. While the CFAR version is not the complete Focusing technique, we have seen it be useful for a majority of our alumni. If you’ve ever felt your throat go suddenly dry when a conversation turned south, or broken out into a sweat when you considered doing something scary, or noticed yourself tensing up when someone walked into the room, or felt a sinking feeling in the pit of your stomach as you thought about your upcoming schedule and obligations, or experienced a lightness in your chest as you thought about your best friend’s upcoming visit, or or or or ... If you’ve ever had those or similar experiences, then you’re already well on your way to understanding the Focusing technique. The central claim of Focusing (at least from the CFAR perspective) is that parts of your subconscious System 1 are storing up massive amounts of accurate, useful information that your conscious System 2 isn’t really able to access. There are things that you’re aware of “on some level,” data that you perceived but didn’t consciously process, competing goalsets that you’ve never explicitly articulated, and so on and so forth. Focusing is a technique for bringing some of that data up into conscious awareness, where you can roll it around and evaluate it and learn from it and—sometimes—do something about it. Half of the value comes from just discovering that the information exists at all (e.g. noticing feelings that were always there and strong enough to influence your thoughts and behavior, but which were somewhat “under the radar” and subtle enough that they’d never actually caught your attention), and the other half comes from having new threads to pull on, new models to work with, and new theories to test. The way this process works is by interfacing with your felt senses. The idea is that your brain doesn’t know how to drop all of its information directly into your verbal loop, so it instead falls back on influencing your physiology, and hoping that you notice (or simply respond). Butterflies in the stomach, the heat of embarrassment in your cheeks, a heavy sense of doom that makes your arms feel leaden and numb—each of these is a felt sense, and by doing a sort of gentle dialogue with your felt senses, you can uncover information and make progress that would be difficult or impossible if you tried to do it all “in your head.” On the tip of your tongue We’ll get more into the actual nuts and bolts of the technique in a minute, but first it’s worth emphasizing that Focusing is a receptive technique. When Eugene Gendlin was first developing Focusing, he noticed that the patients who tended to make progress were making lots of uncertain noises during their sessions. They would hem and haw and hesitate and correct themselves and slowly iterate toward a statement they could actually endorse: “I had a fight with my mother last week. Or—well—it wasn’t exactly a fight, I guess? I mean—ehhhhhhh—well, we were definitely shouting at the end, and I’m pretty sure she’s mad at me. It was about the dishes—or at least—well, it started about the dishes, but then it turned into—I think she feels like I don’t respect her, or something? Ugh, that’s not quite right, I’m pretty sure she knows I respect her. It’s like—hmmmmm—more like there are things she wants—she expects—she thinks I should do, just because—because of, I dunno, like tradi...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Looking back on my alignment PhD, published by TurnTrout on July 1, 2022 on LessWrong. The funny thing about long periods of time is that they do, eventually, come to an end. I'm proud of what I accomplished during my PhD. That said, I'm going to first focus on mistakes I've made over the past four years. Mistakes I think I got significantly smarter in 2018–2019, and kept learning some in 2020–2021. I was significantly less of a fool in 2021 than I was in 2017. That is important and worth feeling good about. But all things considered, I still made a lot of profound mistakes over the course of my PhD. Social dynamics distracted me from my core mission I focused on "catching up" to other thinkers I figured this point out by summer 2021. I wanted to be more like Eliezer Yudkowsky and Buck Shlegeris and Paul Christiano. They know lots of facts and laws about lots of areas (e.g. general relativity and thermodynamics and information theory). I focused on building up dependencies (like analysis and geometry and topology) not only because I wanted to know the answers, but because I felt I owed a debt, that I was in the red until I could at least meet other thinkers at their level of knowledge. But rationality is not about the bag of facts you know, nor is it about the concepts you have internalized. Rationality is about how your mind holds itself, it is how you weigh evidence, it is how you decide where to look next when puzzling out a new area. If I had been more honest with myself, I could have nipped the "catching up with other thinkers" mistake in 2018. I could have removed the bad mental habits using certain introspective techniques; or at least been aware of the badness. But I did not, in part because the truth was uncomfortable. If I did not have a clear set of prerequisites (e.g. analysis and topology and game theory) to work on, I would not have a clear and immediate direction of improvement. I would have felt adrift. But there is not yet any "rationality tech tree" set of well-defined prerequisite rationality skills such that you can learn them in order and grow way stronger. Like, you can't just do the calibration exercises, and then the noticing-confusion exercises, and then other things. Those tools help, but they aren't enough. There won't be a clear and immediate direction of improvement, at first. But you may want to get stronger anyways. I focused on seeming smart and defensible I figured this point out this spring. When I started working on alignment, I didn't know what to do at first, and I felt insecure about my credentials. As far as I remember, I figured I'd start off by becoming respected, since other people's feedback was initially a better guide than my own taste. Unfortunately, I didn't realize how deeply and subtly this goal would grow its roots. I worried about upvotes, I worried about winning arguments, I worried about being defensible against criticism. I was so worried that someone would comment on one of my posts and tear everything down, because I hadn't been careful enough, because I had left myself open by not dotting all my 'i's. (Not that anyone has ever done that on LessWrong before...) I think it was this year that I had my (second) "oh man, don't forget the part where everyone is allowed to die to AI" moment. To illustrate the new mindset this gut-realization gave me, I'll detail a recent decision with social consequences, and then compare the old and the new mindsets. A few months back, Quintin Pope approached me with (what he claimed to be) a new alignment paradigm, which blossomed from asking the following kind of questions: We clearly prefer future AIs to generalize in the way that neuroscientists generalize, so it seems worthwhile to ask: "why don't neuroscientists wirehead themselves?" It's clearly not because humans evolved away...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Failing to fix a dangerous intersection, published by alyssavance on June 30, 2022 on LessWrong. Over the last few years, many people have written about why America can't build things anymore (eg. here, although this is just one of hundreds of relevant essays). Ten years ago, when I was 21 and had just graduated, a friend told me about a dangerous intersection in Berkeley. I tried writing to the city and asking for it to be fixed; I'm posting the email and reply here as a useful data point. My request: To the Berkeley Department of Transportation: I write to urge the City to address a dangerous intersection, at San Pablo and Gilman Streets. When approaching the intersection from the west (on Gilman), there are two lanes of roadway, of which one is left-and-straight and the other is right-and-straight. However, on the opposite side of the intersection (east of San Pablo), there is only one lane of roadway. Hence, cars going straight are forced to merge in the middle of the intersection (with no warning), which is time-consuming and hazardous. I and some fellow Berkeley residents propose that the lane arrows be modified, such that either the left lane is left only, or the right lane is right only. This way, there is only one lane of forward traffic, and cars do not have to merge in the intersection. This modification would cost virtually nothing, and would make driving easier for the thousands of City residents who use this intersection daily, as well as preventing a potentially fatal accident. We greatly appreciate your consideration. Their reply is below. (Email is generally private, but in this case the communication, from a city employee about a government issue, should be an open record under the California Public Records Act.) Dear Alyssa, I forwarded your email to the Supervising Traffic Engineer, and have been asked to respond to your request with the following information. San Pablo Avenue (State Highway 123) is under Caltrans jurisdiction, and any significant changes to the intersection must be approved by Caltrans. Recent communications from Caltrans indicate they have no immediate plans to upgrade their non-freeway facilities. In theory the City could develop a modified striping plan concept and ask for approval from Caltrans to proceed. If Caltrans were to agree to the concept, the City would need to assign resources to develop the conceptual design, then a detailed design, obtain approvals from Caltrans, then put out and pay for a contract for implementation. Though it seems simple enough as an idea, it does involve closing a state highway intersecting with a major collector street, and would be a significant resource-intensive task for the City to plan and execute a temporary traffic management plan. We regret that we are unable to proceed any further with your request at this time, as it is not currently in our budget or in our approved Work Plan. We do keep a "wish list" of unfunded projects from which we can draw should the appropriate funding opportunity arise, perhaps through mitigation funding from a future development project. We already have a signal upgrade (with left turn phase) for the Gilman/San Pablo intersection on this wish list, and will add your less- costly suggestion for a lane-restriping/reconfiguration project to the list. Unfortunately, that is all we are able to do at this time. As of 2021, the intersection has still not been fixed. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Kurzgesagt – The Last Human (Youtube), published by habryka on June 29, 2022 on LessWrong. TLDW: Big Youtube channel Kurzgesagt released a video on the potential of humanity's future that's I think my favorite videos of theirs yet. I have a pretty high bar for posting video content to LessWrong, but this video seemed better than most other videos in terms of capturing some of the things that make me interested in a lot of things on LessWrong. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: It’s Probably Not Lithium, published by Natália Mendonça on June 28, 2022 on LessWrong. A Chemical Hunger (a), a series by the authors of the blog Slime Mold Time Mold (SMTM) that has been received positively on LessWrong, argues that the obesity epidemic is entirely caused (a) by environmental contaminants. The authors’ top suspect is lithium (a), primarily because it is known to cause weight gain at the doses used to treat bipolar disorder. After doing some research, however, I found that it is not plausible that lithium plays a major role in the obesity epidemic, and that a lot of the claims the SMTM authors make about the topic are misleading, flat-out wrong, or based on extremely cherry-picked evidence. I have the impression that reading what they have to say about this often leaves the reader with a worse model of reality than they started with, and I’ll explain why I have that impression in this post. (Preamble) A brief summary of their hypotheses The SMTM authors have recently (a) summarized their hypotheses on how lithium exposure could explain the obesity epidemic. The first hypothesis is that trace exposure is responsible: One possibility is that small amounts of lithium are enough to cause obesity, at least with daily exposure. And the second one is that people are intermittently exposed to therapeutic doses: [E]ven if people aren’t getting that much lithium on average, if they sometimes get huge doses, that could be enough to drive their lipostat upward. I am going to argue that neither of those is plausible. I address the plausibility of the second hypothesis in the next section, and the plausibility of the first one in the rest of the post. Lithium exposure in the general population is extremely low, even at the tails, in the majority of countries for which we have data A few days ago, the SMTM authors published a literature review (a) on the lithium content of food. They conclude that, whereas the existing literature isn’t great, “[i]t seems like most people get at least 1 mg [of lithium] a day from their food, and on many days, there’s a good chance you’ll get more.” They also say it seems plausible that people are intermittently exposed to doses of lithium within the therapeutic range through their diet. However, their literature review pretty much only includes studies that are outliers in the literature. Moreover, they use a misleading threshold for the therapeutic range of lithium. I’ll explain. The studies in SMTM’s literature review of lithium levels in food are pretty much all outliers In 2006, France conducted its second Total Diet Study (henceforth TDS). Across 1,319 food samples, the highest lithium concentration found was 0.6 mg/kg, in water. That’s not the highest average concentration among food groups – it’s the highest concentration of any single sample they tested. (For context, a standard clinical dose of elemental lithium is about 200 mg/day, or 1 gram/day of lithium carbonate.) Similarly, New Zealand’s 2016 TDS examined 1,056 food samples and the highest concentration it found in any single sample was 0.54 mg/kg (in mussels). Canada makes the raw data of its Total Diet Study publicly available (a), and they too measure the lithium content of their food. The maximum level reported is 1.1 mg/kg (in table salt, which is presumably rarely consumed in kilogram quantities) across 479 food samples, with the mean being 25 µg/kg and the median 11 µg/kg. Here’s a histogram of the data: Excluding table salt, the maximum value in the rest of the dataset (N = 476) is 0.4 mg/kg, in mineral water. Total Diet Studies in other countries report similarly low levels. Using data from the UK’s 1994 TDS (which included 400 food samples), the mean daily lithium intake among adults was estimated to be 17 µg/day, more than 50 times lower than SMTM’s estima...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Some reflections on the LW community after several months of active engagement, published by M. Y. Zuo on June 25, 2022 on LessWrong. There seems to be some folks who might derive useful insights from a third-party, and mostly neutral, perspective of how the community appears after an honest and sustained effort of engagement, someone who doesn't really place AI risk as their top priority but who also doesn't completely ignore it like some critics, or opponents, of LW might. Notably I've encountered some folks who had strong personal opinions one way or another but refrained from writing them in a public, or even pseudo-anonymous, manner. There also appears to be a large group of lurkers or once-in-a-blue-moon posters who nonetheless have some views of the community and might benefit from someone willing to take the risk to do a write-up. First off, addressing the popular critiques and praises: There has definitely been some evaporative cooling of the community in the past decade or so. Some of the most insightful members have gone on to do bigger things, and the average quality of new posts is somewhat less than where it was a decade ago. Or so far as I can tell via the archives. This isn't very surprising as this is the common trajectory of every community that rapidly grows in size. It would have taken a super-human effort to retain the same level of quality going from 100 to 1000 users, let alone from 1000 to 10000, and so on. So I don't think that would have been a fair expectation to place on the moderators, or anyone else involved, of a decade ago. On the flip side, there is a larger cross-section of society represented in the 2022 userbase, And there has been a correspondent softening of the hard edges that may have been off putting to some a decade ago. Relatedly, the proportion of really bizarre or challenging writing has gone down, for better and for worse. For example, there definitely does appear to be some unique benefit from a community with lots of oddballs with jarring writing styles, but the downside is obvious because nobody really desires to have their norms be constantly challenged in every paragraph. There has been growing focus, and emphasis, on advancements in AI risk and alignment, so LW does seem to be less of a catch-all forum than before. This is clearly better for those who wish to focus and really get into the details. But I can sympathize with the critique that there's less charm compared to a group of unconstrained folks exploring in a hundred different directions. Some of the common terminology is indeed puzzlingly unique, and thus promotes a distinctive writing style that does, in extreme cases, seem to be satire compared to the academic norm. I found it personally difficult to adjust to, and as you can tell by my somewhat varying writing style, I haven't found the best way to write in a conversational yet concise manner while incorporating the terminology. There is some charm, and exciting challenge, in trying to craft writing that isn't dry and aloof yet still remains accurate enough to describe highly complex and technical ideas. So many of the critics seem to be missing the forest for the trees. And I can see very little direct harm in being an outlier in writing style, and quite a lot of benefit from having such a unique differentiator. Certainly some the best Fanfiction I've ever read came from community members, and I highly doubt they would be nearly as popular if standard academic jargon were used. That being said, the critics do have a reasonable argument when it comes to comparisons with other notable online forums. LW does seem slightly more insular in some respects than SSC, Overcoming Bias, HN, etc., though that is understandable given the unique origins of the community and the developments that have taken place since....
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Popular education in Sweden: much more than you wanted to know, published by Henrik Karlsson on May 17, 2022 on LessWrong. Growing up on the Swedish seaside, I had a five-minute walk to four open learning facilities – not counting the library and the youth center. It was very Christopher Alexander. One of the premises was an abandoned church that my friends and I used as a recording studio; we'd renovated it ourselves with funding from a study association. There we played distorted pop. In another, I learned French from an émigré of Montpellier. We arranged public lectures – once, to our great surprise, we even managed to book then general secretary of the United Nations Ban-Ki Moon for a lecture in Uppsala. I analyzed Soviet cinema with a group of whom an unsettling number sang Sång för Stalin before the screenings. Since leaving Sweden, I have realized that not everyone grows up like this. And I miss it. In fact, if the whole of Sweden was about to burn down and I could only save one thing, I might grab just folkbildningsrörelsen. Folkbildningsrörelsen: that is the name we have for this movement of self-organized study groups, resource centers, maker spaces, public lectures, and free retreats for personal development. These types of things exist in other countries too – but not at the same scale. Or even close. To get a sense of how comprehensive folkbildningsrörelsen is, it helps to remember that Sweden has a population roughly comparable to New York City. If NYC had as many free resource centers per inhabitant as the municipality where I grew up, Manhattan would look like this: I was going to do all of New York but my hand started hurting from making all these dots, so I only managed the tip of Manhattan. At every other intersection, there would be a few rooms where you could go in and get some money to buy literature or access tools you needed. (In practice, the resource centers in cities tend to be lumped together in larger units, but the map still captures a lived reality for the 7.5 percent of Sweden's population who regularly take part in study associations.) Experientially, the spaces I have been part of have felt more like niche internet forums than schools. There were plenty of trolls, witches, and freaks – but we were also able to sustain a depth of conversation which was out of scope at school. When I entered university, seminars often felt like play-acting in comparison. In our often quite dilapidated buildings (as in internet communities), we hadn’t thought about what we were doing as learning. We were just obsessing about things. How did this all come about? In the 19th century, when these houses and the financing that enables them began to be built out, the main impetus came from the German Bildung tradition. Bildung etymologically refers to shaping yourself in the image (das Bild) of God. God in this context should be imagined as a highly self-possessed spectral being – in control of its emotions, with mind and heart in harmony, and willing to take individual moral responsibility. Think Bertrand Russell but less atheist, and sitting on a cloud. This is the look. In its original formulation, Bildung had a somewhat bourgeois flavor. It smelled of tweed and leather elbow patches. But in the early 1800s, thinkers such as Johann Heinrich Pestalozzi, N.F.S. Grundtvig, and Johann Friedrich Herbart figured out how to sell Bildung to farmers and day laborers – a folk Bildung, or folkbildning in Swedish. This was the tradition that took root in Sweden: the popular movement to shape yourself in the image of Bertrand Russell. The English language version of folkbildning’s Wikipedia page refers to it as popular education. This translation is not entirely correct. The term "popular education" has a strong political connotation – the Wikipedia page talks about "c...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: On saving one's world, published by Rob Bensinger on May 17, 2022 on LessWrong. If the world is likeliest to be saved by sober scholarship, then let us be sober scholars in the face of danger. If the world is likeliest to be saved by playful intellectual exploration, then let us be playful in the face of danger. Strategic, certainly; aware of our situation, of course; but let us not throw away the one mental mode that can actually save us, if that's in fact our situation. If the world is likeliest to be saved by honest, trustworthy, and high-integrity groups, who by virtue of their trustworthiness can much more effectively collaborate and much more quickly share updates; then let us be trustworthy. What is the path to good outcomes otherwise? CFAR has a notion of "flailing". Alone on a desert island, if you injure yourself, you're likelier to think fast about how to solve the problem. Whereas injuring yourself around friends, you're more likely to "flail": lean into things that demonstrate your pain/trouble to others. To my eye, a lot of proposals that we set aside sober scholarship, or playful intellectual exploration, or ethical integrity, look like flailing. I don't see an argument that this setting-aside actually chains forward into good outcomes; it seems performative to me, like hoping that if our reaction "feels extreme" enough, some authority somewhere will take notice and come to the rescue. Who is that authority? If you have a coherent model of this, we can talk about it and figure out if that's really the best strategy for eliciting their aid. But if no one comes to mind, consider the possibility that you're executing a social instinct that's adaptive to threats like tigers and broken legs, but maladaptive to threats like Unfriendly AI. If you feel scared about something, I generally think it's good to be honest about that fact and discuss it soberly, rather than hiding it. I don't think this is incompatible with rigorous scholarship or intellectual play. But I would clearly distinguish "being honest about your world-models and feelings, because honesty is legitimately a good idea" from "making it your main strategy to do whatever action sequence feels emotionally resonant with the problem". An "extreme" key doesn't necessarily open an "extreme" lock. A dire-sounding key doesn't necessarily open a dire-feeling lock. A fearful or angry key doesn't necessarily open a lock that makes you want to express fear or anger. Rather, the lock's exact physical properties determine which exact key (or set of keys) opens it, and we need to investigate the physical world in order to find the right key. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Why I'm Optimistic About Near-Term AI Risk, published by harsimony on May 15, 2022 on LessWrong. I'm not worried about AI posing an existential risk in the next 10-20 years. Recent developments in AI capabilities actually make me feel more optimistic about this. The fact that relatively simple models can perform a wide array of tasks suggests that we can build satisfactory AI without the need to use sophisticated, potentially dangerous agents in the near-term. My expectation for how AI will develop over the next decade is that companies will continue to focus on transformer-based foundation models. The general capability of these models will increase for a while simply by using more data, improving training procedures, and leveraging specialized hardware. Eventually, companies will start hitting bottlenecks in the amount of data required for optimal training at a given capability level. But before that, deployment of these systems will favor smaller, faster, and more auditable models leading companies to focus on distilled models specializing in specific tasks. These specialized models will be oriented towards augmenting human productivity, producing entertainment, or automating specific tasks. The slow pace at which industries change their practices and utilize the benefits of a new technology will moderate the adoption of AI. As adoption increases, these AI services will gain autonomy, producing more value at lower cost. Continued specialization will result in mostly autonomous AI's derived from generally capable foundation models that are distilled down for variety of tasks. I'm not claiming that these Tool AI's won't eventually be dangerous, but I can't see this path leading to high existential risk in the next decade or so. I think most people in the AI safety field would agree with me on this, so why write it up? I want to make this point explicit and foster a discussion about near-term AI safety. If AI will become dangerous soon, the field needs to act very quickly. Researchers would have to consider eschewing movement building, trading goodwill for influence, and gambling on near-term approaches. People who would take more than a decade to have an impact would have less reason to join in the first place while investments in infrastructure in the field would become less valuable. It's important for those concerned to be on the same page about near-term risks in order to avoid the unilaterialist's curse. Recent, pessimistic takes about AI risk make it seem superficially as if the consensus has shifted, but I don't think this is representative of the field as a whole. I remain optimistic that innovations in foundation models will produce a lot of value without a large increase in risk, providing more time to build. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Is AI Progress Impossible To Predict?, published by alyssavance on May 15, 2022 on LessWrong. People seem to be continually surprised, over and over again, by the new capabilities of big machine learning models, such as PaLM, DALL-E, Chinchilla, SayCan, Socratic Models, Flamingo, and Gato (all in the last two months!). Luckily, there is a famous paper on how AI progress is governed by scaling laws, where models predictably get better as they get larger. Could we forecast AI progress ahead of time by seeing how each task gets better with model size, draw out the curve, and calculate which size model is needed to reach human performance? I tried this, and apparently the answer is no. In fact, whether AI has improved on a task recently gives us exactly zero predictive power for how much the next model will improve on the same task. The sheer consistency of this unpredictability is remarkable, almost like a law of statistical thermodynamics. No matter what I plug in, the correlation is always zero! For example, does a task improving rapidly when you go from a small model to a 7B parameter model predict similar improvement when you go from a 7B model to Gopher's 280B? No: I tried making the same graph with MMLU tasks instead of BIG-bench, same result: What about DeepMind's new Chinchilla? Did rapid improvement of a task on Gopher predict continued improvement going from Gopher to Chinchilla? Nope: What about Google's PaLM? The full results of PaLM on BIG-bench don't seem to have been published yet, so I couldn't directly compare to Chinchilla or Gopher, but the PaLM paper described an 8B parameter model, a 62B model and a 540B model. Did fast improvement from 8B to 62B predict improvement from 62B to 540B? Not really, R^2 = 0.04: PaLM also provides data on 30 different NLU benchmark tasks. Plot those and you get the same thing: The results here seem pretty clear, but I'm honestly not sure how to interpret them. Before trying this, I assumed you would find that some tasks are "easy" and scale quickly, while others are "hard" and scale slowly. But that would get you high predictability, since fast progress between one pair of models would imply that the task is inherently "easy", and predict (perhaps with some noise) fast progress on the next pair. I didn't see that. You could also have a theory where tasks scaled similarly (all are of comparable "difficulty"), but there was some noise between model training runs, so that task performance on any given run would bounce up and down around some "true" average value. (Since if you did badly on one run, you'd expect to regress to the mean, and do unusually well on the next.) But I didn't see that either. The two effects (some tasks being intrinsically easier, and individual model runs being noisy) could also cancel out, since one implies a positive correlation and the other implies a negative one... but it seems unlikely that they would exactly cancel every time! Is AI task performance a type of submartingale, like a stock market index that goes up over time, but where each particular movement is intrinsically unpredictable? Maybe we can compare it to the growth in company profits, where the literature says that companies might grow slowly or quickly, but whether a company has grown fast recently has zero predictive power for future growth. I guess if we knew what we were doing, it wouldn't be called research. EDIT: By request, here's a Google sheet with the raw data, copy-pasted from the Gopher, PaLM and Chinchilla papers: EDIT 2: Several people suggested using logits instead of raw percentages. I tried that with the Gopher numbers, still got zero correlation: EDIT 3: Tamay noted that if you try to predict 7B Gopher from 1B Gopher, you get a negative correlation: If the models become small enough, maybe this means that scale is...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: "Tech company singularities", and steering them to reduce x-risk, published by Andrew Critch on May 13, 2022 on LessWrong. The purpose of this post (also available on the EA Forum) is to share an alternative notion of “singularity” that I’ve found useful in timelining/forecasting. A fully general tech company is a technology company with the ability to become a world-leader in essentially any industry sector, given the choice to do so — in the form of agreement among its Board and CEO — with around one year of effort following the choice. Notice here that I’m focusing on a company’s ability to do anything another company can do, rather than an AI system's ability to do anything a human can do. Here, I’m also focusing on what the company can do if it chooses rather than what it actually ends up choosing to do. If a company has these capabilities and chooses not to use them — for example, to avoid heavy regulatory scrutiny or risks to public health and safety — it still qualifies as a fully general tech company. This notion can be contrasted with the following: Artificial general intelligence (AGI) refers to cognitive capabilities fully generalizing those of humans. An autonomous AGI (AAGI) is an autonomous artificial agent with the ability to do essentially anything a human can do, given the choice to do so — in the form of an autonomously/internally determined directive — and an amount of time less than or equal to that needed by a human. Now, consider the following two types of phase changes in tech progress: A tech company singularity is a transition of a technology company into a fully general tech company. This could be enabled by safe AGI (almost certainly not AAGI, which is unsafe), or it could be prevented by unsafe AGI destroying the company or the world. An AI singularity is a transition from having merely narrow AI technology to having AGI technology. I think the tech company singularity concept, or some variant of it, is important for societal planning, and I’ve written predictions about it before, here: 2021-07-21 — prediction that a tech company singularity will occur between 2030 and 2035 2022-04-11 — updated prediction that a tech company singularity will occur between 2027 and 2033. A tech company singularity as a point of coordination and leverage The reason I like this concept is that it gives an important point of coordination and leverage that is not AGI, but which interacts in important ways with AGI. Observe that a tech company singularity could arrive before AGI, and could play a role in preventing AAGI, e.g., through supporting and enabling regulation; enabling AGI but not AAGI, such as if tech companies remain focussed on providing useful/controllable products (e.g., PaLM, DALL-E); enabling AAGI, such as if tech companies allow experiments training agents to fight and outthink each other to survive. after a tech company singularity, such as if the tech company develops safe AGI, but not AAGI (which is hard to control, doesn't enable the tech company to do stuff, and might just destroy it). Points (1a) and (1b) are, I think, humanity’s best chance for survival. Moreover, I think there is some chance that the first tech company singularity could come before the first AI singularity, if tech companies remain sufficiently oriented on building systems that are intended to be useful/usable, rather than systems intended to be flashy/scary. How to steer tech company singularities? The above suggests an intervention point for reducing existential risk: convincing a mix of scientists regulators investors, and the public . to shame tech companies for building useless/flashy systems (e.g., autonomous agents trained in evolution-like environments to exhibit survival-oriented intelligence), so they remain focussed on building usable/useful systems (e.g., D...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: What DALL-E 2 can and cannot do, published by Swimmer963 on May 1, 2022 on LessWrong. I got access to DALL-E 2 earlier this week, and have spent the last few days (probably adding up to dozens of hours) playing with it, with the goal of mapping out its performance in various areas – and, of course, ending up with some epic art. Below, I've compiled a list of observations made about DALL-E, along with examples. If you want to request art of a particular scene, or to test see what a particular prompt does, feel free to comment with your requests. DALL-E's strengths Stock photography content It's stunning at creating photorealistic content for anything that (this is my guess, at least) has a broad repertoire of online stock images – which is perhaps less interesting because if I wanted a stock photo of (rolls dice) a polar bear, Google Images already has me covered. DALL-E performs somewhat better at discrete objects and close-up photographs than at larger scenes, but it can do photographs of city skylines, or National Geographic-style nature scenes, tolerably well (just don't look too closely at the textures or detailing.) Some highlights: Clothing design: DALL-E has a reasonable if not perfect understanding of clothing styles, and especially for women's clothes and with the stylistic guidance of "displayed on a store mannequin" or "modeling photoshoot" etc, it can produce some gorgeous and creative outfits. It does especially plausible-looking wedding dresses – maybe because wedding dresses are especially consistent in aesthetic, and online photos of them are likely to be high quality? Close-ups of cute animals. DALL-E can pull off scenes with several elements, and often produce something that I would buy was a real photo if I scrolled past it on Tumblr. Close-ups of food. These can be a little more uncanny valley – and I don't know what's up with the apparent boiled eggs in there – but DALL-E absolutely has the plating style for high-end restaurants down. Jewelry. DALL-E doesn't always follow the instructions of the prompt exactly (it seems to be randomizing whether the big pendant is amber or amethyst) but the details are generally convincing and the results are almost always really pretty. Pop culture and media DALL-E "recognizes" a wide range of pop culture references, particularly for visual media (it's very solid on Disney princesses) or for literary works with film adaptations like Tolkien's LOTR. For almost all media that it recognizes at all, it can convert it in almost-arbitrary art styles. [Tip: I find I get more reliably high-quality images from the prompt "X, screenshots from the Miyazaki anime movie" than just "in the style of anime", I suspect because Miyazaki has a consistent style, whereas anime more broadly is probably pulling in a lot of poorer-quality anime art.] Art style transfer Some of most impressively high-quality output involves specific artistic styles. DALL-E can do charcoal or pencil sketches, paintings in the style of various famous artists, and some weirder stuff like "medieval illuminated manuscripts". IMO it performs especially well with art styles like "impressionist watercolor painting" or "pencil sketch", that are a little more forgiving around imperfections in the details. Creative digital art DALL-E can (with the right prompts and some cherrypicking) pull off some absolutely gorgeous fantasy-esque art pieces. Some examples: The output when putting in more abstract prompts (I've run a lot of "[song lyric or poetry line], digital art" requests) is hit-or-miss, but with patience and some trial and error, it can pull out some absolutely stunning – or deeply hilarious – artistic depictions of poetry or abstract concepts. I kind of like using it in this way because of the sheer variety; I never know where it's going to go with a prompt...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Salvage Epistemology, published by jimrandomh on April 30, 2022 on LessWrong. A funny thing happens with woo sometimes, in the rationality community. There's a frame that says: this is a mix of figurative stuff and dumb stuff, let's try to figure out what the figurative stuff is pointing at and salvage it. Let's call this "salvage epistemology". Unambiguous examples include the rationality community's engagement with religions, cold-reading professions like psychics, bodywork, and chaos magic. Ambiguous examples include intensive meditation, Circling, and many uses of psychedelics. The salvage epistemology frame got locally popular in parts of the rationality community for awhile. And this is a basically fine thing to do, in a context where you have hyper-analytical programmers who are not at risk of buying into the crazy, but who do need a lens that will weaken their perceptual filters around social dynamics, body language, and muscle tension. But there's a bad thing happens when you have a group that are culturally adjacent to the hyper-analytical programmers, but who aren't that sort of person themselves. They can't, or shouldn't, take for granted that they're not at risk of falling into the crazy. For them, salvage epistemology disarms an important piece of their immune system. I think salvage epistemology is infohazardous to a subset of people, and we should use it less, disclaim it more, and be careful to notice when it's leading people in over their heads. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Increasing Demandingness in EA, published by jefftk on April 29, 2022 on LessWrong. In thinking about what it means to lead a good life, people often struggle with the question of how much is enough: how much does our morality demand of us? People have given a wide range of answers to this question, but effective altruism has historically used "giving 10%". Yes, it's better if you donate a larger fraction, switch to a job where you can earn more, or put your career to use directly, but if you're giving 10% to effective charity you're doing your share, you've met the bar to consider yourself an EA, and we're happy to have you on board. I say "historically", because it feels like this is changing; I think EAs would generally still agree with my paragraph above, but while in 2014 it would have been uncontroversial now I think some would disagree and others would have to think for a while. EA started out as a funding-constrained movement. Whether you looked at global poverty, existential risk, animal advocacy, or movement building, many excellent people were working as volunteers or well below what they could earn because there just wasn't the money to offer competitive pay. Every year GiveWell's total room for more funding was a multiple of their money moved. In this environment, the importance of donations was clear. EA has been pretty successful in raising money, however, and the primary constraint has shifted from money to people. In 2015, 80k made a strong case for focusing on what people can do directly, not mediated by donations, and this case is even stronger today. Personally, I've found this pretty convincing, though in 2017 I decided to return to earning to give because it still seemed like the best fit for me. What this means, however, is that we are now trying to build a different sort of movement than we were ten years ago. While people who've dedicated their careers toward the most critical things have made up the core of the movement all along, the ratio of impact has changed. Imagine you have a group of people donating 10% to the typical mix of EA causes. You are given the option to convince one of them to start working on one of 80k's priority areas, but in doing so N others will get discouraged and stop donating. This is a bit of a false dilemma, since ideally these would not be in conflict, but let's stick with this for a bit because I think it is illustrative. In 2012 I would have put a pretty low number for N, perhaps ~3, partly because we were low on money, but also because we were starting a movement. In 2015 I would have put N at ~30: a factor of 6 because of the difference between 10% and the most that people in typical earning to give roles can generally donate (~60%) and a factor of 5 because of the considerations in Why you should focus more on talent gaps, not funding gaps. With the large recent increases in EA-influenced spending I'd roughly put N at ~300 [1], though I'd be interested in better estimates. Unfortunately, a norm of "10% and you're doing your part" combines very poorly with the reality of 100% of someone's career having ~300x more impact than 10%. This makes EA feel much more demanding than it used to: instead of saying "look at the impact you can have by donating 10%", we're now generally saying "look at the impact you can have by building your entire career around work on an important problem." (This has not applied evenly. People who were already planning to make EA central to their career are generally experiencing EA as less demanding: pay in EA organizations has gone up, there is less stress around fundraising, and there is less of a focus on frugality or other forms of personal sacrifice. In some cases these changes mean that if someone does decide to shift their career it is less of a sacrifice than it would've been, t...
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Jaan Tallinn's 2021 Philanthropy Overview, published by jaan on April 28, 2022 on LessWrong. to follow up my philantropic pledge from 2020, i've updated my philanthropy page with 2021 results. in 2021 i made $22M worth of endpoint grants — exceeding my commitment of $14.4M (20k times $718.11 — the minimum price of ETH in 2021). notes: this number includes $1.9M to orgs that do re-granting (LTFF, EAIF, impetus grants, and PPF) — so it's likely that some of that $1.9M should not be included in the "endpoint grants in 2021" total. regardless, i'm comfortably above my commitment level for that not to matter; i have an ongoing substantial charitable project that's not reflected in the 2021 numbers — it's possible (and likely if ETH price holds, as SFF's s-process alone can't handle such amount) that i will report it retroactively next year or in 2024. Thanks for listening. To help us out with The Nonlinear Library or to learn more, please visit nonlinear.org.
Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Lies Told To Children, published by Eliezer Yudkowsky on April 14, 2022 on LessWrong. Growing up, as a kid, I was always told that every sapient life is precious, everything that thinks and knows itself - Yes, this is a story about lies-told-to-children. You'll probably figure it out yourself before too long. For now, just listen. Where was I? Right. As children, we were always told that every sapient life is precious. It was told to us by the teachers, and shown to us in children's television - though I saw less children's television than most children in our age cohort - children's TV was censored where I grew up, though, of course, I didn't find that out until much later - I see you're starting to guess under what sort of circumstances I grew up. Go ahead, write down the prediction if you want. Maybe you already see where this entire thing is headed. But you asked me for a story about the lies I was told as a child, and that's what you're getting. It's not my fault, if a lot of stories like that are predictable; people who lie to children have other things to optimize for than unpredictability. So where was I? Right. I grew up in a remote village of about three thousand people, the sort that's more hills than houses. Charming travel-pathways that cut through forests. Not everyone knows everyone, but you sure know somebody who knows anybody. Children's television in my region was censored, though of course they didn't tell us that as children. But the children's television that we saw had aliens and monsters and creatures of fantasy, with four legs or fourteen legs, three faces or no face at all, and all of them were treated by the television show as having lives that meant something. Sometimes in the children's show there were alien monsters who only thought their own kind of life was valuable, and then maybe you couldn't trade with them as friends, maybe they'd already lied to you once and you couldn't trust them enough to bargain with them, maybe you couldn't talk to them at all. But their lives still had meaning to the story's human protagonists, even some aliens whose lives had no meaning to themselves. You didn't cause them pain if there was any way to avoid it; you didn't kill them unless their biology was sufficiently similar to human that you were confident in your ability to cryopreserve them afterwards. The shows never spelled it out, never said, 'And this is because of a universal rule in every case that sapient life has value.' Our teachers said that explicitly, though. And they treated every one of us children, too, as if our lives had meaning. Except the children with the red hair; those dirty reds. You're nodding along with a knowing look, I see. Was it what you predicted? Not exactly, maybe, but rough ballpark? I suppose I'll find out when we open your prediction afterwards. The red-haired children hardly needed the red hair, as their targeting-mark; they looked different from the rest of us in other ways too. When I was old enough to first ask, I was told that they were the children's children of people who'd been exiled from a faraway city for committing terrible crimes there, who'd been given sanctuary by the grace and mercy of our own benevolent kind. The red-haired children tended bigger than the rest of us, with more adult facial structures, to the point where you could've maybe mistaken them for very small adults in disguise. The red-haired adults, what few of them we ever saw, were correspondingly huge and muscular. You could see, in retrospect - if you were actually trying to think at all, which we weren't really - how somebody might have felt threatened by such big muscular people, even while graciously granting them sanctuary. There weren't many of the red-haired children being educated alongside us; a handful, four or six. I can't recal...