Link to original article
Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: How I think about alignment, published by Linda Linsefors on August 13, 2022 on The AI Alignment Forum. This was written as part of the first Refine blog post day. Thanks for comments by Chin Ze Shen, Tamsin Leake, Paul Bricman, Adam Shimi. Magic agentic fluid/force Somewhere in my brain there is some sort of physical encoding of my values. This encoding could be spread out over the entire brain, it could be implicit somehow. I’m not making any claim of how values are implemented in a brain, just that the information is somehow in there. Somewhere in the future a super intelligent AI is going to do some action. If we solve alignment, then there will be some causal link between the values in my head (or some human head) and the action of that AI. In some way, whatever the AI does, it should do it because that is what we want. This is not purely about information. Technically I have some non-zero causal influence over everything in my future lightcone, but most of this influence is too small to matter. More relevant, but still not the thing we want is deceptive AI. In this case the AI’s action is guided by our values in a non-neglectable way, but not in the way we want. I have a placeholder concept which I call magic agentic fluid or magic agentic force. (I’m using “magic” in the traditional rationalist way of tagging that I don’t yet have a good model of how this works.) MAF is a high-level abstraction. I expect it to not be there when you zoom in too much. Same as how solid objects just exist at a macro scale, and if you zoom in too much there are just atoms. There is no essence, only shared properties. I think the same way about agents and agency. MAF is also in some way like energy. In physics energy is a clearly defined concept, but you can’t have just energy. There is no energy without a medium. Same as how you can’t have speed without there being anything that is moving. This is not a perfect analogy. But I think that this concept does point to something real and/or useful, and that it would be valuable to try to get a better grasp on what it is, and what it is made of. Let’s say I want to eat an apple, and later I am eating an apple. There is some causal chain that we can follow from me wanting the apple to me eating the apple. There is a chain of events propagating through spacetime, and it carries with it my will and it enacts it. This chain is the MAF. I would very much like to understand the full chain of how my values are causing my actions. If you have any relevant information, please let me know. Trial and error as an incomplete example of MAF I want an apple -> I try to get an apple and if it doesn't work, I try something else until I have an apple -> I have an apple This is an example of MAF, because wanting an apple caused me to have an apple through some causal chain. But also notice that there are steps missing. How was “I want an apple” transformed into “I try to get an apple and if it doesn't work, I try something else until I have an apple”? Why not the alternative plan “I cry until my parents figure out what I want and provide me with an apple”. Also, the second step involves generating actual things to try. When humans execute trial and error, we don’t just do random things, we do something smarter. This is not just about efficiency. If I try to get an apple by trying random actions sequences, with no learning other than “that exact sequence did not work”, then I’ll never get an apple, because I’ll die first. I could generate a more complete example, involving neural nets or something, and that would be useful. But I’ll stop here for now. Mapping the territory along the path When we have solved Alignment, there will be a causal chain, carrying MAF from somewhere inside my brain, all the way to the actions of the AI. The MAF has to survive passa...