Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Paper-Reading for Gears , published by johnswentworth on the AI Alignment Forum. Lesswrong has a fair bit of advice on how to evaluate the claims made in scientific papers. Most of this advice seems to focus on a single-shot use case - e.g. a paper claims that taking hydroxyhypotheticol reduces the risk of malignant examplitis, and we want to know how much confidence to put on the claim. It’s very black-box-y: there’s a claim that if you put X (hydroxyhypotheticol) into the black box (a human/mouse) then Y (reduced malignant examplitis) will come out. Most of the advice I see on evaluating such claims is focused around statistics, incentives, and replication - good general-purpose epistemic tools which can be applied to black-box questions. But for me, this black-box-y use case doesn’t really reflect what I’m usually looking for when I read scientific papers. My goal is usually not to evaluate a single black-box claim in isolation, but rather to build a gears-level model of the system in question. I care about whether hydroxyhypotheticol reduces malignant examplitis only to the extent that it might tell me something about the internal workings of the system. I’m not here to get a quick win by noticing an underutilized dietary supplement; I’m here for the long game, and that means making the investment to understand the system. With that in mind, this post contains a handful of thoughts on building gears-level models from papers. Of course, general-purpose epistemic tools (statistics, incentives, etc) are still relevant - a study which is simply wrong is unlikely to be much use for anything. So the thoughts and advice below all assume general-purpose epistemic hygiene as a baseline - they are things which seem more/less important when building gears-level models, relative to their importance for black-box claims. I’m also curious to hear other peoples’ thoughts/advice on paper reading specifically to build gears-level models. Get Away From the Goal Ultimately, we want a magic bullet to cure examplitis. But the closer a paper is to that goal, the stronger publication bias and other memetic distortions will be. A flashy, exciting result picked up by journalists will get a lot more eyeballs than a failed replication attempt. But what about a study examining the details of the interaction between FOXO, SIRT6, and WNT-family signalling molecules? That paper will not ever make the news circuit - laypeople have no idea what those molecules are or why they’re interesting. There isn’t really a “negative result” in that kind of study - there’s just an open question: “do these things interact, and how?”. Any result is interesting and likely to be published, even though you won’t hear about it on CNN. In general, as we move more toward boring internal gear details that the outside world doesn’t really care about, we don’t need to worry as much about incentives - or at least not the same kinds of incentives. Zombie Theories Few people want to start a fight with others in their field, even when those others are wrong. There is little incentive to falsify the theory of somebody who may review your future papers or show up to your talk at a conference. It’s much easier to say “examplitis is a complex multifactorial disease and all these different lines of research are valuable and important, kumbayah”. The result is zombie theories: theories which are pretty obviously false if you spend an hour looking at the available evidence, but which are still repeated in background sections and review articles. One particularly egregious example I’ve seen is the idea that a shift in the collagen:elastin ratio is (at least partially) responsible for the increased stiffness of blood vessels in old age. You can find this theory in review articles and even textbooks. It’s a nice theory: new elastin...