What's New Under the Sun: Recent Episodes

Matt Clancy

Research on Innovation

mattsclancy.substack.com

View Details

A frequent worry is that our scientific institutions are risk-averse and shy away from funding transformative research projects that are high risk, in favor of relatively safe and incremental science. Why might that be?

Let’s start with the assumption that high-risk, high-reward research proposals are polarizing: some people love them, some hate them. If this is true, and if our scientific institutions pay closer attention to bad reviews than good reviews, then that could be a driver of risk aversion. In this podcast, I look at three channels through which negative assessments may have outsized weight in decision-making, and how this might bias science away from transformative research.This podcast is an audio read through of the (initial version of the) article Biases Against Risky Research, originally published on New Things Under the Sun.Articles mentionedGross, Kevin, and Carl T. Bergstrom. 2021. Why ex post peer review encourages high-risk research while ex ante review discourages it. PNAS 118(51) e2111615118. https://doi.org/10.1073/pnas.2111615118

Krieger, Joshua, and Ramana Nanda. 2022. Are Transformational Ideas Harder to Fund? Resource Allocation to R&D Projects at a Global Pharmaceutical Firm. Harvard Business School Working Paper 21-014.

Jerrim, John, and Robert Vries. 2020. Are peer reviews of grant proposals reliable? An analysis of Economic and Social Research Council (ESRC) funding applications. The Social Science Journal 60(1): 91-109. https://doi.org/10.1080/03623319.2020.1728506

Lane, Jacqueline N., Misha Teplitskiy, Gary Gray, Harder Ranu, Michael Menietti, Eva C. Guinan, and Karim R. Lakhani. 2022. Conservatism Gets Funded? A Field Experiment on the Role of Negative Information in Novel Project Evaluation. Management Science 68(6): 3975-4753. https://doi.org/10.1287/mnsc.2021.4107

This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit mattsclancy.substack.com

View Details

This is an audio read-through of the initial version of How Common is Independent Discovery?

Like the rest of New Things Under the Sun, the underlying article upon which this audio recording is based will be updated as the state of the academic literature evolves; you can read the latest version here.

This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit mattsclancy.substack.com

View Details

This is an audio read-through of the initial version of Science is getting harder.

Like the rest of New Things Under the Sun, the underlying article upon which this audio recording is based will be updated as the state of the academic literature evolves; you can read the latest version here.

This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit mattsclancy.substack.com

View Details

This is an audio read-through of the initial version of When Extreme Necessity is the Mother of Invention. To read the initial newsletter text version of this piece, click here.

Like the rest of New Things Under the Sun, this underlying article upon which this audio recording is based will be updated as the state of the academic literature evolves; you can read the latest version here.

This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit mattsclancy.substack.com

View Details

This is an audio read-through of the initial version of Steering Science with Prizes. To read the initial newsletter text version of this piece, click here.

Like the rest of New Things Under the Sun, this underlying article upon which this audio recording is based will be updated as the state of the academic literature evolves; you can read the latest version here.

This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit mattsclancy.substack.com

View Details

This is an audio read-through of the initial version of Progress in Programming as Evolution. To read the initial newsletter text version of this piece, click here.

Like the rest of New Things Under the Sun, this underlying article upon which this audio recording is based will be updated as the state of the academic literature evolves; you can read the latest version here.

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

Note: An audio version of New Things Under the Sun is now available on all major podcast platforms. Apple, Spotify, Google, Amazon, Stitcher

Think of new technologies as proceeding through a set of stages:

Basic scientific research that explores phenomena

Applied research to better understand how to harness certain phenomena

Technology development to capture and orchestrate phenomena for a purpose

Marketing and diffusion of the new technology

The real world can be more complicated with back-and-forth interplay between the stages, but this is a fine place to start. If you want to shape the direction of technology, you can intervene early in this process and try to push the kinds of technology you want onto the market, by subsidizing research. Or you can intervene at the end of the process and try to pull the kinds of technology you want into existence by shaping how markets will receive different kinds of technology.

One specific context where we have some really nice evidence about the efficacy of pull policies is the automobile market. Making fuel more expensive or just flat out mandating carmakers meet certain emissions standards seems to pretty reliably nudge automakers into developing cleaner and more fuel efficient vehicles. We’ve got two complementary lines of evidence here: patents and measures of progress in fuel economy. In this post, first I’ll go over the evidence, and then I’ll talk a bit what I think we should take away from it. In my view, we have strong evidence that pull policies work well for incremental progress, but the case for their efficacy at promoting radical innovation is a lot shakier.

Patents, Progress, and Pull Policies

One pull policy is a tax on undesirable technologies, since, all else equal, that makes the taxed technology less profitable to develop and alternatives more profitable to develop. A carbon tax is the most famous example of this kind of policy. Aghion et al. (2016) are interested in how a carbon tax might change innovation in the auto sector, but given how rarely anyone actually tries to implement a carbon tax, they can’t directly study the question. Instead, they do the next best thing. From the perspective of a carmaker, one of the main effects of a carbon tax is to raise the price of fuel. So how do carmakers respond to higher fuel prices?

To answer that question, they need a way to measure innovation. More specifically, they want to measure the kind of innovation a firm decides to do: do carmakers focus on clean technology (electric, hydrogen fuel cell, and hybrid vehicle) or conventional fossil fuel innovation? Patents are a useful dataset for this kind of problem, since they are correlated with inventive effort and can be easily categorized into different kinds of technology. One problem though is that patents vary tremendously in how valuable they are. Some are very important, but a lot are junk. So Aghion and coauthors focus on the subset of patents for which the patent-holder sought patent protection in three big markets: the USA, the European Union, and Japan. Since it’s costly to apply for a patent in each market, doing so in all these markets is a signal that the inventor thinks the patented invention is sufficiently valuable to be worth protecting in multiple large markets. So in this paper, you can think of them measuring innovation by counting the number of valuable patents for different kinds of automobile technology.

In an ideal setting, if we really wanted to assess the impact of fuel prices on the innovation decisions of carmakers, we would want to randomly assign some carmakers to face higher fuel prices than others. Then we could compare the subsequent patenting behavior across groups facing different fuel prices. And if we wanted to establish this robustly, we would want to do this kind of experiment many times.

We can’t do that. But Aghion and coauthors do something that gets you closer to this ideal. The price of fuel varies a lot from year-to-year, thanks to fluctuations in the price of oil (see left figure below), but it also varies a lot from country-to-country, because countries vary substantially in the size of their taxes on fuel (see right figure below).

Moreover, carmakers typically sell to multiple countries, but have different footprints in different countries. Aghion and coauthors reason that carmakers are more sensitive to (tax-inclusive) fuel prices in the countries where they have a larger share of their total sales. For example, in the figure above it’s clear that the UK raised taxes pretty substantially over the 1990s, while taxes remained flat in the USA. In other words, if we have two carmakers, one with most of its sales in the USA and some in the UK, and another with most of its sales in the UK rather than the USA, then these carmakers are effectively facing different fuel prices. The carmaker selling primarily to the UK sees an increase in effective fuel prices, and we can compare their behavior to the one selling primarily in the USA.

For every carmaker, Aghion and coauthors construct an “effective” fuel price that is specific to that carmaker, by weighting the fuel price in each country by the carmaker’s exposure to that country. They estimate this exposure from the share of patents the carmaker seeks protection for in that country over 1965-1985, because generally you don’t bother seeking patent protection in countries where you don’t plan to operate in the future. They then look to see how the patents of carmakers differ in the subsequent 20 years (1985-2005) as each carmaker faces a different effective fuel price. We can then compare the behavior of firms that were established in markets that went on to have higher fuel prices to the behavior of firms that were established in markets that went on to have lower fuel prices.

(Note that they estimate the markets where a firm is established over 1965-1985, but look at the effect of fuel prices over 1985-2005; this prevents their results from being driven by innovative fuel efficient companies entering markets when they raise taxes, or from their measure of exposure to different markets from being whipped around by subsequent patenting activity)

Aghion and coauthors find fuel prices exert a powerful impact on innovation. A 10% increase in the effective price of fuel (that a specific carmaker is exposed to) is associated with roughly 5% fewer patents of the conventional “fossil fuel” type, and a 10% increase in clean energy patents.

Rozendaal and Vollebergh (2021) adapt this strategy to study the impact of emissions standards on auto innovation (in addition to fuel prices). Emissions standards are essentially a requirement that a carmaker’s average CO2 emissions per mile fall below some target by some date. It’s not actually quite that simple; but that’s the gist of the idea. Just as fuel prices differ across time and space, so to do regulatory standards. And just as a carmaker is more likely to care about the fuel prices in markets where it has a lot of sales, so too is a carmaker more likely to care about the emissions standards of markets where they have a lot of sales.

Rozendaal and Vollebergh make one of those kinds of observations that is obvious in retrospect but which some previous papers apparently missed. In terms of its impact on the rate and direction of innovation, what matters is not whether a country has an emissions target, or even if this target is high or low. What matters is (1) how high or low is this target relative to today’s average emissions and (2) how long do carmakers have to meet the target? If the standard is high, but you’ve already cleared it, then even though it’s high it imposes no extra incentive to innovate. On the other hand, if the standard is high and you have not met it yet, but actually have a long way to go to meet it, then it matters whether you’ve got ten years or one year to get there. If you don’t account for this kind of thing, it might look like standards don’t have much of an impact on innovation.

Rozendaal and Vollebergh construct a measure of standards that takes all this into account: it’s basically the difference between the average emissions of a country and the target, divided by the number of years left to meet the target. When the number is large, it means the average car in the country has a long way to go before it meets the target, and not a lot of time to get there. Here’s how their measure looks for the big three markets.

As with Aghion and coauthors, Rozendaal and Vollebergh construct estimates of each carmaker’s exposure to these three major markets, as of the year 2000. Those heavily exposed to Japan, for example, faced stronger incentives to innovate in the 2000s, compared to those heavily exposed to US and EU markets. But after 2010, the situation has largely reversed. Again, Rozendaal and Vollebergh are going to look at valuable patenting of clean and dirty technology in response to these measures of emission standards stringency (in addition to fuel prices).

And they find these pull policies work. A 10% increase in the stringency standards is associated with 2% more clean patents. Also importantly, Rozendaal and Vollebergh confirm Aghion and coauthor’s finding that fuel prices matter, albeit the strength of the relationship is weaker than Aghion and coauthors find. This might be because they are looking at a very different time frame, 2000-2016 compared to 1985-2005.

What I like about these studies is that they have thousands of different inventive entities, and each of these entities is, at least in principle, facing a “pull policy”of a different strength (based on the mix of markets they operate in). This lets you estimate pretty precisely how effective these policies are. Of course empirical economics is hard and these papers have a few potential weaknesses too. Among these is their reliance on patents as a measure of innovation. So let’s turn to a complementary strand of the literature that doesn’t rely on patents at all. The trade-off will be that we no longer have measures of the rate of innovation for quite as many different organizations facing different policies. Then, after that, I’ll point to a few more reasons for caution about these results.

Directly Measuring Fuel Efficiency

Presumably, the reason policymakers are interested in shaping the direction of progress in automobiles is that they want to encourage innovation that makes cars more fuel efficient. Ideally, we would want cars to be zero emission vehicles. So why not simply measure fuel efficiency directly, if that’s what we care about? We don’t actually care about patents, but results.

The challenge here is that you can improve fuel efficiency without innovating. You can make lighter cars, or cars with smaller engines with less horsepower. In general, cars have a lot of attributes and some of those attributes are associated with a reduced fuel efficiency. So if you want to measure technological progress, what we really need to do is see if cars get more fuel efficient without changing any of these other attributes. And that is precisely what another strand of literature tries to do.

Knittel (2011) has data on the models of practically all consumer vehicle models sold in the USA between 1980 and 2006: approximately 14,000 cars and 13,000 light trucks. For each of these vehicles he has data on their attributes, most importantly weight, horsepower, torque, and fuel efficiency (as measured in miles per gallon). As illustrated in the figure below, there is a tradeoff between each of these attributes and fuel efficiency. In general, models with higher weight, horsepower, or torque, have lower fuel efficiency. But what’s interesting is to compare the shape of this tradeoff at the beginning of the sample (in 1980) to the end (in 2006). For any of these attributes, if you pick a level you want, it tends to be feasible to get much higher fuel efficiency in 2006 than in 1980. The gap between these curves is a fairly direct measure of technological progress.

What Knittel ends up with is year-by-year estimate of the state of US auto fuel efficiency technology; that is, on average, how much more fuel efficient can you make a car, compared to one with the same attributes in 1980. He can then look at how this measure changes over time to assess the rate of technological progress. The results of Aghion et al. (2016) and Rozendaal and Vollebergh (2021), which were based on patents as a measure of innovation, imply that we should see the fastest technological change during periods of elevated US fuel prices and when US fuel economy standards were most stringent (relative to current levels). And that’s basically what Knittel finds - a 10% increase in the fuel price was associated with 0.3% faster technological change, and a 10 percentage point increase in fuel economy standards was associated with one percentage point faster technological progress.

Klier and Linn (2016) builds on this approach to run the sort of larger-scale statistical tests that Aghion et al. (2016) and Rozendaal and Vollebergh (2021) did with patent data. Klier and Linn estimate technological progress over 2000-2012 for US car models and 2005-2010 for Europe, a period during which the US and EU began implementing some new emissions regulations. Unfortunately for the economists, each of these regulations applied to all cars or trucks, which made them less useful for testing their efficacy, since ideally we would like to compare the behavior of firms that face regulations to those that don’t. But as with Rozendaal and Vollebergh, Klier and Linn exploit the fact that the same regulation can impose different burdens on different car companies. Specifically, manufacturers whose models are far from meeting the standards have a lot more work to do than those who are close to meeting the standard (or who already have met them). So Klier and Linn’s idea is to construct manufacturer-specific measures of stringency based on how far each manufacturer needs to tug up the fuel efficiency of its fleet, and then see if we observe faster technological progress among the cars of manufacturers who faced more stringent regulations from their point of view.

And we do. For every standard deviation increase in the stringency standards faced by a firm, technological improvement in US cars improved by an extra 0.5 percentage points per year, US trucks improved by an extra 1.5 percentage points per year, and EU cars improved by an extra 0.3 percentage points. Compared to baseline rates of improvement in the range of 1.4-2.0% per year, this isn’t bad.

So this is overall reassuring. We’re not touching patents at all with these two papers, but we still observe faster technological progress during the periods that the other patent-based papers suggest we should see it. But maybe we’re still worried that we haven’t properly controlled for everything in our statistical models. There were also recessions and sky-high fuel prices during 2000-2012, after all. Klier and Linn attempt to control for this, but it’s challenging. So let’s look at one last paper, Kiso (2019), which tries to exploit an unusually clean setting to test the effect of emission standards.

Kiso’s basic idea is to identify a setting that is very close to an experiment where we have two identical groups, and then we change one variable for one group but not the other. He looks at cars sold in the USA over the period 1985-2004. This was a twenty-year period during which US fuel prices were relatively stable and emissions standards in the US and EU were also fairly stable.

But during this period, Japan instituted emissions standards on vehicles sold there. These policies were announced in January 1993, with a deadline to hit the targets by the year 2000. In March 1999, a new set of targets were established for the year 2010. As indicated below, different vehicle weights had to hit different targets.

During this era, US and EU sales into Japan were tiny as a share of their total sales, whereas Japan was obviously a large and important market for Japanese carmakers. So obviously Japan was more incentivized over this period to improve fuel efficiency than US and EU carmakers, who are facing relatively stable prices and no major new emissions standards in their major markets. But the thing is, Japan is under no obligation to sell the same cars in Japan and the USA (and in fact it does not). So it could have been that Japan sold small fuel efficient cars in the Japanese market, and larger less fuel efficient cars in the US market, where emissions standards were relatively lax.

But the thing about ideas is they have a tendency to spill over into new applications. If you discover ways to make cars more fuel efficient without sacrificing other desirable vehicle characteristics, there is no reason you couldn’t apply those ideas to all the cars in your fleet. So Kiso’s idea is to use the methods of Knittel (2011) and Klier and Linn (2016) to compute the rate of technological progress (in fuel economy) for cars sold in the USA for Japanese carmakers and then to compare that to the rate of progress for cars sold in the USA from US and EU carmakers. If there’s a systematic difference beginning after 1993, when Japan introduced tighter standards, that’s evidence that the standards induced technological progress which spilled over into Japanese carmaker’s models in other countries.

And that is basically what we see. The figure below is the fuel efficiency gap, holding fixed car characteristics, between the US models of Japanese carmakers and the US models of US/EU carmakers. Beginning a few years after the policy announcement in 1993, the cars of Japanese carmakers have persistently better fuel economy than the US and EU cars. I’ve highlighted the three different relevant time periods in the figure below. Note the region highlighted in gray, which I call “out of sample” corresponds to a period when rising fuel prices and tightening emissions standards in the US muddy our interpretation of Japan’s emissions standards (since we now have multiple things changing at the same time).

Taken together, over 1995-2004, fuel economy technology was 2.3-3.3% higher per year for Japanese carmakers (depending on whether you include or exclude the lightest and heaviest cars) as compared to US and EU carmakers.

How far can we extrapolate this?

I think this is pretty compelling. Aghion et al. (2016) tells us fuel prices exert a strong effect on the direction of technological progress, as measured by patents. Five years later, Rozendaal and Vollebergh (2021) use the same approach on a newer slice of the data, showing that effect is still there, albeit weaker. They also show that emissions standards, measured properly, exert the same kind of effect. Then, Knittel (2011) and Klier and Linn (2016) show that measures of technological progress which are based on the actual attributes of vehicles, rather than counting patents, also accelerate during times of high fuel prices or more stringent standards. Finally, Kiso (2019) shows that even when you restrict your attention to the set of cars sold in one market with broadly stable fuel prices and emissions standards, you can detect technological progress speeding forward among the cars of automakers who are, in separate markets, facing pressure to meet higher emissions standards.

At least for cars, pull policies work.

Moreover, we can extend this conclusion beyond auto markets. There is also a slew of studies (some of them briefly mentioned here) that find higher energy prices are associated with more patenting in various renewable energy technologies (like solar, wind, or battery technologies). And in healthcare, a lot of work has also shown that firms increase R&D on diseases that become more profitable to treat. Lastly, we may want to think of Operation Warp Speed, which successfully accelerated the development of highly novel mRNA vaccines by placing huge advance orders for vaccines that had not yet been validated (among other things).

How far should we extrapolate these results? Should we conclude that pull policies usually work? That we should use them for pretty much everything?

I think there are a few things to keep in mind before going that far.

First, one big caveat with all these studies on emissions standards is we can’t quite think of these standards as being just randomly selected. It may be that more ambitious standards are set precisely when new technological opportunities make it feasible to reach those targets. That doesn’t mean the standards don’t work, since technological advances that were possible might not have happened without the standards. But it does mean we need to be cautious about extrapolating from this setting to settings where technology is in a murkier state.

In fact, at least in the context of the EU standards, a 2021 paper by Reynaert provides some evidence that the ambitions of these standards were carefully matched to what experts at the time believed was technologically feasible. For example, he notes:

The EU Commission relied on several studies to support the design of the EU emission standard [....] The policy report only includes possible technology adoptions that should be readily available for the car makers at no fixed or development costs. In designing the regulation, the policymaker clearly had the channel of technology adoption in mind.

I think this suggests the results might be overstated to some degree: to some extent more stringent policies lead to more rapid technological advance, but it’s also true that more stringent policies are also selected when more rapid technological advance is feasible.

Second, and related, recall the simplistic model of technology development I mentioned at the beginning of this piece. New technologies start with basic exploratory science, move into applied research and then technology development, before finally marketing and diffusion. In this other piece, I looked at some evidence that profit incentives worked quite well for the latter stages of health-related R&D, but relatively poorly for initial stages. Specifically, the profit motive’s effect on the development of drugs farther from being market ready appears to be pretty weak.

Can we say the same thing about cars? That this policy is great for speeding up incremental change which is already close to being commercially viable, but poorly suited for more transformative and radical technological change?

As we would expect if the regulations were designed to enable the adoption of already existing technologies , it does seem that technological change in this context was largely incremental. That said, it’s a bit tough to be sure since few papers look specifically at this. However, both Reynaert (2021) and Knittel (2011) discuss the kinds of technological changes that helped bring about higher fuel efficiency. Knittel has the following charts, showing how various new technologies get integrated into more and more cars over time. This is innovation of a sort, but this chart implies a lot of the technology gains came from more and more cars adopting existing technologies that had already been successfully proven elsewhere.

Klier and Linn (2016), as well as Reynaert (2021) doesn’t even call this “innovation” in their papers, preferring to use the phrase “technology adoption” to describe the process of car models becoming more fuel efficient. I tend to think it’s still a form of innovation - the vehicles had to be redesigned to incorporate these changes - but it’s innovation at its most incremental.

We can also turn to some other literature that attempts to describe how much different technologies rely on science for innovation. Ahmadpoor and Jones (2017) is a paper that computes how “close” a technology is to science by looking at how long is the chain of citations connecting an academic article and a patent. Auto-related technologies have pretty long chains by this metric. Patents pertaining to the internal combustion engine, motor vehicles, and electricity transmission to vehicles all having a mean distance to science exceeding 4, meaning on average the shortest link between a patent in one of these fields and a scientific journal is the patent cites a patent that cites a patent that cites a patent that cites a journal article. And a large share of patents never reach a scientific paper at all.

You can see how this stacks up against other technologies in the figure below; the horizontal axis is the length of a citation chain between patents and papers (the further the right, the further from science) and the vertical axis is the share of patents with any linkage to science (the lower, the less connected to science).

Both axes suggest auto technology is historically not dependent on science. And if you don’t like patent citations, this is also indicated by a 1994 survey of lab managers, where just 13% of projects used publicly funded science, one of the lowest shares of any industry surveyed (and as compared to a 20% average).

So although pull policies seem to have worked really well, as a general rule, innovation in auto tech has largely been of the incremental type where pull policies usually work well (at least, according to evidence from health). But there may be one important exception: electric vehicles.

The transition to an all electric vehicle fleet is a a more radical kind of innovation than the kind of incremental changes discussed above. Aghion et al. (2016) and Rozendaal and Vollebergh (2021) are the only papers that address the impact of pull policies on electric vehicles. Both papers include electric vehicle patents among their measures of clean technology, though only Rozendaal and Vollebergh actually split out electric vehicle patents from the rest to see how emissions policies affect them specifically. But when they do, they find the effect of emissions standards on electric vehicles is actually stronger than it is for other technologies. A 10% increase in standards stringency is associated with 3% more electric vehicle patents, compared to an average of 2% across all clean technologies.

One possible answer to this question is that by the time of Rozendaal and Vollebergh’s study, electric vehicle technology was, in fact, commercially viable and so pull policies worked well. Electric vehicles have been around a long time, but what made the technology much more promising in the 21st century (when their study is set) seems to have been the development of a new class of powerful lightweight batteries (lithium ion batteries). The development of these batteries, meanwhile, seems to have been largely a case of a classic knowledge spillover from the consumer electronics sector. As far as I can determine, it wasn’t the case that these batteries were developed by carmakers spurred on by the promise of profits for developing radically more fuel efficient vehicles. They just got lucky; though luck like this happens all the time in innovation, and is why you should generally push forward innovation along lots of dimensions at once.

But in terms of the efficacy of pull policies, I worry that Rozendaal and Vollebergh’s finding that emissions standards worked very well to promote more radical electric vehicle innovation is another example of high standards being set when policymakers believe the technology is already capable of meeting them. Good luck brought the auto sector good batteries, and then after Tesla motors proved you could make desirable cars on the platform, policymakers decided to push the sector to adopt this new technology. But if the sector had not had the luck to get gifted better batteries from consumer electronics, or put another way, if these pull policies had been implemented before these batteries had been developed, maybe we would not see such good results.

But it’s tough to say. I think we have unusually good evidence here that pull policies certainly work for pushing forward incremental innovation. I’m skeptical that they work nearly so well for pushing forward radical technology, but the evidence we have isn’t so great on that question.

Thanks for reading! For other related articles by me, follow the links listed at the bottom of the article’s page on New Things Under the Sun (.com). And to keep up with what’s new on the site, of course subscribe!

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

Note: An audio version of New Things Under the Sun is now available on all major podcast platforms. Apple, Spotify, Google, Amazon, Stitcher

There’s this idea that technology is characterized by path dependency: once you start going down one technology trajectory, you kind of get locked in and it’s hard to switch to another, possibly better trajectory. That can happen for lots of reasons, but one possibility is that it’s something about the nature of knowledge itself. The more you know, the more you can learn: knowledge begets more knowledge. So whichever technology trajectory we start on becomes the one we know the most about, and therefore the one it makes most sense to stick with.

One line of evidence about this comes from dynamics of patenting. You know what’s a pretty good predictor of patent activity in the future? Patent activity in the recent past. In this post, I want to see what we can learn from a literature that directly or indirectly looks at this dynamic. But while I think this line of evidence is useful, I want to be up front that it also has significant limitations. Most importantly, in this literature we almost never get anything like an experiment. Instead, we’re reading the tea leaves in observational data.

Patent Stocks

Specifically, we’re going to look at the conceptual category of a “patent stock” (also frequently called a “knowledge stock”). To illustrate just what a patent stock is, let’s start with a practical implementation of the notion in a paper.

Aghion et al. (2016) is a paper interested in three kinds of technological progress among automobiles: clean tech (think electric cars, hydrogen fuel cells, and hybrid vehicles), fossil fuel technology (think internal combustion engine), and “grey” tech (think more fuel efficient fossil fuel technology). The paper measures innovation in these different technologies by counting valuable patents. There are about 2,500 companies and 1,000 individuals who hold one of these patents, and Aghion and coauthors construct three patent stocks for each of these inventive entities, one for each of these three flavors of technological progress. (Constructing patent stocks is hardly the only thing this paper does, but it’s what I’m focusing on today)

Constructing a patent stock basically means adding up all the patents that an inventive entity has taken out in the past, giving more weight to the more recent patents. To be very explicit, suppose Ford Motor Company had a clean technology patent stock of 1,000 last year and obtains 250 new patents for clean technology this year. Then, to construct the patent stock for current year, we take last years’ patent stock, multiply it by 80%, and then add to that the number of new patents. So this year’s patent stock is 0.8 x 1000 + 250 = 1050. Suppose we get 150 patents next year. Then the patent stock next year is 0.8 x 1050 + 150 = 990. The exact value of 80% isn’t that important and people use different numbers, though always less than or equal to 100%. The key idea is it’s telling us, in a single number, something about the prior patent activity of the Ford Motor Company. Note that last year’s patents are worth less than this year’s patents (in this case, 80% as much). And since we apply the 80% discount each year, patents from two years ago are worth even less (80% x 80% = 64%).

For the data that Aghion et al. (2016) have, for clean technology, a 10% increase in last year’s clean tech patent stock is associated with a 3% increase in this year’s new clean tech patents (for that firm). For fossil fuel technology, the link is even better: if a firm has a 10% higher patent stock in fossil fuels last year, that’s associated with 5% more new fossil fuel patents this year.

This is quite a robust (though not universal) finding. Rozendaal and Vollebergh (2021) look at the same context (clean and dirty innovation in automobiles), but with a more recent slice of data (2000-2016, instead of Aghion and coauthors’ 1986-2005). In their sample, they find an even stronger link: roughly speaking, a 10% increase in last year’s patent stock is associated with a 10% increase in patenting this year. Noailly and Smeets (2015) do something similar for innovation in renewable energy and fossil fuel energy (not automobiles). They also find that firms sought many more clean or fossil fuel patents when last year’s patent stock (of the appropriate type) was higher.

You can also go beyond firms and look at whole countries. Looking specifically at US patents, a famous 2002 paper by David Popp uses a slightly different approach to create patent stocks for 11 different technologies related to energy. He finds that a 10% increase in last year’s patent stock for a particular technology was associated with 7% more patents for that technology this year. And Porter and Stern (2000) compute patent stocks for 17 different OECD countries (looking at patents they seek in the USA, to hold consistent the definition of a patent). Again, a 10% increase in a country’s patent stock last year is associated with 8-11% more patents this year.

This isn’t a universal finding. Lazkano, Nøstbakken, and Pelli (2017) look at innovation in renewable energy, conventional energy, and storage (battery) technology. Unlike the work discussed above, they typically find the opposite result: a higher patent stock for energy storage technology last year is associated with less energy storage patenting today. Similarly for renewable energy. That said, this paper is the outlier (I’ll return to it briefly later). For now, let’s proceed with the understanding that a positive link between yesterday’s patent stock and today’s patenting is a pretty robust correlation. But before jumping from a correlation to a conclusion, we need to think a bit harder about what’s really going on here. Why is there this correlation?

(And if you feel like saying “patents don’t measure innovation!” I hear you, but bear with me for a bit)

What’s Going on Under the Hood?

One potential explanation is quite interesting. Suppose:

Patent stocks measure how much knowledge we have about a technology

When we have more knowledge, it’s easier to discover new things

The new knowledge we discover gets added to our existing knowledge

If this is what’s “really” going on, then it explains why we have a positive link between last year’s patent stock and this year’s patenting. Knowledge begets more knowledge! And this is especially true of more recent knowledge, for which we have not already wrung out all the possible implications (hence the higher weight on recent patents).

If we really believe this story, we can even use the estimated statistical models to make neat little forecasts of how technologies will develop. For every year, we can use last year’s patent stock to predict how many new patents will be discovered. We then use that to predict what next year’s patent stock will be. Rinse and repeat.

There’s also a policy implication. If we can temporarily accelerate the accumulation of knowledge in a field, it can pay huge dividends. That’s because the benefits will compound, since they’ll enable more knowledge creation next year, which will enable more knowledge creation in the following year and so on.

Moreover, this model implies path dependence is a powerful force in technology. If one technology gets a minor head start, it might keep it’s lead for a very long time. In fact, if the relationship is strong enough (a 10% increase in the patent stock increases patenting by 10% or more), then all else equal a technology that is a bit behind can never catch up!

But before we get too far ahead of ourselves, we need to strongly consider some other potentially important explanations.

Let’s jettison the conceptual category “patent stock” for a minute. So far all we have really shown is that patenting in the recent past is correlated with patenting today. To think through why that might be the case, we need to think through what kinds of factors determine the number of patents in a given year (for a firm, technology, or country).

Since I’m an economist, I’m going to divide the potential factors into two categories: supply and demand. Supply factors are anything that affects the cost of getting patents (where cost is broadly construed). Demand factors are anything that affects the value of getting a patent. If it patents get cheaper or more valuable, then we should expect inventors to seek more patents. And if the things that made them cheaper or more valuable persist over time, then patenting in the recent past predicts patenting today.

Yes, it’s true that knowledge is one kind of things that makes it cheaper to discover new things and get patents. But many other non-knowledge factors might also be important. On the supply side, that might include things like more scientists/inventors; more physical capital (computers, laboratories, etc); improved access to financing; more patent lawyers; and so on. On the demand side that might include demand from consumers; new regulations favoring certain kinds of technology; or even just oddball idiosyncratic stuff like demand from a new CEO who thinks the firm should be patenting more of its existing inventions.

If we get any of these factors, we’ll get more patenting today and tomorrow (if the factor sticks around), which will deliver the correlation that patenting today predicts patenting tomorrow. But unlike the interesting “knowledge begets knowledge” theory, this doesn’t necessarily have the same policy implications, nor does it necessarily mean we have strong path dependency in technologies. If these other factors are driving the correlation, then if we increase patenting by hiring more scientists, subsidizing R&D, passing some new regulation, or whatever, the boost in patenting will last as long as the policy, but it won’t necessarily have any further knock-on effects.

So I think it’s worth poking at this correlation to see how much of it can be explained by these various factors. And while our evidence base isn’t fantastic, I think a couple lines of evidence suggest the “knowledge begets knowledge” idea is a big part of the story.

Controlling for Supply and Demand

The most straightforward way to parcel out the drivers of this patent stock/patenting correlation is to try and directly measure the plausible factors that might drive it and adjust for them wherever possible.

Let’s start with supply factors other than knowledge itself. For example, a firm that hires a lot of new R&D workers and invests in new labs might be able to crank out a lot of new (patented) inventions. As long as the firm has this elevated set of inputs, patenting might be elevated too, regardless of what’s happening to knowledge.

There certainly is a correlation between patent stocks and what we might call “inputs to invention.” Park and Park (2006)construct patent stocks for 23 different US manufacturing industries over 1976-1996 and compare them to the number of R&D workers in these sectors over the same time period, finding a correlation coefficient of about 0.3. Park and Park don’t have great measures of the non-labor inputs to research (lab equipment, computers, etc), but they do have total R&D spending, which should encompass that as well as the salaries of R&D workers. The correlation coefficient between patent stocks and R&D spending is higher, at 0.5.

So it’s true that patent stocks are partially but not totally explained by other supply side factors. Indeed, precisely one reason many papers use patent stocks is to proxy for these supply-side innovation factors, since it’s often easier to observe a firm’s patents than it’s employee headcount and R&D spend. But the correlation is a lot less than 1; we’re not explaining 100% of the variation just by looking at these supply factors.

Another strand of evidence comes from a paper that explicitly tries to adjust for the scientific labor force. Porter and Stern (2000) constructs patent stocks for 17 different OECD countries and finds a strong correlation between last year’s patent stock and this year’s patenting, even after adjusting for each country’s R&D workforce. Even within the same country, comparing two years where the scientific workforce is unchanged, if there is a 10% larger patent stock in one of those two years, they find that’s associated with 11% more patents in the following year.

So we might consider that some evidence that this isn’t entirely about supply side factors. What about demand?

We know from a variety of evidence in the life sciences that biomedical R&D is pretty responsive to changing market demand. When a disease becomes more prevalent or more profitable to treat, private sector companies respond by increasing related research. There’s no reason to think biomedical research is special in this regard, and so we should expect demand side drivers to matter for lots of technologies.

A lot of these papers are specifically concerned with the transition from fossil fuels to clean energy. In that context, the relevant demand-side factors are the price of fossil fuels and government regulations that make it more attractive to invent and use renewable energy. And lots of papers have specifically controlled for these demand side factors, and still find this patent stock / patenting correlation:

For 11 different energy technologies, Popp (2002) finds past patent stocks are strongly correlated with present patenting after controlling for energy prices

For renewable energy, Noailly and Smeets (2015) find the same after controlling for energy prices and the size of the energy market

For clean and dirty auto innovation, Aghion et al. (2016) find past patent stocks are strongly correlated with present patenting after controlling for fuel prices and some proxies of government regulation

Rozendaal and Vollebergh (2021) find the same using more recent data, fuel prices, and an improved measure of the stringency of government regulation

These measures of demand-side factors are always going to be incomplete and imperfect, but they seem like plausible candidates for the first-order drivers of demand for different kinds of energy-related innovation. So that should tilt us a bit to the side of thinking the correlation between a patent-stock and subsequent patenting isn’t just about unobserved demand.

So where do we end up? In this article, we’re interested in how much the correlation between patent stocks and subsequent patenting is driven by knowledge generating more knowledge. We’ve seen this correlation persists when we try to control for some various confounders, like the number of scientists or changing demand for different kinds of innovation. But it’s always possible we’ve missed something. Moreover, it’s a bit frustrating that we have one set of paper that looked at supply and another set that looked at demand. It might be that if we could control for both supply and demand at once, the patent stock / patenting correlation would be much smaller, or even disappear. So to close, let’s turn to a few papers that try to look more directly for positive evidence on knowledge itself.

Measuring Knowledge, not just Patents

To begin, there are a few papers that have attempted to show patent stocks really capture something like “technological know-how”. As I noted at the beginning, a potential critique up to now might have been, “patents don’t measure innovation!” For sure lots of patented stuff is junk and lots of brilliant stuff is not patented. But I think there’s signal in the noise.

For example, Park and Park (2006) and Porter and Stern (2000) both look to see how well patent stocks predict not just patenting but also “total factor productivity.” Total factor productivity - hereafter TFP - is a common (albeit highly imperfect) measure of technological capability. The basic idea behind it is if you can squeeze more economic outputs out of the same number of inputs (for example, capital and labor), then that’s a reflection of better technology. The challenge is you can’t measure TFP directly; instead, you directly measure outputs and inputs and try to predict the outputs with the inputs. Something like the gap in your predictions is your measure of technology. It’s a nice concept because it’s a general purpose measure of technology, but it’s very noisy and can be affected by things other than technological progress.

Nonetheless, it’s one of the few broad measures of technology we have. Park and Park (2006) look at how correlated the growth of TFP in an industry is with its patent stock. The correlation is there, though a bit weak; a 10% increase in the patent stock is associated with a 1.5% increase in the TFP growth rate. And Porter and Stern find similar results at the level of the country; a 10% increase in the national patent stock is associated with a 0.5% increase in the national TFP growth rate. To the extent patent stocks are correlated with a completely different measure of technological know-how (that is computed with no reference to patents), that’s some evidence patent stocks are picking up information about technological know-how. That know-how apparently affects economic performance, and so it’s not such a leap to imagine it also affects subsequent technological progress. To the extent the link is weak though, that’s either evidence that the signal about knowledge from patent stocks is pretty small, or that TFP is a crappy measure of technology, or - most likely, in my opinion - a bit of both.

Finally, a few papers try to go beyond merely counting patents, to get at some proxies for the knowledge associated with patented inventions. To the extent these papers actually do a better job of measuring knowledge (which is debatable), they tend to show more knowledge is associated with more patenting.

One of these is the previously mentioned Popp (2002). Popp attempts to use patent citations to measure the value of the knowledge created by patents. The basic idea is that if patents in a particular year have an unusually high probability of being cited, that suggests the patents in that year are particularly valuable as sources of knowledge. And if your knowledge stock contains lots of these high-knowledge patents, then it should generate more patents than another patent stock with the same number of patents but fewer of the high-knowledge type.

This is a tricky exercise; patent citations are a pretty noisy indicator of knowledge flows, though I tend to think they have some signal. Moreover, you need to be careful not to get the causality backwards, so that a surge of patents in later years causes more citations to patents in earlier years, creating this spurious correlation between the number of highly cited patents in a year and subsequent patent activity. Popp avoids this by using the probability a patent gets cited, which takes into account the number of patents that might cite it.

In any event, Popp shows yesterday’s value of this more sophisticated knowledge stock is more tightly correlated with today’s patenting than the standard knowledge stocks that make no such allowances.

Another paper that tries to measure the knowledge in patent stocks is one of my own, Clancy (2017). This paper is based on the idea that innovation is a combinatorial process - new ideas are all about finding useful combinations of older pre-existing ideas. The underlying intuition is that it’s not the number of patents in a technology field that matters, but the number of useful combinations of ideas that the field knows about.

The challenge is finding sensible proxies for these concepts, and then to test them correctly. Fortunately, the US patent and Trademark Office (USPTO) has developed a highly granular system for classifying the technologies described in patents. When I was doing the research, the USPTO had 450 different major technological “classes” and about 15,000 “mainline subclasses” which provided a more detailed description of technology. Importantly, most patents are assigned more than one of these mainline subclasses, which is one indication that two distinct “ideas” (each of which is described by one of these subclasses) have been successfully combined in the single patented invention.

Clancy (2017) defines two different measures of technological know-how for each of the 450 broad technology “classes.” The first of these is just a count of the number of patents assigned to the class, which is basically the same as the patent stocks we’ve been discussing so far. The second is a measure based on the number of pairs of mainline subclasses that are assigned to patents in that class, adjusted in a way to try and capture inventors’ depth of experience with combining the mainline subclasses. The paper then uses both of these measures to predict subsequent patenting.

Here’s an example to illustrate the basic idea. Suppose we have two different technologies: solar energy and fusion energy. Suppose they each belong to a different technology class. Pretend each has roughly the same number of patents associated with it, so that they have similar patent stocks. But suppose solar energy patents combine dozens of different technological categories and the fusion patents only combine a few different categories again and again. That might imply it would be easier to come up with new solar energy patents than fusion patents, because we only know how to do a few different kinds of things with fusion, and those things have been largely explored, while with solar we have dozens of avenues left to explore.

And that’s basically what Clancy (2017) finds. Within a given technological class, a 10% jump in this measure of “knowledge about how to combine pairs of technology” is associated with a subsequent 9% increase in patenting in that field. But now a 10% jump in the number of patents - holding fixed our measure of how much the field knows about how to combine pairs of technologies - is associated with a 30% drop in patenting. The notion here is that new patents can be good for future innovation (if they show how to do new things like combine rarely paired technologies) or bad for future innovation (if they merely exhaust possibilities that are already known). Perhaps this kind of argument can help explain why we occasionally find a higher patent stock implies less patenting, such as in Lazkano, Nøstbakken, and Pelli (2017)?

(As far as I know, Popp (2002) and Clancy (2017) are some of the only papers that adopt the patent stock approach, but experiment with developing improved measures of the “knowledge” in them. I suspect you could now do a lot more now to develop measures of knowledge in a field, for example, by using the text in patents to measure the diversity of ideas in a field, or looking at how much the field is citing new scientific research versus old research.)

The Upshot

So, to step back and appraise the situation, we’ve got this correlation that’s pretty general: yesterday’s patenting tends to predict today’s. One interesting hypothesis to explain this stylized fact is that “knowledge begets knowledge;” a policy that pumps up knowledge production today could have a long echo into the future, as new knowledge enables yet more discovery. Indeed, if the strength of this relationship is strong enough, so that a 10% increase in knowledge leads to at least a 10% increase in the production of new knowledge, then you obtain strong path dependency effects. Once you get a technological paradigm started, such as fossil fuel derived energy, innovation is self-propelling and it becomes very hard for rival paradigms to ever catch up (without active assistance).

But we should be careful about leaping to that interpretation. We don’t have anything like a nice experiment or even quasi-experiment here; just messy correlations in observational data. In particular, it’s quite likely part of this correlation is driven by factors that don’t necessarily have this self-propelling character. It might be that the correlation is driven by the fact that fields that do a lot of patenting in the recent past have a lot of scientists working in them, and so we should expect those scientists to generate a lot of patents in the future as well. But it looks like this isn’t the whole story.

At the same time, it might be that the correlation is driven by the fact that technologies that are in high demand yesterday attract a lot of inventive effort yesterday, and so long as that demand remains elevated then they will continue to attract inventive effort today. But again, at least for energy innovation, this patent stock to patenting correlation seem to exist even after taking into account those demand-side factors.

Lastly, when we do try to drill down and measure knowledge more directly, we do find some evidence consistent with this view that the knowledge is a big part of what matters. Patenting is correlated with some other measures we have of technological know-how, and improved measures of knowledge seem to predict future patenting better than unimproved measures.

Despite all that, nature is tricky and it could well be that all the confounders I’ve already mentioned, plus things we’ve overlooked, drives this correlation. But lastly, we have a few additional suggestive lines of evidence that knowledge matters for innovation, which are discussed at length elsewhere on New Things Under the Sun. For example, morescientific knowledge also seems to facilitate innovation, and knowledge spillovers seem quite important for innovation. That suggests we should approach this line of evidence already pretty open to the notion that we’ll find knowledge is an important explanatory factor.

So here is my main take-away from all this. Some part of the correlation between yesterday’s patent stock and today’s patenting is probably driven by knowledge itself, but not all of it. At the same time, we rarely find a 10% increase in yesterday’s patent stock is associated with a greater than 10% increase in today’s patenting, and when we do, it’s usually not much more than 10%. More commonly, we find a 10% increase in yesterday’s patent stock is associated with a positive but less than 10% increase in today’s patenting. Since I think knowledge probably only accounts for part of this correlation, I suspect, on average a 10% increase in knowledge (if we could measure knowledge correctly) leads to a lessthan 10% increase in knowledge discovery.

In other words, knowledge begets more knowledge, but with diminishing returns. Innovation tends to get harder. Policies that pump up knowledge do echo into the future; but they don’t echo forever. And technologies that get a head start will probably keep them for a long time; but in general, not forever.

Thanks for reading! For other related articles by me, follow the links listed at the bottom of the article’s page on New Things Under the Sun (.com).And to keep up with what’s new on the site, of course subscribe!

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

Suppose we think there should be more research on some topic: asteroid deflection, the efficacy of social distancing, building safe artificial intelligence, etc. How do we get scientists to work more on the topic?

Buy it

One approach is to just pay people to work on the topic. Capitalism!

The trouble is, this kind of approach can be expensive. To estimate just how expensive, Myers (2020) looks at the cost of inducing life scientists to apply for grants they would not normally apply for. His research context is the NIH, the US’ biggest funder of biomedical science. Normally, scientists seek NIH Funding by proposing their own research ideas. But sometimes the NIH wants researchers to work on some kind of specific project, and in those cases it uses a “request for applications” grant. Myers wants to see how big those grants need to be to induce people to change their research topics to fit the NIH’s preferences.

Myers has data on all NIH “request for applications” (RFA) grant applications from 2002 to 2009, as well as the publication history of every applicant. RFA grants are ones where NIH solicits proposals related to a prespecified kind of research, instead of letting investigators propose their own topics (which is the bulk of what NIH does). Myers tries to measure how much of a stretch it is for a scientist to do research related to the RFA by measuring the similarity of the text between the RFA description and the abstract of each scientist’s most similar previously published article (more similar texts contain more of the same uncommon words). When we line up scientists left to right from least to most similar to a given RFA, we can see the probability they apply for the grant is higher the more similar they are (figure below). No surprise there.

Myers can also do the same thing with the size of the award. As shown below, scientists are more likely to apply for grants when the money on offer is larger. Again, no surprise there.

The interesting thing Myers does is combine all this information to estimate a tradeoff. How much do you need to increase the size of the grant in order to get someone with less similarity to apply for the grant at the same rate as someone with higher similarity? In other words, how much does it cost to get someone to change their research focus?

This is a tricky problem for a couple reasons. First, you have to think about where these RFAs come from in the first place. For example, if some new disease attracts a lot of attention from both NIH administrators and scientists, maybe the scientists would have been eager to work on the topic anyway. That would overstate the willingness of scientists to change their research for grant funding, since they might not be willing to change absent this new and interesting disease. Another important nuance is that bigger funds attract more applicants, which lowers the probability any one of them wins. That would tend to understate the willingness of scientists to change their research for more funding. For instance, if the value of a grant increases ten-fold, but the number of applicants increases five-fold, then the effective increase in the expected value of the grant has only doubled (I win only a fifth as often, but when I do I get ten times as much). Myers provides some evidence that the first concern is not really an issue and explicitly models the second one.

The upshot of all this work is that it’s quite expensive to get researchers to change their research focus. In general, Myers estimates getting one more scientist to apply (i.e., getting one whose research is typically more dissimilar than any of the current applicants, but more similar than those who didn’t apply) requires increasing the size of the grant by 40% or nearly half a million dollars over the life of a grant!

Sell it

Given that price tag, maybe a better approach is to try and sell scientists on the importance of the topic you think is understudied. Academic scientists do have a lot of discretion in what they choose to study; convince them to use it on the topic you think is important!

The article “Gender and what gets researched” looked at some evidence that personal views on what’s important do affect what scientists choose to research: women are a bit more likely to do female-centric research then men, and men who are exposed to more women (when their schools go coed) are more likely to do gender-related research. But we also have a bit of evidence from other domains that scientists do shift priorities to work on what they think is important.

Perhaps the cleanest evidence comes from Hill et al. (2021), which looks at how scientists responded to the covid-19 pandemic. In March 2020, it became clear to practically everyone in the world that more information on covid-19 and related topics was the most important thing in the world for scientists to work on. The scientific community responded: by May 2020 and through the rest of the year, about 1 in every 20-25 papers published was related to covid-19. And I don’t mean 1 in every 20-25 biomedical papers - I mean all papers!

This was a stunning shift by the standards of academia. For comparison, consider Packalen and Bhattacharya (2011), which looks at how biomedical research changed over the second-half of the twentieth century. Packalen and Bhattacharya classify 16 million biomedical publications, all the way back to 1950 and look at the gradual changes in disease burden that arise due to the aging of the US population and the growing obesity crisis. As diseases associated with being older and more obese became more prevalent in the USA, surely it was clear that those diseases were more important to research. Did the scientific establishment respond by doing more research related to those diseases?

Sort of. As diseases related to the aging population become more common, the number of articles related to those diseases does increase. But the effect is a bit fragile - it disappears under some statistical models and reappears in others. Meanwhile, there seems to be no discernible link between the rise of obesity and research related to diseases more prevalent in a heavier population.

Further emphasizing the extraordinary pivot into covid-related research, most of this pivot preceded changes in grant funding. The NIH did shift to issuing grants related to covid, but with a considerable lag, leaving most scientists to do their work without grant support. As illustrated below, the bulk of covid related grants arrived in September, months after the peak of covid publications (the NSF seems to have moved faster).

On the one hand, I think these studies do illustrate the common-sense idea that if you can change scientists beliefs about what research questions are important, then you can change the kind of research that gets done. But on the other hand, the weak results in Packalen and Bhattacharya (2011) are a bit concerning. Why isn’t there a stronger response to changing research needs, outside of global catastrophes?

It’s hard

I would point to two challenges to swift responses in science; these are also likely reasons why Myers (2020) finds it so expensive to induce scientists to apply for grants they would not normally apply for. Both reasons stem from the fact that a scientific contribution isn’t worth much unless you can convince other scientists it is, in fact, a contribution.

The first challenge with convincing scientists to work on a new topic is there need to be other scientists around who care about the topic. This is related to the model presented in Akerlof and Michaillat (2018). Akerlof and Michaillat present a model where scientists’ work is evaluated by peers who are biased towards their own research paradigms. They show that if favorable evaluations are necessary to stay in your career (and transmit your paradigm to a new generation), then new paradigms can only survive when the number of adherents passes a critical threshold. Intuitively, even if you would like to study some specific obesity-related disease because you think it’s important, if you believe few other scientists agree, then you might choose not to study it, since it will be such a slog getting recognition. There’s a coordination challenge - without enough scholars working in a field, scholars might not want to work in a field. (This paper is also discussed in more detail here)

The second challenge is that, even if there is a critical mass of scientists working on the topic, it may be hard for outsiders to make a significant contribution. That might make outsiders reluctant to join a field, and hence slow its growth. We have a few pieces of evidence that this is the case. Hill et al. (2021) quantify research “pivots” by looking at the distribution of journals cited in a scientists career and then measuring the similarity of journals cited in a new article to the journals cited in the scientist’s last three years.

For example, my own research has been in the field of economics of innovation and if I write another paper in that vein, it’s likely to cite broadly the same mix of journals I’ve been citing (e.g., Research Policy, Management Science, and various general economics journals). Hill and coauthors’ measure would classify this as being a minimum pivot of close to 0. I also have written about remote work, and that was a bit of a pivot for me; the work cited a lot of journals in fields I didn’t normally cite up until this point (Journal of Labor Economics, Journal of Computer-Mediated Communication, but also plenty of economics journals). Hill and coauthors’ measure would classify this as an intermediate pivot, greater than 0 but a lot less than 1. But if I were to completely leave economics and write something on the biology of covid-19, I might not cite any journals I’ve ever cited before; that would be measured as a the pivot maximum of 1. By this measure, most covid-related research involved a much bigger pivot than average.

Hill and coauthors then look to see what is the probability a given paper is in the top 5% for citations received in its field and year. The greater the pivot of the paper, the less likely a paper is to be highly cited.

Arts and Fleming (2018) provides some additional evidence on the difficulty of outsiders making major intellectual contributions, but among inventors instead of academics. As a simple measure of inventors entering new fields, they look at patents that are given a technology classification that has never been given to the inventor’s previous patents. As with Hill et al. (2021) they find these patents tend to receive fewer citations. One thing I quite like about this paper though is that they also go beyond citations and look at alternative measures of the value of a patent, such as whether the inventor or assignee chooses to pay the renewal fees to keep the patent active. By this measure too, patents from outsiders are less valuable. (This paper is discussed a bit more here)

Science is pretty competitive and if it’s harder to do valuable work in a new field, then it may well be in any given scientist’s best interest to stay in their lane. But that can make it hard for the system overall to respond to changing research needs.

It’s Not Impossible

While the above barriers make it harder for new scientific fields to emerge, clearly it does happen. To close, let’s look at two factors that might make change easier.

First, if you can initially solve the coordination challenge of getting a critical mass of scholars to focus attention on a new topic, then you can create a new equilibrium where pivoting into that field can be in any given individual’s self-interest. Covid-19 provides just such an example of a new equilibrium. It is too early to learn much from citations to covid-19 research, but as an early indicator Hill and coauthors look at the journals where covid-19 research gets published. They assign each journal a score based on the historical probability an article published there becomes a top-5% most cited publication for its field and year. The figure below compares the size of the pivot for covid-related and non-covid related research to the historical hit rate of the journal that publishes it.

In blue (for non-covid research) and red (for covid-related research), we can see the same pivot penalty as we observed before; article that involve a bigger research pivot are less likely to place in journals that tend to get highly cited. But the gap between these lines is also informative - because there seems to be a new consensus that covid-related research is so important, covid-related research tended to publish much better. Indeed, a pivot to covid that’s measured at around 0.7 appears to have about the same likelihood of becoming highly cited as a non-covid paper that executes a minot pivot measured at around 0.3. All else equal it’s better to make a smaller pivot to covid-related research, but large pivots are not nearly as unattractive as they previously were.

Second, if career incentives constrain scientists to stay in their lane and avoid branching out into new topics, then changing those incentives might also help. Evidence here is more mixed though.

On the one hand, we have a well-known paper by Azoulay, Graff-Zivin, and Manso (2011), which compared the recipients of Howard Hughes Medical Institute (HHMI) support to a control sample of early career prize winners in the life sciences. A key difference between these groups is that HHMI winners are relatively more insulated from the typical academic grant system: they receive at least 5 years of support that is not tied to any specific project, they have a relatively lax initial review, and are typically renewed at least once for 5 more years (and if not renewed, they get two years of funding to help keep the lab open while they search for new grants). In contrast, someone getting support on an NIH grant would often receive just three years of funding, with a comparatively low probability of being renewed.

Azoulay and coauthors use the MeSH lexicon - a standardized set of keywords assigned to biomedical papers by experts - to show HHMI investigators are more likely to explore new research topics than the control group. The MeSH words assigned to their papers tend to be of a more recent vintage, and there tends to be less overlap between the MeSH words assigned to their post-HHMI support papers and their pre-HHMI support papers when compared to the control group. That suggests being (somewhat) freed from the need to secure grant support via normal channels gives scientists the autonomy to explore new fields.

On the other hand though, we have a 2018 paper by Brogaard, Engelberg, and Van Wesep which looked at another set of career incentives that insulates a scientist from the pressure to conform: tenure. Brogaard, Engelberg, and Van Wesep look at 980 economists who at one point belonged to a top 50 economics or finance department between 1996 and 2014, and who were granted tenure by 2004. They track down these economists complete publication record in order to assess how they publish before they get tenure and in the ten years after.

For our purposes, one of their most interesting results is about whether economists use the autonomy of tenure to branch out into new journals or new areas. Unfortunately, Brogaard, Engelberg, and Van Wesep find no evidence that they do. Indeed, in some versions of their statistical tests, economists are slightly less likely to branch out after receiving tenure. And this isn’t just because economists take a breather for a year or two after getting tenure. This effect persists.

This leaves us in a bit of a muddle. For elite scientists, HHMI style support seems to have encouraged them to branch out and try new things. But for economists, the insulation of tenure did not. Is that because tenure is different from HHMI support, or because economists are different from biomedical scientists, or because elite scientists are different from everyone else? We don’t know.

Until then though, I think we can say a few things with confidence. First, the direction of research does respond to money and perceptions of intrinsic value. But it doesn’t appear to be super responsive. That might be because of some incentives scientists face to stay focused: fields may need a critical mass of sympathetic peers before it is individually rational to enter them, and even when a critical mass exists, it is challenging for outsiders to do top work in them, at least initially. Building a new field is probably pretty hard for these reasons; but if you can get the ball rolling, it’s also possible that it can continue going on it’s own momentum.

But more research is needed.

Thanks for reading! If you liked this, you might also enjoy this article about some of the forces that bias science away from novelty and towards conservatism. For other related articles by me, follow the links listed at the bottom of the article’s page on New Things Under the Sun (.com). And to keep up with what’s new on the site, of course subscribe!

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

How do scientists and inventors decide what to work on? Part of it comes down to what they find personally meaningful. And that, in turn, can be informed by your specific life experiences.

We can actually see this in data, if we look in the right places. Every human life is unique, but in this article we’re going to collapse diverse life experiences down into something quite crude: binary gender. This can be a useful exercise because, as we’ll see in a minute, there are a variety of non-controversial strands of evidence that women are slightly more likely to work on some things than men, and that this may well reflect different perceptions on what questions are meaningful (likely derived from lived experience). But it’s also an exercise fraught with risks, because there may well be a variety of other factors that lead men and women to work on different topics.

For example, it turns out that the share of scientists who are women differs a lot across fields. Over 1990-2011, West et al. (2013) find that 41% of published authors in sociology are women, compared to just 14% of authors in economics. Should we read this as evidence that different life experiences lead men and women to have different interests in sociology, relative to economics? I don’t think so, as there are many other possible barriers to entry in one field relative to another: discrimination or bias, culture, access to networks, and so on.

But cognizant of these risks, we can still look at the choices people make within a specific, narrow, field. That won’t completely rule out the possibility that the choices of research question is constrained by bias and discrimination (even within a specific subfield), but it’s an informative place to start.

Gender and the Direction of Technological Innovation

Koning, Samila, and Ferguson (2021) analyze a set of ~400,000 US biomedical patents from 1976-2010. For each patent, they classify the gender of the inventors based on their names (they can match 98% of inventors with 95% confidence; note they have a binary classification of gender, which they point out is a not appropriate in all cases). They next run a sample of the text of these biomedical patents through an algorithm designed to classify academic articles into different biomedical categories. This algorithm classifies about 13% of these biomedical patents as being related to “female organs, diseases, physiologic processes, genetics, etc.” and another 13% as being related to male analogues. They also cross-check these classifications with a variety of alternative metrics: are patents classified as having a female focus more likely to be based on all-female clinical trials? Are they more likely to be related to diseases that have much higher incidence in women than men?

They find that patents where the majority of the inventors are men are more likely to patent male-focused patents and patents where the majority of the inventors are women are more likely to patent female-focused patents.

Einiö, Feng, and Jaravel (2019) find something similar using a variety of other datasets.

Via Nielsen data, they have information on a sample of households that purchase a wide range of specific consumer products, as well the barcode associated with each product. They can link these barcodes back to the manufacturing firm, which they can then link to patents owned by these firms. Finally, they can use the same strategy as Koning, Samila, and Ferguson to classify the inventors on these patents as men or women. This lets them draw a connection between the gender of inventors and the gender of the people purchasing products manufactured by the companies that own these patents. They find, indeed, that products that are disproportionately purchased by female households also have a higher share of female inventors.

Stepping away from patents, they also use data from Crunchbase to identify startups with female founders, and use the same Nielsen data to assess the gender of purchasers of these startups’ products. They find that startups with a female founder are also more likely to sell products with a disproportionately female consumer base.

So along three different axis, we see that inventors who are women (or at least, have names traditionally given to women) are more likely to develop new products and new technologies that appeal to women, as compared to inventors who are probably men. We can see a similar effect when we look at research papers.

Gender and the Direction of Science

Koning, Samila, and Ferguson (2021) also perform a similar exercise on biomedical science. They extract data on about 2 million biomedical research articles published between 2002 and 2020 and once again classify the authors as male or female based on their names. They focus this analysis on original research published in journals whose articles tend to be highly cited by patents (because this paper’s main focus was on patents).

There’s no need to process these articles through a machine-learning algorithm to see if they focus on female or male related topics (i.e., pertain to female/male organs, diseases, physiologic processes, genetics, etc.), since the articles had already been classified into biomedical categories as part of the publication process. Again, they find articles with more women as authors are more likely to focus on female topics. The following figure is how much more likely a given article is to be female-focused, estimated with two different statistical models.

(If interested - the black estimate comes from a basic regression, controlling for journal-year and team size-year commonalities, the gray one from a model where teams are matched to all-male control teams that are observationally similar).

Nielsen et al. (2017) looks at the influence of gender on biomedical research in a different way. Rather than seeing if papers are classified as pertaining to male or female biomedical categories, Nielsen and coauthors look to see if the gender of the author affects the probability papers incorporate gender and sex analysis, defined by them as “scientific approaches that are aimed at understanding how social and behavioral differences between women, men and gender-diverse people (gender analysis) and biological differences between female and male research subjects (sex analysis) relate to health outcomes.” The GenderMed dataset identifies approximately 5,000 papers, published between 2008 and 2015, incorporating such an analysis. Over the same period, about 1.5 million papers were published on diseases for which such an analysis can be relevant, so a gender and sex analysis is a pretty rare event (3 in 1,000 papers). Nonetheless, Nielsen and coauthors find papers with more women coauthors were more likely to include such an analysis, even within a specific disease category and country.

First and Second Order Effects of Representation

These differences are never huge, but they are persistent across quite a range of evidence. It seems to me pretty likely this reflects either a relative lack of awareness or a relative lack of empathy by male scientists/inventors about the importance of these issues. That, in turn, supplies one reason why representation in science is important. If we want research to broadly reflect the priorities of society at large, and if a part of society is not represented in the research space, the above evidence suggests we’re not going to get the kind of research we want.

That’s a first order effect: if you want research related to a specific group, the chances of getting it may be higher if that group is able to participate in the research process. But there is also a second order effect. We also have some evidence that, as representation improves, everyone’s research priorities shift.

For example, let’s look again at the probability majority male inventor teams works on a female-focused biomedical patent compared to a male-focused one (left figure below). The gap between the two has been closing. Could that be because the number of female inventors has been rising (right figure below), which is increasing male awareness of female-focused research?

Nielsen et al. (2017) also look beyond the impact of female coauthors, to the broader effect of more women in a particular subfield. They find that when there are more women studying a particular disease, this also increases the probability that studies include a gender and sex analysis, independent of the composition of the team of coauthors on any individual study. In other words, a team of men is more likely to include a gender and sex analysis if they are working on a disease where more of the scientists studying the same disease are women.

That could be because as women attain positions of influence - serving as peer reviewers, grant reviewers, PhD supervisors, etc. - the rest of the field becomes more responsive to their concerns. As discussed in “Conservatism in Science” we do have evidence that scientists constrain the feasible choice of research topics for their peers through a variety of channels like this.

But it could also be that people’s awareness and empathy can change when they are around people different from themselves.

An exposure experiment?

It’s tough to experimentally vary “awareness” and “empathy” but a 2021 dissertation by Truffa and Wong look at a situation that has parallels to this. Between 1960 and 1990 76 all-male US universities went coed and began to admit women as undergraduates. Truffa and Wong look to see what happens to the research of these universities before and after they make the switch.

They identify all the academic papers (in the Microsoft Academic Graph) with authors affiliated with these universities, and then they classify papers as being related to gender if the title or abstract includes various keywords (for example “lady”, “female”, “misogyny”, “mothers”, “sexism”). Note that this approach covers all fields, not just the life sciences. A manual inspection of 100 random papers identified as being gender-related by this keyword approach finds the method works pretty well, with only about 8% of the papers identified in this way not really related to gender. They also get similarly encouraging results when they benchmark this against other techniques (for example, by training a machine-learning algorithm on gender studies papers and papers in gender-related journals; or focusing just on biological papers and comparing the results to the biomedical classification systems discussed earlier).

When these universities began to admit women, they also began to produce more research related to gender, as illustrated in the following figure (with the year of going coed centered at zero). The effect is small in absolute terms, but large relative to the pre-existing number of papers related to gender (just 2% in 1960 at these universities).

Why does this happen?

Well, it could have nothing to do with the changing composition of the undergraduates. Maybe these institutions had already decided to take gender more seriously, and that was reflected simultaneously in research and the choice of which students to admit. But Truffa and Wong argue the actual reason for the transition from all-male undergraduates to mixed was not so high-minded. Instead, all-male schools found they were having a harder time attracting “top boys”, who increasingly preferred to attend coed universities. So, to continue attracting “top boys” these institutions (grudgingly?) began to accept undergraduate women too. In other words, the decision to go coed does not seem to have been driven by faculty itching to take gender more seriously in their research.

So why did research begin to take gender more seriously after the schools went coed? It could be that going coed was accompanied by increased hiring of women faculty. As we’ve seen above, it shouldn’t be that surprising that institutions that hire more women will end up getting more women-centered research. Truffa and Wong do find that the share of new assistant professors who were women rose from about 13% to 17% after schools went coed. But the effect goes well beyond these new hires.

Instead, a large part of this effect seems to be driven by pre-existing faculty revising their research preferences. As the figure below shows, even incumbent male professors were more likely to do gender related research after the undergraduate body included women.

These effects were bigger at colleges that admitted more women after the change. Truffa and Wong also present some evidence about a few very concrete channels through which these changes might have occurred. For example, the number of gender-related classes increased significantly after colleges went coed, and in preparing and teaching these classes, faculty might have begun to engage more with these ideas. Second, undergraduates may sometimes participate in the research process. For example, in psychology research of this era, it was common to perform experiments with undergraduate volunteers, and a coed pool of volunteers would have made it easier to do experiments that examined gender differences. Truffa and Wong show, in psychology, the increase in gender-related papers does indeed stem from experimental papers.

The Importance of Representation

The takeaway then, is that representation matters. When a group previously excluded from research enters, it may carry with it research priorities that are better aligned with the needs of the excluded group. Fortunately, as those ideas enter the bloodstream of a research community, it looks like the priorities do get taken up more widely.

But we remain a long way from gender parity in most of research. Holman, Stuart-Fox, and Hauser (2018) examine the rate at which the gender gap in science is closing. Across 115 disciplines that publish to the PubMed or arXiv paper repositories, 87 had fewer than 45% women authors in 2016. In almost every one of these fields the share of women was increasing over time, but at the current rate in most cases it would be well over a decade for authors to come within 45%, sometimes far longer (and longer still for more senior positions). Lastly, there is no reason to believe similar dynamics don’t play out with other underrepresented groups.

Thanks for reading! If you liked this, you might also enjoy this article about some of the forces that bias science away from novelty and towards conservatism. For other related articles by me, follow the links listed at the bottom of the article’s page on New Things Under the Sun (.com). And to keep up with what’s new on the site, of course subscribe!

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

Isaac Asimov’s Foundation series imagines a world where there are deep statistical regularities underlying social history, which “psychohistorian” Hari Seldon uses to forecast (and even alter) the long-run trajectory of galactic civilization. Foundation is science fiction, but what if we could develop a model of the “deep laws” governing society? What might they look like?

Well, as a starting principle, we might reflect that in the very long run, material prosperity is driven primarily by technological progress and innovation. So long-run forecasts might depend on the future rate of technological change. Will it accelerate? Slow? Stop?

But this poses a problem. By definition, innovation is the creation of things that are presently unknown. How can we say anything useful about a future trajectory that will depend on things we do not presently know about?

In fact, I think we can say very little about the long-run outlook of technological change, and even less about the exact form such change might take. Psychohistory remains science fiction. But a certain class of models of innovation - models of combinatorial innovation - does provide some insight about how technological progress may look over very long time frames. Let’s have a look.

Strange Dynamics of Combinatorial Innovation

Twenty years ago, the late Martin Weitzman spelled out some of the interesting implications of combinatorial models of innovation. In Weitzman (1998), innovation is a process where two pre-existing ideas or technologies are combined and, if you pour in sufficient R&D resources and get lucky, a new idea or technology is the result. Weitzman’s own example is Edison’s hunt for a suitable material to serve as the filament in the light bulb. Edison combined thousands of different materials with the rest of his lightbulb apparatus before hitting upon a combination that worked. But the lightbulb isn’t special: essentially any idea or technology can also be understood as a novel configuration of pre-existing parts.

An important point is that once you successfully combine two components, the resulting new idea becomes a component you can combine with others. To stretch Weitzman’s lightbulb example, once the lightbulb had been invented, new inventions that use lightbulbs as a technological component could be invented: things like desk lamps, spotlights, headlights, and so on.

That turns out to have a startling implication: combinatorial processes grow slowly until they explode.

Let’s illustrate with an example. Suppose we start with 100 ideas. With a little math, we can show the number of unique pairs of ideas that can be created from these 100 is 4950. But most of these ideas born in this way, from combining all possible ideas, are going to be garbage like “chicken ice cream.” Let’s assume only 1% are viable. Rounding down, our supply of 100 ideas yields 49 new viable ideas (1% of 4950) after extensive R&D to test all possible combinations.

In the next period, we’ll have 149 ideas available. We can go through the whole exercise, calculating how many unique pairs are possible (11,026). Of course, 4950 of those we’ve already investigated. But there are still 6076 new pairs, and if 1% of them are useful, we’ve added 61 more ideas to our stock of ideas. Now we have 210 ideas, and we can go through the process again. If we grind this process forward in time, the number of good ideas in every period of R&D evolves as follows.

Up through period 5, the growth of ideas via this combinatorial process looks roughly like an exponential process. But by the time we get to period 6, an explosion is underway, where the number of new ideas in the next period always becomes so large that it renders the cumulative number of all prior ideas negligible.

Do such processes happen in the real world though? Well, here is an estimated figure of GDP per capita over two millenia. Looks familiar!

The Industrial Revolution

Weitzman alludes in passing to the fact that combinatorial innovation’s prediction of a long period of slow growth followed by an explosion seems to fit the history of innovation quite well. A 2019 working paper by Koppl, Devereaux, Herriot, and Kauffman expands on this notion as a potential explanation for the industrial revolution. For Weitzman, innovation is a purposeful pairing of two components, but for Koppl, Devereaux, Herriot, and Kauffman, this is modeled as a random evolutionary process, where there is some probability any pair of components results in a new component, a lower probability that triple-combinations result in a new component, a still lower probability that quadruple-combinations result in a new component, and so on. They show this simple process generates the same slow-then-fast growth of technology.

Now; their point is not so much that the ultimate cause of the industrial revolution is now a solved question. They merely show a process where people occasionally combine random sets of created artifacts around them and notice when they yield useful inventions inevitably transitions from slow to very rapid growth. It’s not that this explains why the industrial revolution happened in Great Britain in the 1700-1800s; instead, they argue an industrial revolution was inevitable somewhere at some time, given these dynamics. Also interesting to me is the idea that this might have happened regardless of the institutional environment innovators were working in. If random tinkering is allowed to happen, with or without a profit motive, then you can get a phase-change in the technological trajectory of a society once the set of combinatorial possibilities grows sufficiently large.

The Birth of Exponential Growth

Weitzman and Koppl, Devereaux, Herriot, and Kauffman both show that innovation under a combinatorial process accelerates, which is consistent with a very long run view of technological progress. But it doesn’t seem to fit today’s world. Indeed, for one hundred years, growth in GDP per capita has been remarkably consistent. What’s going on here?

Again, Weitzman proposes a solution. The reason technological progress does not accelerate in all times and places is because in addition to ideas, Weitzman assumes innovation requires R&D effort. In the beginning, we will usually have enough resources to fully fund the investigation of all possible new ideas. So long as that’s true, the number of ideas is the main constraint on the rate of technological progress and we’ll see accelerating technological progress. But in the long run, the number of possible ideas explodes and growth becomes constrained by the resources we have available to devote to R&D, not by the supply of possible ideas.

The basic idea is easy to see. Start with the previous example, but imagine it costs, say, $1mn to combine two ideas and then see if they work out. In 2020 US GDP was $21 trillion, so if the US economy was 100% devoted to R&D it could investigate 21mn ideas per year. Returning to our illustrative example from earlier, that means the US could fully fund R&D on every possible idea for periods 0-5 (that’s not necessarily obvious, but if you do the math, you’ll see in each of these periods there are less than 21mn possible ideas). Indeed, even in period 5, when there are 1.6mn possible ideas to investigate, the US could fully investigate every idea with less than 10% of the economy’s resources. During this time, technological progress would feel like an accelerating exponential process.

But in period 6, there are 153mn possible new ideas, far more than the US has the resources to investigate. From that point going forward, the economy must grow at an exponential rate. The number of ideas has become irrelevant, since there are so many that we’ll never begin to explore them all. All that’s relevant is the amount of resources you can throw at R&D, and that can never be more than 100% of all resources. Think of how an animal population grows at a constant exponential rate when food is not a constraint, because each member of the population spawns a fixed number of children. Just so, an economy grows at a constant rate when the number of ideas is not a constraint, because each technology generates some income, which can be reinvested in R&D to generate a fixed number of “children” technologies.

Why might technological progress be exponential?

But there is a buried assumption in here that we might want to scrutinize. In our illustrative example, we assumed 1% of possible ideas are viable. We don’t need to assume it’s 1%, but if we want to observe constant exponential progress over the very long run in this model, we do need to assume there is some constant proportion of ideas that are viable. And that hardly seems to be guaranteed. Maybe innovation gets harder, for example, because combining ideas that are themselves large hierarchies of combined ideas, get progressively harder to combine. Or maybe innovation gets easier, for example, because eventually we invent things like artificial intelligence to help us efficiently find the viable combinations.

Weitzman explores some of these possibilities, but the best he can really do is establish some possible cases and point to the kinds of criteria that lead to one case or another. If it gets progressively harder to combine ideas, and if ideas are sufficiently important for progress, then growth can halt. On the other hand, if ideas get easier to combine over time then you can get a singularity - infinite growth! And then there are cases where growth remains steady and exponential. That doesn’t help us much to predict the long-run trajectory of society, a la Hari Seldon in the Foundation series.

But Jones (2021) maps out an alternative plausible scenario. So far we have assumed some ideas are “useful” and others are not, and progress is basically about increasing the number of useful ideas. But this is a bit dissatisfying. Ideas vary in how useful they are, not just if they’re useful or not. For example, as a source of light, the candle was certainly a useful invention. So was the light bulb. But it seems weird to say that the main value of a light bulb was that we now had two useful sources of light. Instead, the main value is that light bulbs are a better source of light than candles.

Instead, let’s think of an economy that is composed of lots and lots of distinct activities. Technological progress is about improving the productivity in each of these activities: getting better at supplying light, making food, providing childcare, making movies, etc. As before, we’re going to assume technological progress is combinatorial. But we’re now going to make a different assumption about the utility of different combinations. Instead of just assuming some proportion of ideas are useful and some are not, we’re going to assume all ideas vary in their productivity. Specifically, as an illustrative example, let’s assume the productivity of combinations is distributed according to a normal distribution centered at zero.

This assumption has a few attractive properties. First off, the normal distribution is a pretty common distribution. It’s what you get, for example, if you have a process where we take the average of lots of different random things, each of which might follow some other distribution. If technology is about combining lots of things and harnessing their interactions, then some kind of average over lots of random variables seems like not a bad assumption.

Second, this model naturally builds in the assumption that innovation gets progressively harder, because there are lots of new combinations with productivity a bit better than zero (“low hanging fruit”), but as these get discovered the share of combinations with better productivity get progressively less common. That seems sensible too.

But this model also has a worrying implication. Normal distributions are “thin-tailed.” As productivity gets higher, the difficulty of finding a combination with even higher productivity gets much much harder. For example, suppose productivity for some economic activity is currently zero - basically, we can’t do this activity at all. If we do a little R&D and build a new combination, the probability it will have productivity higher than zero is 1/2. So, on average, we’ll find one useful innovation for every two we explore. If productivity is currently “1”, it gets harder to find a new combination with productivity higher than that. On average, we’ll find one useful innovation for every six we explore. If we grind this forward, we get the following figure, which plots the expected number of combinations that must be explored in order to find one combination with higher productivity than the current best.

By the time we advance to a technology with productivity of 5, on average we’ll have to explore more than 3 million combinations before we find one better!

It looks like we have a process that initially grows very slowly and then explodes. But that’s precisely the shape of a combinatorial explosion as well! Indeed, the point of Jones’ paper is to show these processes balance each other out. Under a range of common probability distributions (such as the standard normal, but also including others), finding a new technology that’s more productive than the current best gets explosively harder over time. However, the range of options we have also grows explosively, and the two offset each other such that we end up with constant exponential technological progress. Which is a pretty close approximation to what we’ve observed over the last 100 years!

This is interesting, but it does conflict with Weitzman’s story of why combinatorial innovation eventually leads to steady exponential growth. In Weitzman, steady growth arose from the fact that doing R&D required some real resources. So even though the set of possible ideas explodes in a way that can offset explosively harder innovation, in Weitzman’s paper it’s not possible to explore all those possibilities because R&D resources are limited. The result is steady exponential growth in Weitzman’s model.

If we had to spend real economic resources to explore possible combinations in Jones’ model, we wouldn’t end up with steady exponential growth. That’s because we would only be able to increase our resources to explore ideas at an exponential rate, but the number of ideas we would need to explore in order to find a technology better than the current best practice increases at a faster than exponential rate. The net result would be slowing technological progress.

But Jones has a different notion of R&D in mind. In Jones’ model, we still need to spend real R&D resources to build new technologies. But it’s sort of a two-stage process, where we costlessly sort through the vast space of possibilities and then proceed to actually conduct R&D only on promising ideas. This way of thinking reminds me an essay on mathematical creation by famed mathematician Henri Poincaré. He describes his two-stage process of creating new math as follows:

One evening, contrary to my custom, I drank black coffee and could not sleep. Ideas rose in crowds; I felt them collide until pairs interlocked, so to speak, making a stable combination. By the next morning I had established the existence of a class of Fuchsian functions, those which come from the hypergeometric series; I had only to write out the results, which took but a few hours.

Poincaré has a two-stage process of invention. He explores the combinatorial space of mathematical ideas, and when he hits on a promising combination, he does the work of writing out the results. But is it really possible for a human mind, even one as good as Poincaré’s, to consider all possible combinations of mathematical ideas?

No. As a mathematician, Poincaré is aware of the fact that the space of possible combinations is astronomical. Mathematical creation is about choosing the right combination of mathematical ideas from the set of possible combinations. But, he notes:

To invent, I have said, is to choose; but the word is perhaps not wholly exact. It makes one think of a purchaser before whom are displayed a large number of samples, and who examines them, one after the other, to make a choice. Here the samples would be so numerous that a whole lifetime would not suffice to examine them. This is not the actual state of things. The sterile combinations do not even present themselves to the mind of the inventor. Never in the field of his consciousness do combinations appear that are not really useful, except some that he rejects but which have to some extent the characteristics of useful combinations. All goes on as if the inventor were an examiner for the second degree who would only have to question the candidates who had passed a previous examination.

So long as we have a way to navigate this vast cosmos of possible technologies and zero in on promising approaches, we can obtain constant exponential growth - even in an environment where pushing the envelope and finding something more useful than the current best practice gets progressively harder.

Long-run Growth and AI

But what if there are limits to this process? Human minds may have some unknown process of organizing combinations, to efficiently sort through them. But there are quite a lot of possible combinations. What if, eventually, it becomes impossible for human minds to efficiently sort through these possibilities? In that case, it would seem that technological progress must slow, possibly a lot.

This is essentially the kind of model developed in Agrawal, McHale, and Oettl (2018). In their model, an individual researcher (or team of researchers) has access to a fraction of all human knowledge, whether because it’s in their head or they can quickly locate the knowledge (for example with a search engine). As a general principle, they assume the more knowledge you have access to, the better it is for innovation.

But as in all the other papers we’ve seen, they assume research teams combine ideas they have access to in order to develop new technologies. And initially, the more possible combinations there are, the more valuable discoveries a research team can make. But unlike in Jones (2021), Agrawal and coauthors build a model where the ability to comb through the set of possible ideas weakens as the set gets progressively larger. Eventually, we end up in a position like Weitzman’s original model, where the set of possibilities is larger than can ever be explored, and so adding more potential combinations no longer matters. Except, in this case, this occurs due to a shortage of cognitive resources, rather than a shortage of economic resources that are necessary for conducting R&D.

As we suspected, they show that as we lose the ability to sort through the space of possible ideas, technological progress slows (though never stops in their particular model).

But if the problem here is we eventually run out of cognitive resources, then what if we can augment our resources with artificial intelligence?

Agrawal and coauthors are skeptical this problem can be overcome with artificial intelligence, at least in the long run. They argue convincingly that no matter how good an AI might be, there is always a number of components where it becomes implausible for a super intelligence to search through all possible combinations efficiently. If that’s true, then in the long run any acceleration in technological progress driven by the combinatorial explosion must eventually stop when our cognitive resources lose the ability to keep up with it.

To illustrate the probable difficulty of searching through all possible combinations of ideas, let’s think about a big number: 10^80. That’s about how many atoms there are in the observable universe. That would seem like a difficult number of atoms for an artificial intelligence to efficiently search over. Yet if we have just 266 ideas, then the number of possible combinations is about equal to 10^80, i.e., the number of atoms in the universe!

Still, maybe there are ways for the AI to efficiently organize this gigantic set, so that it does not actually have to search through more than a tiny sliver of the combinations?

The trouble is, the number of ideas we have available must be much, much larger than 266. As a starting count, we could turn to the US Patent and Trademark Office’s technology classification system, which attempts to classify patents into distinct technological categories. It has about 450 “three digit” classifications, but these categories seem to broad to be much use to an inventor. They include categories like “bridges” or “artificial intelligence: neural networks.” If we instead turn to a more detailed set of classifications, there are over 150,000 subclassifications which get quite specific. If we were to take that as our measure of the number of possible combinations, then we’re talking about 10^45,181 possible combinations of technology. If you took every atom in the universe and stuffed another universe worth of atoms inside each atoms, and then did that again with every one of those new atoms, and then repeated that process in a nesting structure 500 levels deep, you would be getting in the ballpark of how many combinations there are to consider. Seems challenging?

That said, the point is not that an AI can’t accelerate innovation. Perhaps there are domains where the number of possible combinations is larger than human minds can search but not artificial ones. In such domains, AI could accelerate innovation for awhile, via the combinatorial approach. But they would argue, eventually, any given domain would probably accrue too many ideas for any supercomputer to search through all possible combinations.

And lastly, there is more to innovation than searching through combinations. Agrawal and coauthors also point out that AI can serve like a recommender, helping you access more knowledge, and in their model it’s always good to have access to more ideas (even if you aren’t going to be searching through all combinations). I think of it as more like the ability to access a really good google for existing knowledge, so you can borrow and adapt solutions developed elsewhere. In their model, the ability of AI to put you in touch with more of knowledge that’s out there, the better it is for innovation in the short run and as far into the future as you care to go.

Psychohistorians of Technological Progress

To return to the motivation for this article: combinatorial dynamics are a plausible candidate for a deep law of innovation. Those dynamics will tend to deliver slow growth that accelerates for a time, as the combinatorial explosion heats up. But beyond that, two kinds of process may work to stymie continued explosive growth.

Innovation might get harder. If the productivity of future inventions are like draws from some thin-tailed distribution (possibly a normal distribution), then finding better ways of doing things gets so hard so fast that this difficulty offsets the explosive force of combinatorial growth.

Exploring possible combinations might take resources. These resources might be cognitive or actual economic resources. But either way, while the space of ideas can grow combinatorially, the set of resources available for exploration probably can’t (at least, it hasn’t for a long while).

If I can get a bit speculative, ultimately, these ideas seem linked to me. To start, it seems to me that it must take resources to explore the space of possible ideas, whether those resources are cognitive or economic. It may be that, we are still in an era where human minds can efficiently organize and tag ideas in the combinatorial space so that we can search it efficiently. But I suspect, even if that’s true, it can’t last forever. Eventually the world gets too complicated. That suggests the ultimate rate of technological progress depends on how rapidly we can increase our resources for exploring the space of ideas. (If we need more cognitive resources, then that means resources in the form of artifical intelligence and better computers). And our ability to increase our resources with new ideas is a question that falls squarely in the domain of Jones (2021): how productive are new ideas?

If the productivity of new ideas follows a thin-tailed distribution, then we would seem to be in trouble: according to Jones (2021) only so long as we can hunt through the combinatorial landscape are we able to increase productivity at an exponential rate. But an exponential rate isn’t good enough if the resources needed to explore a combinatorial space grow at a faster-than-exponential rate (which they must if they’re proportional to the number of possible ideas). In that case, it would seem that technological progress would need to slow down over time.

On the other hand, suppose the productivity of new ideas follows a fat-tailed distribution. That’s a world where extremely productive technologies - the kind that would be weird outliers in a thin-tailed world - are not that uncommon to discover. Well, in that world, Jones (2021) shows that the growth rate of the economy will be faster than exponential, at least so long as it can efficiently search all possible combinations of ideas. And faster than exponential growth in resources is precisely what we would need to keep exploring the growing combinatorial space.

Which outcome is more likely? Well, the future has no obligation to look like the past, but given that innovation seems to be getting harder, I am hoping for the best but bracing for the worst.

Thanks for reading! If you liked this, you might also enjoy these much more data driven articles arguing that innovation has tended to get harder, and that this might be because of the rising burden of knowledge. For other related articles by me, follow the links listed at the bottom of the article’s page on New Things Under the Sun (.com). And to keep up with what’s new on the site, of course…

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

It might seem obvious that we want bold new ideas in science. But in fact, really novel work poses a tradeoff. While novel ideas might sometimes be much better than the status quo, they might usually be much worse. Moreover, it is hard to assess the quality of novel ideas because they’re so, well, novel. Existing knowledge is not as applicable to sizing them up. For those reasons, it might be better to actually discourage novel ideas, and to instead encourage slow and incremental expansion of the knowledge frontier. Or maybe not.

For better or worse, the scientific community has settled on a set of norms that appear to encourage safe and creeping science, rather than risky and leaping science.

Blocking New Approaches

How might you identify whether a scientific community blocks new ideas? Merely observing the absence of new ideas isn’t enough; it could be that there just aren’t any new ideas worth pursuing. But what if you could identify sets of scientific fields that are quite similar to each other? Next, imagine you suddenly and randomly exile the dominant researchers in half the fields in order to stymie their ability to block new research. You could then compare what happens in fields that lost these dominant researchers to ones that didn’t. In particular, do new people enter these fields and pursue new and novel ideas? If so, that suggests there were new ideas worth investigating, but the dominant researchers in the field were blocking investigation of them. Alternatively, if new people enter these fields but basically continue doing the same thing as the previous generation, that suggests the opposite; the dominant researchers were not blocking anything and everyone in the field was just pursuing what they thought were the best ideas available at the time.

You can’t really do this experiment because you’re not going to get IRB approval to exile dominant researchers. It’s a bad thing to do. But Azoulay, Fons-Rosen, and Graff Zivin (2019) try to learn something from the sad fact that we live in a world where bad things do happen all the time. They begin by identifying a large set of superstar researchers in the life sciences. It turns out, 452 of these superstars died in the midst of an active research career. This is going to be analogous to the “exile” I described above in this hypothetical idealized experiment, though applying only to one prominent researcher per field, rather than the full set of dominant ones. (Note – as I’ll discuss later, the findings discussed below should not be interpreted as implying these superstars were behaving unethically or something. It’s more complicated than that.)

Azoulay, Fons-Rosen and Graff Zivin next identify a large set of small microfields in which these deceased researchers were active in the five years before their deaths. Each of these microfields consists of about 75 papers on average. For each of these microfields, they find sets of matching microfields that are similar in various ways. Most importantly, these matched microfields are ones that include active participation by a superstar researcher who did not pass away. The idea is to compare what happens in a microfield where a superstar researcher passed away to what happens in similar microfields where a superstar researcher did not pass away.

The first thing that happens is new people do begin to enter the field. Following the death of a superstar researcher, there is a slight increase in the number of new articles published in that microfield by people other than the superstar, as compared to microfields where the superstar lives. These new papers are disproportionately written by people who were not previously publishing in the microfield, rather than by people who were active publishing more.

Most relevant for our inquiry today, these new researchers appear to bring new ideas and new approaches with them. They don’t cite the existing work in this microfield, including the work of the superstar. And their work is assigned a greater share of novel scientific keywords from the MeSH lexicon (a standardized classification system used in biomedicine). Also important – the new work tends to be well cited. There’s a greater influx of highly cited work than low-citation work.

Can we say anything more about the specific mechanisms going on here?

Well, we can say it’s not simply a case of these dominant researchers vetoing grant proposals and publication from rival researchers. Only a tiny fraction of them were in positions of formal academic power, such as sitting on NIH grant review committees or serving as journal editors, when they passed away. So what else could it be? I’m going to tentatively suggest it’s about the dominance of the ideas the superstar researcher promoted. We’ll return to this notion later. But first, let’s look at some evidence about how scientists resist the influence of new ideas in a field.

Citing novel work

Let’s start at the end. Suppose we’ve recently published an article on an unusual new idea. How is it received by the scientific community?

Wang, Veugelers, and Stephan (2017) look at academic papers published across most scientific fields in 2001 and devise a way to try and measure how novel these papers are. They then look at how novel papers are subsequently cited (or not) by different groups.

To measure the novelty of a paper, they rely on the notion that novelty is about combining pre-existing ideas in new and unexpected ways. They use the references cited as a proxy for the sources of ideas that a paper grapples with, and look for papers that cite pairs of journals that have not previously been jointly cited. The 11.5% of papers with at least one pair of journals never previously cited together in the same paper are called “moderately” novel in their paper.

But they also go a bit further. Some new combinations are more unexpected than others. For example, it might be that I am the first to cite a paper from a monetary policy journal and an international trade journal. That’s kind of creative, maybe. But it would be really weird if I cited a paper from a monetary policy journal and a cell biology journal. Wang, Veuglers, and Stephan, create a new category for “highly” novel papers, which cite a pair of journals that have never been cited together in the past, and also are not even in the same neighborhood. Here, we mean journals that are not well “connected” by some other pair of journals.

For example, if journals A and B have never been cited together that’s kind of novel. But suppose A and C have been jointly cited, and also B and C have also been jointly cited. That means there is a journal that acts as a kind of bridge from A to B, even if they’ve never actually been paired themselves. More novel would be a pair of journals X and Y that don’t even share this kind of bridge (or at least, there are fewer bridges than usual). Wang, Veugelers, and Stephan designate the 1% of all papers that make the most unusual novel connections, defined in this way, as highly novel.

Echoing what I said at the beginning, highly novel work seems desirable; it’s much more likely to become one of the top 1% most cited papers in its field.

But importantly, this isn’t really because people in the field quickly recognize it’s a breakthrough. In fact, this recognition disproportionately comes from other fields, and with a significant delay. Restricting attention to citations received from within the same field, novelty isn’t rewarded.

And even looking across all fields, these citations come late. Restricting attention to citations received in just the first three years, again, novelty isn’t really rewarded.

Moreover, Wang, Veugelers, and Stephan also show moderately and highly novel papers are less likely to be published in the best journals (as measured by the journal impact factor). That suggests to me that highly novel work is also less likely to be published at all, and so the results we’re seeing might be over-estimating the citations received, since they rely on novel work that was good enough to clear skeptical peer review.

From Citations to Grants

All this is important because the reception of your prior work plays a role in your ability to do future work. Li (2017) provides some evidence on this by looking at about 100,000 NIH grant applications filed between 1992 and 2005. Funding for research from the NIH is limited - only 20% of applicants succeed - and so the NIH tries to prioritize funding projects that will have the biggest impact. The trouble here is that assessing the quality of frontier research is hard; those with enough knowledge to serve as good evaluators tends to be people who are also using that knowledge to actually do frontier research. Accordingly, grant applications are assessed by a review committee drawn from other active researchers in the area. They’re looking for proposals they think will generate new and significant scientific findings, taking into account the proposed research questions, methodologies, and capability of the researcher to follow through.

But what constitutes the most impactful work? Li begins by showing a proposal is more likely to be funded if the reviewer has previously cited the applicant’s prior work. Across applicants who are judged by the same committee, and have similar numbers of prior citations and grant awards and similar coarse demographic characteristics, the probability of funding increases by 3.3 percentage points for every (permanent) committee member who has cited the applicant’s prior work. That’s not too surprising on it’s own; if reviewers liked the applicant’s prior work enough to cite it, they probably like proposals for more work in the same vein.

But this result does imply something worrying. Proposals are not assigned to reviewers at random. Instead, they are usually given to whoever has the closest and most relevant expertise. That means it’s likely that the reviewers of a grant application will be researchers drawn from the applicant’s own scientific field. And if the applicant has a history of doing highly novel research, Wang, Veugelers, and Stephan’s work suggests these reviewers are less likely to have previously cited it. And Li (2017) suggests that means these applicants will face a tougher time getting funding.

Evidence of less indirect bias against novel ideas come from Boudreau et al. (2016). They convinced a university to run an experiment in grant-making for them, getting 150 applicants to submit a short proposal for a research project on endocrine-related disease, dangling a $2,500 grant for successful applicants and higher probability of winning a second stage grant with significantly more money available. Boudreau and coauthors use a similar strategy as Wang, Veugelers, and Stephan to measure the novelty of these proposals, looking for the number of novel combinations of scientific keywords (again, from the MeSH lexicon) attached to the grant applications. While a little novelty is desirable, overall reviewers looked less favorably on more novel proposals.

Where does anti-novelty bias come from?

Scientists are curious folk; why would they be biased against novel ideas?

Li (2017) and Boudreau et al. (2016) both find some evidence that this bias stems less from malice than from a greater ability of reviewers to identify high quality proposals when they are closer to their existing knowledge. In other words, when proposals are less novel and closer to existing work, reviewers have an easier time separating the really good proposals from the mediocre and poor ones. But when proposals are further from existing work, it’s hard for reviewers to separate the good from the bad, and all such proposals are more likely to be judged as just average. But when grant making is really competitive, only applications judged as being really good might be funded, leaving the hard-to-evaluate novel proposals persistently underfunded.

To show reviewers are better able to identify the quality of less novel papers, Li (2017) and Boudreau et al. (2016) need some kind of “objective” measure of the quality of grant proposals. This isn’t possible, but each paper uses an interesting proxy. In Boudreau et al. (2016)’s experiment with grant-making, they created a way to measure the reviewer’s intellectual “distance” from the proposal, based on the MeSH keywords describing the proposal and MeSH keywords attached to the reviewer’s prior published work. Then, rather than assigning each grant application to the reviewers with the most relevant prior expertise, they randomized who the reviewers were. Each proposal was reviewed by 15 reviewers, and Boudreau and coauthors compare the score assigned by the reviewer who is “closest” to the proposal (i.e., has the most relevant domain expertise) with the average of all the other reviewers. In general, they find there was more disagreement between these expert reviewers and everyone else about which proposals should get high scores rather than low ones. Those with close expertise seemed to know “something” that led them to more divergent conclusions about high scoring proposals.

I think what might be going on is even more clearly illustrated in Li (2017). Li exploits an unusual feature of the grant process to try and find an “objective” measure of NIH grant proposal quality. In the process of preparing an NIH grant proposal, researchers do a lot of work that is publishable. Indeed, Li argues for the period she studies, the arms race in NIH application quality had progressed to the point where the majority of the application was actually based on work that was nearly completed and ready for publication. This means a lot of the material that was in the grant application is already barreling towards publication, whether or not the proposal is funded (presumably the proposal builds on and extends this work). Li tries to identify this proposal material that is spun out into academic articles by looking at work from the lab that is published in a short interval after the proposal is submitted but before new experiments could plausibly have been completed, and which is about the same topic as the grant application (as judged by shared keywords). She then looks at how well this “spinoff” work gets cited. This provides a proxy for the quality of the grant applications, both those that are ultimately funded and those that are not.

The takeaway is well summarized in the following figure. On the vertical axis, we have the probability an application gets funded, after taking into account the applicant’s prior record (how much they’ve been cited, how many grants they’ve won, their demographic and educational background, and so on). On the horizontal axis we’ve got an estimate of the number of citations spinoff work from the grant application receives. The scatterplot drifts up and to the right, indicating that grants that are higher quality (i.e., spinoff publications get more citations) are more likely to get funded. But Li divides these into two groups. In red, we have applications that have also been cited by a permanent member of the grant review committee, and in blue we have ones that weren’t. Note for the red dots there is a stronger relationship between quality and funding than for the blue. If the reviewers have previously cited your work, they are better at discerning high and low quality proposals, compared to when they have not cited it.

One way to interpret this is that committee members are better able to judge quality for work that is closely related to their own work. If you work on similar stuff as a committee member, you are more likely to be cited by them. If you go on to propose a bad research project, then it doesn’t really help you that the committee member has previously cited you. You get rejected, and your spinoff publications receive few citations. On the other hand, if you propose a great idea that will garner lots of citations, it’s probability of getting funded isn’t much better than your bad idea if no one on the committee is familiar with your work, but it is much more likely to be funded if they are.

Uncertainty + Competition = Novelty Penalty?

To sum up, while I don’t doubt scientists have biases towards their own chosen fields of study (how could they not? They chose to study it because they liked it!), I suspect at least part of the bias against new ideas comes from a more subtle process. When assessing a new idea, we can make more confident assessments when the idea is closer to our own expertise. All else equal, this gives an edge to less novel ideas when we’re ranking ideas from most to least promising. But in a world with scarce scientific resources (whether those be funds or attention), we can’t give resources to every idea that merits them. Instead, we start at the top and work our way down. But that can mean valuable novel ideas are less likely to get resources than less valuable but less novel ideas.

(As an aside, this bias against novel work might be a reason why, paradoxically, novel work is actually more likely to be highly cited. If you know there’s a bias against novelty, then you also might believe it’s not worth working on novel stuff unless it has a decent shot of ultimately being really important)

(As another aside, the anti-novelty bias is probably an additional reason we should fund more R&D; maybe it’ll increase out willingness to fund novel stuff when competition for resources isn’t so fierce. End asides.)

To close, I think we can see something like this dynamic in Azoulay, Fons-Rosen, and Graff-Zivin’s study of the impact of superstar deaths. Recall again, that it’s probably not the case that superstar’s are blocking rival research by personally denying grants or publications, since they weren’t actually sitting in positions of authority when they died in most cases. In the paper, they actually go through the effort of dropping everyone in one of these positions of authority from the analysis, in order to show it doesn’t drive their results. But even if they aren’t personally blocking research, their ideas might be.

For one, the ability of researchers to enter with new ideas seems weaker when the superstar’s former collaborators retain positions of influence, sitting on more NIH grant review committees and editorships (though the latter can only be assessed really imperfectly). Second, Azoulay, Fons-Rosen, and Graff-Zivin provide some suggestive evidence that when the microfield has more firmly consolidated itself around the ideas promoted by the superstar, it remains hard for newcomers to enter even after the superstar has passed away. They attempt to measure this by looking at measures of closely connected coauthorship networks are, how intensely the field cites its own work, and how similar papers in the field are to each other according to an algorithm.

But when a field is not highly consolidated around the superstar’s ideas and there is still some active debate about the way a microfield might go, there appears to be a big effect on having a superstar active in the field prematurely pass away. After the superstar researcher passes away, clearly they aren’t around anymore to publish, thereby expanding and clarifying the idea. And indeed, the entry of newcomers is strongest in fields where the superstar had an outsized role in the microfield, as measured by the share of publications, citations, and research funding they garnered. But the death of a superstar has additional knock-on effects that might further shake the dominance of their ideas. Even the former collaborators of the superstar publish significantly less work after the superstar dies. And this effect is stronger for collaborators who work on more similar topics as the superstar.

Taken together, I think it suggests that when a superstar prematurely dies, their ideas may fade in importance and salience, as compared to fields where the superstar stuck around. That could create space for alternative approaches to get resources.

To close, I want to be sure and caution against a sort of fatalism. This isn’t absolute, it’s probabilistic. Novel ideas to get published, they can attract attention, and old paradigms do fall. It just happens less often than we might like. Research has an inertia that is real, but not insurmountable.

Thanks for reading! If you liked this, you might also enjoy this article from early 2021 talking about how these forces of conservatism can be overcome, using the credibility revolution in economics (for which another set of Nobel prizes was just awarded) as an illustrative example: How a field fixes itself - the applied turn in economics

For other related readings, follow the links listed at the bottom of the article’s page on New Things Under the Sun (.com).

Subscribe at mattsclancy.substack.com

View Details

Hello! Here’s what’s new this week on New Things Under the Sun:

New Articles

I’ve put up a new article titled Science as a map of unfamiliar terrain. From the introduction:

Two things seem to be true: more science leads to more technological progress, but only a minority of new technologies directly rely on science. So what kinds of technology benefit most from a scientific foundation to draw on?

One way to conceptualize the difference between science and technology, is that scientific knowledge tells us about how the world works, while technology is about capturing and orchestrating regularities in nature to do something we think is useful. But by definition, a new technology has to step into the unknown and try something new - either relying on novel natural processes, or orchestrating existing ones in novel ways. Since science gives us knowledge about how natural processes work and interact, it can provide an imperfect map of this unknown terrain, helping inventors step wisely. We can see this is indeed the case with a set of papers, each of which looks to patents as a measure of invention, but which take varied approaches to measuring the “reliance on science” and the degree to which the terrain an inventor is exploring is “unknown.”

You can also listen to a podcast version of this article at the top of this email.

Updated articles

The article Are ideas getting harder to find because of the burden of knowledge? has been updated to include a discussion of a 2016 paper by Ajay Agrawal, Avi Goldfarb, and Florenta Teodoris. The previous version of “are ideas getting harder to find because…” presented a lot of evidence that scientists, engineers, and inventors are:

spending more time in training

working in ever larger teams

specializing more and more

This is all consistent with Ben Jones’ “burden of knowledge” hypothesis, which argues these trends are a natural by-product of the tendency for new problems to require the application of ever more knowledge to solve. The way we throw more knowledge at problems is by assembling bigger teams of narrower and more deeply trained specialists.

But the evidence presented is all correlational. It mostly shows trends over time, and it would be nice to have some quasi-experimental evidence about this theory. That’s where Agrawal, Goldfarb, and Teodoris come in:

Suppose we wanted to conduct an experiment to test Jones’ story. What would that look like? What if we could take a set of similar fields and then randomly raise the burden of knowledge in some but not others. Then, we could see if the fields with higher burdens responded by forming bigger teams, specializing, and spending more time in school. But to do an experiment like that, we would need to dump a bunch of new knowledge into some fields but not others. This isn’t easy to do in a lab. But Agrawal, Goldfarb, and Teodoris argue the collapse of the Soviet Union provides just such a quasi-experimental context.

To jump straight to the new discussion, click here.

What else?

The New York Times had a nice write-up of some research on innovation and remote work, featuring links to New Things Under the Sun (and a quote from me).

That’s all for now; thanks!

Subscribe at mattsclancy.substack.com

View Details

What kinds of things make someone decide to try and solve some problem, instead of accepting it? Economists tends to think in terms of broad costs and benefits: if the expected benefits from innovating exceed the expected costs, then a person decides to innovate. But an alternative perspective is well articulated by the economic historian Anton Howes:

The more I study the lives of British innovators, the more convinced I am that innovation is not in human nature, but is instead received. People innovate because they are inspired to do so — it is an idea that is transmitted. And when people do not innovate, it is often simply because it never occurs to them to do so. Incentives matter too, of course. But a person needs to at least have the idea of innovation — an improving mentality — before they can choose to innovate, before they can even take the costs and benefits of innovation into account.

If this is right, where does this “improving mentality” come from? People are social creatures, and often take their cues from the people around them. So one way people could obtain this “improving mentality” is if they see it modeled in other people.

I reviewed some evidence for this notion in last week’s post, “Entrepreneurship is contagious.” That article tried to show two things. First, entrepreneurs are often found in social clusters - if people have worked with entrepreneurs or lived near them, they are more likely to go on to become entrepreneurs. Second, this effect is causal, in the sense that if you expose a random person to entrepreneurs you can “infect” them with entrepreneurship.

In this piece, I want to present two more complementary strands of evidence in favor of the notion that entrepreneurs transmit to their peers the idea that “yes, even someone like you can become an entrepreneur.” Those two strands of evidence are:

Entrepreneurship transmits from peer to peer more readily when peers are similar.

The positive impact of being around entrepreneurs falls off quickly, once that idea has been planted.

Let’s start with the similarity of peers.

Entrepreneurs Just Like Me

In “Entrepreneurship is contagious” I reviewed a 2015 paper by Lindquist, Sol, and Van Praag that showed Swedish adoptees were more likely to become entrepreneurs if either their biological or adoptive parents were entrepreneurs. Moreover, the link between entrepreneurial parents and children was twice as strong for adoptive parents as it was for the biological parents of adopted children, and there was little evidence this was due to factors like inheriting the family business or access to wealth. Instead it seems to have been something the children learned from their parents.

Could it have been the idea that entrepreneurship is the kind of thing “people like us do?” One piece of evidence in favor of that interpretation is the differential effect of parental gender on their children. Adopted sons are more likely to become entrepreneurs if either parent is an entrepreneur, but the effect of fathers on sons is generally more than twice as strong as the effect of mothers. For daughters, the effect is even stronger. It turns out adopted daughters are more likely to become entrepreneurs only if their mother is an entrepreneur - adopted daughters raised by entrepreneurial fathers are no more likely to become entrepreneurs than those raised by non-entrepreneurial fathers.

This gender asymmetry has been seen in other papers as well. Rocha and Van Praag (2020) looks at people who work with an entrepreneur: the employees of startups, this time in Denmark. Everyone who joins one of these startups has some degree of exposure to the founder and therefore everyone is exposed to the “idea” of entrepreneurship. But Rocha and Van Praag focus on how similar the founder is to the employee.

As shown in the table below, as with mothers and fathers, Rocha and van Praag find female employees are more likely to subsequently go on and found a business of their own if the founder is also a woman. While the table is just the raw correlations in the data, this is one of those findings that sticks around when you use more and more sophisticated methods that control for more and more possible confounders.

Moreover, Rocha and Van Praag find the more similar the founder is, the more they seem to influence the decisions of their employees. The effect is stronger if the employee and founder are both women, and they are both mothers (or both not mothers). It’s stronger if they’re both women, and have similar ages; or if they’re both women with similar educational background; or both women from the same place of birth.

Kacperczyk (2013) finds a similar effect. She studies mutual fund managers and their decision to strike out on their own and found a hedge fund, and looks to see if this decision is influenced by exposure to other hedge fund managers who do the same. Similar to a 2010 study by Nanda and Sørensen (discussed in this post), she finds hedge fund managers are more likely to strike out on their own if they have more coworkers who have done the same.

But she also finds the same effect if more university alumni (who also work as fund managers) struck out to found their own funds in the preceding year. One reason this is notable is that, whereas it might be that fund managers interested in striking out on their own gravitate to the same work places (which would lead to a spurious correlation between coworkers starting funds), that’s less likely to be true of alumni. It would mean, for example, that people interested in starting their own hedge fund some day choose to go to the similar universities and then wait until their peers start forming funds to do the same.

For our purposes though, alumni is interesting as another example of indicating the importance of “people like me” doing entrepreneurial things. Kacperczyk also establishes the effect is much stronger for alumni of the same gender, and for alumni who went to school around the same time period.

Taken together, it’s all consistent with the idea that people take their cues about what kinds of options are available to them from people who occupy a similar social position. When people like you are entrepreneurs, it seems to be especially impactful on your future probability of becoming one yourself.

You only need an idea once

A second piece of evidence that peers activate the idea to be an entrepreneur comes from the fact that we see little evidence these peer effects work well when people probably already have the idea of being entrepreneurs. You can only get an idea once after all.

As we’ve seen, the children of entrepreneurs are themselves more likely to become entrepreneurs. Let’s suppose this is because the children of entrepreneurs are much more likely to have the idea of being an entrepreneur in their heads; that would imply further exposure to entrepreneurs doesn’t add that idea to their choice set - it’s already there. It turns out a general finding is that all of these social exposure to entrepreneurship effects are a lot weaker for the children of entrepreneurs.

In Rocha and Van Praag’s study of the impact of working with a female founder on the probability of women going on to start businesses, the effect is only statistically significant for those without an entrepreneurial mother. The daughters of entrepreneurial mothers were more likely to become entrepreneurs, but they weren’t “even more” likely to do so if they also worked for a startup headed by a female founder.

In Nanda and Sørensen’s study showing people are more likely to become entrepreneurs if they work with coworkers who have previously been entrepreneurs, the effect is half as strong for children with an entrepreneurial parent. Moreover, they also find a weaker effect for those who reside in a region where there are more entrepreneurs. In general, those with exposure to entrepreneurship outside work are much less affected by the presence or absence of entrepreneurial coworkers.

In Eesley and Wang (2017), students in a 10-week innovation and entrepreneurship class worked on a startup project with a mentor (the paper was reviewed in more detail here). Students were randomly assigned mentors who were entrepreneurs or not, and those assigned entrepreneur mentors were more likely to join startups in the two years after graduation. But only if they did not have a parent who was an entrepreneur. For the children of entrepreneurs, there was no additional impact from having an entrepreneurial mentor.

Note this is not because the children of entrepreneurs are always entrepreneurs, and therefore there is no way to further increase their chances of being entrepreneurs. Entrepreneurship remains uncommon in all cases, even for the children of entrepreneurs.

I think this idea is also consistent with one of the few studies that finds exposure to entrepreneurial peers does not increase the probability people become entrepreneurs. In a 2013 study by Lerner and Malmendier (reviewed last week) MBA students at Harvard Business School who were randomly assigned to class sections with more classmates who had previously been entrepreneurs were actually less likely to express an intention to become entrepreneurs when they graduated. This might well be because people getting an MBA at Harvard Business School do not lack the confidence and idea of becoming an entrepreneur. They are already weighing the costs and benefits of entrepreneurship, and meeting an entrepreneur classmate doesn’t add anything “new” to their choice set (instead, Lerner and Malmendier suggest it helps them avoid starting businesses that are unlikely to succeed).

The importance of ideas

To sum up, I agree with Anton Howes that innovation may not come naturally to most people (why that might be will have to be a topic for another day). As humans, we have an enormous range of courses of action we can take as we try to live our lives; too many possibilities to consider them all in fact. So, as a shortcut, we use the choices of people in similar situations as ourselves to build a circumscribed choice set. These are the possibilities whose costs and benefits we weigh when deciding how to act.

Now, we actually don’t have much evidence about this as it relates to innovation (with one exception - see postscript below). But we do have a lot of evidence related to a close cousin of innovation: entrepreneurship. Across this article and “entrepreneurship is contagious” I’ve tried to pull together four strands of evidence about this notion that entrepreneurship isn’t normally a default choice that people consider, but they instead have to learn it’s an option from their social context:

Entrepreneurs are often found in social clusters (workplaces, neighborhoods)

Quasi-random exposure to entrepreneurs increases the probability of becoming an entrepreneur

Entrepreneurial influence seems stronger when entrepreneurial peers occupy a more similar social position

The effect of exposure to entrepreneurs is much weaker for the people most likely to already be considering a career in entrepreneurship

Taken together, I think it’s pretty compelling. In addition to the usual things that drive economic growth - institutions, macroeconomic policy, technological opportunities - we should also think about something intangible like what people regard as possibilities for their lives.

Postscript

There is a famous 2018 study by Bell and coauthors with the seemingly very relevant title “Who becomes an inventor in America: the importance of social exposure to innovation.” Despite the subtitle, I don’t think it actually sheds much light on the topic at hand. As the authors state: “the data point to mechanisms such as … access to networks that help children pursue a certain subfield, acquisition of information about certain careers, or role model effects.” My read of their paper is that their evidence is certainly consistent with peers transmitting the “idea” of innovation, but it’s hard to disentangle the importance of this channel from others. I think their best evidence is on things like the importance of networks, advice about how to succeed in a specific industry, and influencing people’s decisions about which industry to work for.

Subscribe at mattsclancy.substack.com

View Details

Is entrepreneurship contagious? Consider a few cases:

Two different teams of scientists make substantially the same discovery at roughly the same time. One of the teams goes on to found a startup based on the idea; the other does not. Why?

Marx and Hsu (2021) study basically this situation. They identify ~20,000 such “twin” discoveries by finding cases where two papers by two different teams are published within a year of each other, share a large portion of subsequent citations, and are cited at least one time in the same parenthetical block, (for example, “(Clancy 2021, Marx and Hsu 2021)”), which is an indication each of the papers can be cited without elaboration to justify the same claim. They go on to identify a couple hundred cases where these discoveries are subsequently commercialized via a startup and look to see what factors are correlated with a decision to commercialize the discovery.

One such factor is if one of the scientists who made the discovery has a history of commercialization - no surprise there. Scientists who have commercialized their research before are probably more likely to do it again. But more intriguing, they also find that if one of the scientists making the discovery has previously collaborated with a scientist with a history of commercialization, the discovery is more likely to be commercialized, even if the discovering scientist does not have a history of commercialization themselves. It’s as if the entrepreneurship bug jumped from one coauthor to another.

We can broaden this line of inquiry to workplaces in general with Nanda and Sørensen (2010). Nanda and Sørensen track approximately 270,000 workers in Denmark over the period 1990-1997, each of whom has no prior history of entrepreneurship before 1990, and were newly hired by an established employer in 1990. For each of these individuals, they count how many of the individual’s co-workers were entrepreneurs during the previous five years (and for how long). They then see how the presence of formerly entrepreneurial coworkers affects the subsequent decision of an individual to start a business.

Again, people whose coworkers have a history of entrepreneurship are more likely to go on to be entrepreneurs themselves. One weakness of this data is that Nanda and Sørensen only know if people worked at the same establishment; they don’t know if the coworkers ever had any actual interaction. But it’s a reasonable guess that people were more likely to interact with their coworkers when they both worked at small establishments. And as the figure below indicates, the impact of entrepreneurial peers is greatest in the settings where the probability of an interaction is highest.

We can broaden this to whole communities as well. Giannetti and Simonov (2009) look at how the rate of entrepreneurship in an individual’s local neighborhood affects their decision to become an entrepreneur. Using a sample of 289 Swedish municipalities, Giannetti and Simonov show residents are more likely to become new entrepreneurs (this year) when there were more existing entrepreneurs in the neighborhood (last year). This is also true for people who did not grow up in the neighborhood, but only moved there later in life.

As with Nanda and Sørensen, a weakness in this data is that we only know where people live; we don’t actually know if people residing in the same neighborhoods interact with each other (though Giannetti and Simonov cite studies showing the bulk of interactions are, in fact highly local). But Giannetti and Simonov also make the plausible suggestion that it is more likely people interact with neighbors who have a similar educational background. If that’s so, then if entrepreneurship “spreads” through social contact, the extent of entrepreneurship among residents of the local community who have the same educational background as an individual should predict whether that person subsequently becomes an entrepreneur. The extent of entrepreneurship among locals with a different educational background shouldn’t. And indeed, when Giannetti and Simonov partition their measure of local entrepreneurship into “same education background” and “different education background” categories, only the former is associated with someone becoming a new entrepreneur.

In all three cases - scientists, workplaces, local communities - did people catch the entrepreneurship bug from their peers?

Do entrepreneurs find each other or create each other?

The big challenge in this literature is establishing that this association is causal; that interacting with entrepreneurs makes people more likely to become entrepreneurs. It could, for example, be merely that aspiring entrepreneurs seek out actual entrepreneurs to be around. Maybe scientists who are interested in someday commercializing their research are drawn to work with serial entrepreneurs; some kinds of companies are appealing both to people aspiring to start a business and people who have started a business; and aspiring entrepreneurs move to places where entrepreneurship is common, in the same way aspiring actors move to Hollywood. In any of those situations, the correlation between “social exposure” to entrepreneurship and subsequent entrepreneurship could be spurious. It’s not that entrepreneurship is contagious, it’s merely that entrepreneurial types cluster.

This is not a problem these papers are unaware of. Marx and Hsu (2021) don’t present their results as causal. Nanda and Sørensen, and Giannetti and Simonov both try a bunch of different tricks to reduce the likelihood the results can be explained by merely by entrepreneurial types preferring to cluster together. To take one example, Nanda and Sørensen show their results hold even if you control for the number of entrepreneurs working at a company at the time an employee decided to work there (which might have influenced the decision to seek a job there) and instead focus just on new people who were hired after the focal individual started. They find the same result, but even so, you could imagine there is something about the workplace that is drawing entrepreneurial types to work there. But it is hard to do this in an airtight way.

One way to rule out the possibility that people predisposed to become entrepreneurs seek each other out is to look at settings where people form relationships without having any say in the matter. One such setting is family. Lindquist, Sol, and Van Praag (2015) study the career decisions of Swedish children born before 1970 to parents born after 1920. They find the children of entrepreneurs are about 12 percentage points more likely to be entrepreneurs at some point in their life than the children of non-entrepreneurs.

Obviously children can’t choose their parents. But with families, we also have to worry about a lot of other possible confounding factors. Specifically, with parents and children we might think this is all about genes. Perhaps the children of entrepreneurs merely inherit their parents’ taste for risk and autonomy, and it’s these shared preferences that lead both parents and their children to disproportionately choose to be entrepreneurs.

The cool thing about Lindquist, Sol, and Van Praag’s paper is that they have very good data on nearly 4,000 Swedish adoptees, which lets them parcel out the effects of genes (which derive from the birth parents) and the effects of being raised by an entrepreneur (which derive from the adoptive parents). Both matter. Adoptees with an entrepreneurial birth parent are about 4 percentage points more likely to be entrepreneurs at some point than those without. But children whose adoptive parent is an entrepreneur are about 8 percentage points more likely to become entrepreneurs themselves.

So genes seems to be part of it - maybe 1/3 - but the rest is coming from being raised by an entrepreneur. But that still doesn’t necessarily mean entrepreneurship is “contagious.” Maybe the children of entrepreneurs are more likely to become entrepreneurs simply because they have a family business to inherit or access to cheap loans from mom or dad. Lindquist, Sol, and Van Praag look a bit at all these options and try to rule them out in various ways (for example, they show inheriting the family business is quite rare and that the wealth of parents doesn’t affect the decision to be an entrepreneur, which suggests it’s not really about access to cheap loans). But we can also look at other papers where the social influence channel might matter, but things like inheritance or cheap loans don’t.

Eesley and Wang (2017) use an experiment to see what happens when people get close contact with an entrepreneur. One of the authors teaches a 10-week university class on innovation and entrepreneurship. As part of the class, students work in small teams with a pair of mentors from industry on a startup project. Randomly, some teams are matched with mentors who are entrepreneurs and others are matched with mentors who are not.

Eesley and Wang track these students for two years after graduating to see if they found or join a small startup. They then compare the rate at which students who were paired with entrepreneurs become entrepreneurs themselves with the control group of students paired with non-entrepreneurs. They find students who were randomly assigned an entrepreneur mentor founded or joined a startup 37% of the time, compared to 28% for those who were randomly assigned a non-entrepreneur mentor.

Are these numbers plausible? Compared to spending an entire childhood with a parent, in this case, the students spent on average just 5-7 hours in collaboration with their mentor. But it seems likely those collaborative hours meant a lot to the students; they self-assessed the mentoring portion of the class as one of the most important and they are part of a population who is curious about entrepreneurship (since they chose to take the class), working with a mentor on an industry they are also interested in.

OK, so those are two cases where people have no say in the decision to form a relationship with an entrepreneur. That’s rare in real life. But it is the case in real life that we may form associations with people for reasons that are quite incidental to their entrepreneurship. Let’s look at one of those studies next.

Azoulay, Stuart, and Liu (2017) look at elite academic life scientists, and study whether being mentored by a scientist who commercializes their research (by getting a patent) leads their postdocs to do the same. As you might expect, given the results from Marx and Hsu above, it does. But Azoulay, Stuart, and Liu’s paper is much more about trying to nail down the argument that this effect is not driven by entrepreneurial scientists seeking each other out.

To establish that students do not select postdoc mentors based on the commercial orientation of the advisor, Azoulay, Stuart, and Liu focus their study on academic life scientists who are selected to be Pew scholars or Searle scholars up through the year 2000. One reason to focus on this group is the existence of the Pew Scholar Oral History and Archives, a set of oral life histories available for 200 Pew scholars. The authors read a sample of 62 such histories (each is long; 100-400 pages) to see what kinds of factors Pew scholars self report as being important in their decision of which postdoc mentor to work with. The overwhelmingly most important factor cited was the scientific topic being investigated, followed by geography (where the lab was), the advisor’s prestige in the field, and interpersonal rapport. None mentioned the commercial orientation of the advisor, or their interest in patenting. And this wasn’t simply because they were shy to talk about non-academic goals; when asked about their own patents, interviewees were apparently quite candid.

Azoulay, Stuart, and Liu use this qualitative analysis to form the basis of some additional quantitative exercises. They come up with measures of scientific similarity, geographical proximity, and prestige, which they use to derive statistical models of the matching process between postdocs and mentors. They can then see if matches that are poorly explained by these stated factors seem to be unusually correlated with the decision to patent, which would be evidence that people left their true motivations - a desire to work with a scientist who patents - unstated. But they don’t really find any evidence of this. The statistics back up what the scholars say: recent graduates don’t really think about patenting when deciding who to work with for their postdocs. But if they “accidentally” end up working with an advisor with a history of patenting, they’re more likely to patent themselves, later in their career.

Not so fast

So far, that’s all pretty consistent with the thesis that entrepreneurship is contagious. Entrepreneurs are found in clusters, at least partially because entrepreneurs exert an influence over their otherwise non-entrepreneurial peers. You can see this when you randomly match students to entrepreneurs, when top scientists get incidentally matched with more entrepreneurial advisors, and when children get placed in the care of entrepreneurial parents. In all three cases, the exposed are more likely to become entrepreneurs themselves, than people who were not exposed.

But to close, let’s look at one more study, whose results muddy or challenge the above. Lerner and Malmendier (2013) exploit another natural experiment, where some groups are more exposed to entrepreneurs than others. In their case, they look at nearly 6,000 students who attended Harvard Business School over 1997-2004. At Harvard Business School, the administration breaks each incoming class into sections of 80-95 students. Students take all their classes with this cohort, and they typically form strong social bonds with other students in their section. Importantly, while the assignment of students to different sections is not random (the school endeavors to give each section a roughly representative cross section of various backgrounds), students have no say in the matter and their prior experience with entrepreneurship is not a factor considered by the school. That means there is some variation across sections; in some sections 0% of the students have a history of entrepreneurship and in others more than 10% of the students do.

Within each section, they then look to see what share of students who were not previously entrepreneurs, state their post-graduation plan is to found a company (or continue working on one they founded during school). In a simple model where entrepreneurship is contagious, we would expect to see more students opt to found companies if they were in a section where more of their peers were entrepreneurs in the past. In fact, the opposite is true!

In the figure below, each dot represents a section of students at Harvard Business School. The horizontal axis is the share of students in the section with prior experience as an entrepreneur. The vertical axis is the share of remaining students (those without prior entrepreneurship experience) who leave as entrepreneurs. The negative relationship is pretty clear.

What’s going on here? One obvious response might be that the kinds of entrepreneurs who quit their company and go to get an MBA are those who weren’t very successful. Maybe that dissuaded their peers?

But no. Actually, the kinds of entrepreneurs who get into Harvard Business School were pretty good at what they did. Significantly better than the population at large, frequently selling their companies for a large profit before going to get an MBA. Moreover, separating out exposure to “successful” entrepreneurs and “failed” entrepreneurs doesn’t really change the above results.

Lerner and Malmendier instead provide a variety of evidence that entrepreneurs attending Harvard Business School serve to discourage their peers from founding businesses that are unlikely to succeed. It turns out the high levels of entrepreneurship seen in sections with few entrepreneurs consist most often of “bad” entrepreneurs. Their businesses fail more often. Lerner and Malmendier suggest, in this case, the primary effect of having entrepreneurs around is they prevent inexperienced students from forming doomed businesses. Those students, instead, seem to slot into alternative non-entrepreneurial careers.

So why doesn’t this happen in all the other cases? Well, probably the best answer is “more research is needed.” We just don’t know.

But I’ll hazard a guess. I suspect entrepreneurship is only “contagious” to those who normally wouldn’t consider it. Maybe, for most people, starting a new business is literally unthinkable, in the sense that they just don’t think of doing it. But being around someone who has done it plants the seed in your mind that it’s a possibility, something you really could do. For most of the studies, the population exposed to entrepreneurship is a population that wouldn’t normally consider it. For them, exposure has a measurable positive effect.

But students getting an MBA at Harvard Business School aren’t typical students; maybe they are a group that is constantly exposed to the idea of starting a business from alumni and teachers, with the self-assured confidence that they are the kind of people who could successfully start a business, if that’s what they wanted to do. This group doesn’t need an injection of confidence that “yes, you too could form a business!” Instead, what has the biggest impact on them is a trusted peer telling them, actually their idea is kind of dumb.

Next time, we’ll look at some other evidence that suggests this role modeling effect is important.

Subscribe at mattsclancy.substack.com

View Details

Note: if you’re a regular reader of New Things of the Sun and have a few minutes, please fill out this short survey. Thanks!

It’s long been assumed that the best sorts of innovation happen when smart people work in an environment where spontaneous face-to-face interaction is the norm. Importantly, if that’s true, it implies the widespread transition to more remote work - where spontaneous face-to-face interaction is not possible - poses a threat to innovation. I’ve written about these ideas a lot before, but this week, I want to look at a case study for a sector that:

engages in frontier knowledge work

has strong incentives to adopt practices that produce better outcomes

has been well studied

has increasingly moved to a model of remote collaboration

I am talking, of course, about academia.

Remote Collaboration in Academia

Academic research is a sector where knowledge workers try to innovate - the whole game is trying to push the knowledge frontier outward. Whether or not the system could work better, it certainly does seem to work, generating new and useful knowledge pretty much every day. It’s also a system that’s highly competitive, where thousands of individuals compete with each other for jobs and space in journals.

And yet, despite strong incentives to use any possible edge to generate new and better research, academics are increasingly forgoing the option to work with their local colleagues.

Agrawal, McHale, and Oettl (2015) is a study of the changing nature of collaboration in evolutionary biology. They find the the number of distinct institutions represented on evolutionary biology papers has steadily increased from 1.4 to 2.4 over 1980-2005, while the average distance between coauthors on papers has risen from 350 to 550 miles over the same period.

Evolutionary biology isn’t some weird outlier. Freeman, Ganguli, and Murciano-Goroff (2015) look at publications in particle and field physics, nanoscience and nanotechnology, and biotechnology and applied microbiology over 1990-2010 and calculate the share of articles by authors who reside in the same city (not necessarily the same university). The share of these colocated articles dropped from over 50% in 1990 to 40% by 2000, where it stayed until 2010 or so. The majority of papers in these fields are written by teams of coauthors who are not local.

Lastly, in Clancy (2020) I pulled data from the National Science Foundation’s science and engineering indicators on the share of US journal articles in Scopus with authors belonging to more than one institution, or more than one country. The results are pretty consistent with Freeman, Ganguli, and Murciano-Goroff (2015): the share of publication done at different universities rose through the early 2000s and then had a decade long “pause” in the trend towards more remote collaboration during the 2000s. Since then though, the trend has been back up, with 76% of papers coauthored by people at different institutions, or just 24% all belonging to the same institution. Maybe we are worried that different institutions might sometimes be close enough to enable frequent face-to-face interactions though? But if you look at international collaborations, the trend is the same.

All told, the share of papers produced by academic teams primarily engaged in remote collaboration - rather than primarily local face-to-face interaction - has been on the rise for decades, and now accounts for the majority of papers.

Are we sure face-to-face matters anymore?

A set of related papers can help us see why this change is underway.

Probably the strongest evidence for benefits from being able to collaborate locally is Agrawal, McHale, and Oettl (2017), another paper studying evolutionary biology. In this one, the authors look at what happens to the research productivity of a department when they successfully recruit a star researcher to join the department (“stars” are defined to be researchers whose citation-weighted publication count is in the top 10%). In one exercise, they take the set of academics working in the department before it gets a star hire, and divides them into two groups: those who have previously cited the star’s work (“related” scientists) and those who haven’t. The productivity (annual citation-weighted publications) of scientists working on related stuff goes up 96% in the years after a star is hired!

Now, a potential objection to this paper is that star researchers don’t just get randomly allocated to different departments. It could be that departments that are already “on-the-up” are the ones that attract stars. Maybe, for example, a department has just secured a big endowment that lets them go on a hiring binge and provide better resources to everyone in the departments. In that case, we might falsely attribute higher research productivity that is actually due to more resources to the appearance of a star researcher. But if that was the case, we would expect to see research productivity for the department to rise for everyone around when a star is hired - and we don’t. Instead, scientists unrelated to the star see no impact from the star’s arrival. (Agrawal, Oettl, and McHale also try to answer this concern in a few other ways, but don’t find anything that alters their core conclusion).

So this result is a bit puzzling, given the trend towards more distributed academic teams. It also contradicts a few other studies. Let’s look at this studies and come back to this one.

Dubois, Rochet, and Schlenker (2014) looks at the same time period for the set of all mathematicians (in the world) who published at least two articles in a mainstream math journal between 1984 and 2006. Rather than look at the impact of an individual researcher on the productivity of a department though, it looks at the impact of a department on the productivity of an individual. By keeping track of the moves of individual mathematicians, they can see how the productivity of researchers changes when they move to top universities. If being able to have frequent face-to-face interactions with other great researchers is really important for producing good research, then we ought to see the research productivity of mathematicians jump up when they move to a department full of great researchers.

While it is true that top departments have more productive faculty, it turns out this is just because they hire more productive faculty. Once you take into account the productivity of a researcher (i.e., from before they get hired to a top department), the authors don’t really find any increase in productivity when these people move to top departments. It seems that the ability to have lots of face-to-face interactions with other top mathematicians is not doing much to raise their productivity (which was already high).

Ultimately, these are both observational studies, and so we might be a bit worried that something funny is going on to deliver these results. In an ideal experiment, we would randomly shuffle around researchers, looking at the impact on departments that randomly receive (or lose) a researcher to those that do not. We can’t do that, but two other papers identify some (grim) natural experiments where faculty in some department suddenly lose the opportunity to have frequent face-to-face interactions with a top researcher. In all cases, there doesn’t seem to be much value to it.

Waldinger (2012) looks at German physics, chemistry, and mathematics departments over 1925-1938. In 1933, there was an awful event that led to some departments losing colleagues and other departments not, for reasons quite unlikely to be correlated with factors driving academic productivity: the Nazi dismissal of Jewish and politically “undesirable” scientists. Like Agrawal, Oettl, and McHale, Waldinger looks at the research productivity of the faculty who were left behind. And, surprisingly, he finds no impact. The subsequent research trajectory of faculty in departments that lost local colleagues were no different than departments that didn’t lose anyone. (This isn’t to say there was no effect; as he documented in Waldinger 2016 - departments that dismissed faculty saw total output drop, and were never able to recover by hiring replacement faculty with the same research output).

Azoulay, Graff Zivin, and Wang (2010) are also interested in the impact of losing a colleague: they use sudden and unexpected deaths as a plausibly random shock to a department. Azoulay and coauthors identify 10,349 elite life scientists, as well as their coauthors. They then further identify a subsample of 112 eminent life scientists who died suddenly and unexpectedly between 1979 and 2003 before the age of 67, and who were still active researchers at the time of their death. It is the coauthors of these eminent scientists that they study. They match each scientist who experienced a sudden and unexpected death of an eminent coauthor to a similar scientist working with another eminent coauthor who did not die.

They find the unexpected death of a colleague reduces output by 5-8% relative to those who did not. Of this group, about 12% of coauthors were colocated with the deceased at the time of their death. But when Azoulay, Graff Zivin, and Wang look to see if this group is more adversely affected by the death than distant coauthors, they find just a small effect (in the opposite direction) that can’t be statistically distinguished from zero. Losing a coauthor hurts; but in terms of it’s impact on measurable research outputs (setting aside the real human cost), it doesn’t look like coauthors able to engage in frequent face-to-face interaction were any more affected than those who couldn’t.

So we have four studies now. One finds that evolutionary biologists benefit (a lot) when a talented colleague working on similar stuff joins the department. The other three find local access to a colleague, however, doesn’t seem to have much impact. And we actually know that long-distance collaboration is becoming the norm in evolutionary biology, precisely the field where another study documents strong benefits to a local star researcher. How can we square these results?

I think there are two things going on. First, I suspect proximity might be quite important for forming working relationships, but less important for maintaining them. Second, the importance of local colleagues is declining over time, and most of the results on evolutionary biology are driven by hires in an era when the effects of locality were stronger than they are today. Let’s start with the former.

Getting to know you

Suppose local face-to-face interactions are important for forming new relationships but less important for doing productive work once a relationship has been formed. In that case, we would expect hiring a new star to boost the output of local researchers working on related topics, because they could form a new collaborative partnership. And that’s basically what Agrawal, Oettl, and McHale (2017) find. Most of the increased output of evolutionary biologists when they hire a new star research seems to stem from publications they coauthored with the star.

Conversely, if the opportunity to collaborate disappears, everyone who would normally collaborate with them - near and far - would be equally adversely affected. That’s essentially what Waldinger (2012) and Azoulay, Graff Zivin, and Wang (2010) find. Waldinger notes that coauthorship across departments was already pretty common in Germany before 1933 because Germany had a well integrated national research community with frequent national conferences and mail correspondence about scientific research. Azoulay, Graff Zivin, and Wang’s study focuses on people who were already coauthors with the deceased - definitionally a group that has likely already cleared the threshold of “forming a relationship” - and no longer needs proximity to work together productively. So when the opportunity to collaborate disappears, due to one tragedy or another, they’re equally negatively impacted, whether they are in the department or not.

We’ve got some complementary evidence that colocation is important for forming academic working relationships. In a figure I’ve reproduced before, Freeman, Ganguli, and Murciano-Goroff illustrate how the academic collaborators in their study first met:

Look at how non-colocated and international collaborators first met: 90% of the time, it was during some kind of face-to-face interaction.

Change over time

In Agrawal, McHale, and Oettl (2017), most of the new hires of star researchers that they observe occur prior to 1995. That might be important, because we have a number of papers that suggest the benefits of local face-to-face interactions where stronger prior to the year 2000.

The most direct test of this is Kim, Morse, and Zingales (2009), which examines the research productivity of academics in economics and finance over 1970-2000. Like Dubois, Rochet, and Schlenker (2014), they’re going to try and identify the productivity bump a researcher gets from moving to a top department by tracking the moves of individual academics over their careers. In this case, research productivity is measured in terms of how many articles an economist/finance professor writes per year, adjusted for the quality of journal. They actually do find that moving to a top department helps researchers be more productive - in the 1970s. But the productivity bonus researchers got for moving to a top 25 university dropped in the 1980s, and had disappeared by the 1990s.

Two other recent papers zero in on the flow of knowledge across space, as measured by citations. Head, Li, and Minondo (2019) looks at the probability that there is a citation between two different mathematics papers. In the figure below, they plot how the elasticity of distance changes over time - roughly speaking, the more negative, the less likely a citation between math papers by authors who are far apart.

In the left figure, in orange, we see the distance penalty has dramatically shrunk between 1990 and 2010. That’s the main finding of interest for today - being close matters less and less, in terms of the probability two researchers know about each other’s work. (See this post for a bit more discussion of this paper). In the right figure, the authors break this out for citations between US papers (in orange) and citations between US and non-US countries. Surprisingly (to me), there has actually never been a strong penalty for distant papers in the USA, at least since the 1990s. But the probability a paper from abroad is cited used to be substantially affected by distance and no longer is.

Hellmanzik and Kuld (2020) perform a similar exercise for economics, but focused entirely on international citations. They begin with a gravity model, which is traditionally used to statistically model the flow of exports and imports between countries, but then adapt it to model the flow of citations between economics papers written by authors in different countries. After controlling for a variety of factors, they find papers in 1979-1988 were nearly 6 times as likely to cite papers from their own country as compared to international ones, but by 2009-2016 they were just 1.4 times as likely to papers from their own country. Measures of the “distance” penalty, wherein countries that are farther away are even less likely to be cited also declined over the same time period.

Hellmanzik and Kuld also introduce a new measure that helps account for why some countries are more likely to cite papers from each other (again, after adjusting for lots of other potentially important stuff like language): the number of hyperlinks from one country to another. As you might expect, they find countries that are more interconnected by the internet are more likely to cite each other’s work, and that this effect has grown stronger over time.

Taken together, I read this as providing a different but complementary line of evidence that local interaction is not as important as it used to be for doing frontier knowledge creation.

The Upsides of Remote Collaboration

So I think proximity was (and perhaps still is) important for forming collaborative working relationships, but these relationships remain productive even when academics are far away from each other. At least, this is increasingly the case.

But at the same time, I think there’s another factor that increasingly pushes academics to collaborate remotely. That is the ever-rising set of knowledge needed to push the frontier, which academics are responding to by drilling into ever narrower specialties and forming teams of specialists to tackle research projects. When academics specialize more and more, it becomes ever less likely that the specialty you need to complete a project happens to reside in the same department.

We’ve got a few pieces of evidence in support of this view.

First, when Freeman, Ganguli, and Murciano-Goroff (2015) survey coauthors about “what they bring” to a collaboration, the most common response (80%) for non-colocated authors was “unique knowledge, expertise, or capabilities”, beating out responses like access to data, funding, or specialized equipment, none of which cleared 60%.

Azoulay, Graff Zivin, and Wang (2010) provide some similar evidence. They perform a series of exercises to try and sniff out what kinds of factors exacerbate the decrease in research productivity that comes when an eminent coauthor passes away. For example, they look to see if the negative impact of a coauthor’s death is stronger when the coauthor has plausibly better access to NIH grant funding. It isn’t. They look to see if it’s stronger when the coauthor has better potential to network in the field. Again, it isn’t.

Lastly, they look to see if the negative effect of a coauthors death is stronger when the coauthor works on more similar topics. This time, they find a robust correlation. That again suggests these collaborative partnerships are formed as a way to acquire the knowledge necessary to tackle hard problems, rather than as a way to share grant funding or curry favor in the scientific hierarchy.

Agrawal, McHale, and Oettl (2017) also get some results that I think also underscore how specialized science is becoming. Recall this is the paper that found evolutionary biologists benefit a lot if they hire a new star researcher, but only if they’ve previously cited their work. It turns out, this isn’t very common: just 9% of the faculty in a typical department are related to the star in this way. The other 91% of scientists in the field of evolutionary biology apparently work on subfields too dissimilar to benefit from the opportunity to collaborate with a star in evolutionary biology.

Agrawal, Oettl, and McHale (2015) also find some additional suggestive evidence consistent with this story. They find that, over time, stars have been producing a growing share of all citation-weighted work. They argue this is because stars build larger pools of potential collaborators (for example, by training more doctoral students) which becomes increasingly important as it gets more important to find partners with just the right specialty. A bigger pool to draw on also lets stars be more selective. Their paper provides evidence that stars do indeed produce papers with a larger set of unique coauthors, and that they have grown more selective over time.

Two cheers for dispersed teams!

In sum then, my story is basically this.

it’s not actually that hard to collaborate productively at a distance in academia, at least once you’ve gotten to know someone.

innovation requires ever more collaboration among specialists as knowledge accumulates.

Over time, falling travel and communication costs have increasingly favored building those teams by turning to remote colleagues with the right specialization.

I think these trends are a harbinger for things to come for the non-academic world (to some extent it’s already happening). Innovation by distributed teams will probably become more and more important as access to just the right specialty becomes ever more important and falling communication costs make it ever easier to do this well. To the extent innovation via dispersed teams has been shown to work pretty well in academia, that’s a reason not to be too concerned about it in other domains.

But there are caveats.

First, the nature of academic work might be particularly well-suited to innovation by distributed teams, as compared to other forms non-academic innovation. Projects are discrete, and finite, and many forms of academic projects have a modular character - one person can do the experiment, another can analyze the data, and so on. They also ultimately hinge on the creation of a digital artifact - the paper - rather than a physical prototype (although in many cases there is a lot of work that happens in the physical world on the way to a paper). Still, it might be that innovation in hardware, for example, doesn’t work nearly so well with a distributed team.

Second, while I think it’s quite feasible to productively work together at a distance, I suspect that academia still relies heavily on face-to-face interaction for one crucial step in the innovation process: meeting new people. If we’re thinking about applying the lessons of academia to non-academic examples of innovation by distributed teams, much hinges on whether colleagues at fully remote organizations get to know each other as well as they would if they were all in the office. As best as I can tell, this is certainly possible, but some deliberate planning is often necessary.

Besides the departmental work-unit, academia has evolved a rich set of mechanisms to circulate researchers and give them opportunities to form relationships with people who are unlikely to be colocated colleagues for the majority of a career. On the longer side, these include training - often in one place for a PhD and another for a postdoc. But they also include the seminar circuit and frequent academic conferences. Collectively, these temporary nexuses account for another 40-50% of non-colocated collaborative relationships in Freeman, Ganguli, and Murciano-Goroff’s survey. Non-academic sectors that hope to be as effective as academia at innovation with distributed teams may need similar distant networking opportunities, whether those take the form of apprenticeships, conferences, online social networks, etc.

If you liked this post, you might also like:

Proximity: more important for meeting than collaborating?

Cities as a platform to mix up knowledge

Are ideas getting harder to find because of the burden of knowledge?

Subscribe at mattsclancy.substack.com

View Details

Science has a problem. There is an element of unavoidable randomness in research. Combine that with a propensity to disproportionately publish notable research, and you end up with two factors that distort our picture of the evidence. First, results that are not sufficiently “interesting” may remain unpublished and unknown. Second, in anticipation of this, researchers might bend their research methods to force their results to be “interesting” and hence, publishable (this practice is sometimes called p-hacking). The resulting biased presentation of evidence presents a challenge for science’s cumulative project to better understand how our world works.

On the other hand, it isn’t necessarily a bad thing that top journal seek to highlight research that challenges our existing beliefs about the world. Even if there aren’t publishing space constraints in a digital world, attention is limited. It makes sense for top journals to curate the research that is most useful to an audience with limited time. That could mean privileging research that provides evidence for some specific theory of how the world works, rather than null results that maybe don’t (Frankel and Kasy 2021 develops this idea more fully).

But not all journals need to fill that role. We could create a different set of publication outlets whose primary purpose is to provide peer review, archiving, and search engine optimization services and not to curate research for casual readers. These outlets could be a home to results not publishable in top journals. In fact, some journals like this already exist, such as the Series on Unsurprising Results in Economics (acronym: SURE). That would enable the full set of results related to a particular topic to be discoverable, with a little bit of effort. We might end up with poorly informed casual readers, who only read top journals, but inventors, policy-makers, and researchers who need to get an accurate idea of the actual state of the evidence would know to dig deeper to get the whole story.

Would that work? One place to get some evidence on this is to look at our experience with preprint servers.

Preprint Servers as a Secondary Publication Outlet?

Preprint servers are places where researchers - typically with some kind of affiliation - can post work that is nominally in progress but usually quite close to a finished product. They are not subject to peer review or editorial discretion, and so could, in principle, serve as a home for research results that don’t end up getting published, perhaps because of publication bias.

Indeed, a non-negligible share of work on preprints is never published. Baumann and Wohlrabe (2020) estimate that about 25% of working papers published on major preprint servers in economics are never published. Lariviere et al. (2014) estimate 36% of working papers on arXiv are never published in a journal listed on the Web of Science, and Tsunoda et al. (2020) estimate 59% of papers posted on bioRxiv during 2013-2019 were not (yet) published.

We have some evidence these preprint servers do help mitigate censoring. Fanelli, Costas, and Ioannidis (2017) obtain 1,910 meta-analyses drawn from all areas of science, and pull from these 33,355 datapoints from original studies. They then look at the size of estimated effects in each of these disciplines for those published in peer-reviewed journals and those published elsewhere (i.e., on preprint servers, but also in conference papers, personal communications, unpublished drafts, graduate theses, etc.). The latter group - which hasn’t been through explicit peer review - is often called “gray literature.” Fanelli, Costas, and Ioannidis find gray literature articles do report smaller effect sizes than those in peer-reviewed journals. That’s consistent with papers on preprint servers facing less publication bias or pressure to engage in p-hacking.

But the effect is pretty small, explaining on the order of 1% of the variation in outcomes. That’s not too surprising, in light of a finding from Franco, Malhotra, and Simonovits (2014) I have discussed elsewhere - in their study of the social sciences, most null results were never written up at all, much less posted publicly on a preprint server. This seemed to be because the researchers believed these studies faced a hopeless path to publication, and hence weren’t worth the effort of writing up.

In fact, restricting attention to the null results that did get written up, they were published at about the same rate as positive results. One interpretation of that is sometimes null results are, in fact, interesting, and so they get written up. But in those cases, publication bias isn’t really a problem, because the results face a decent shot of getting published anyway. Indeed, preprint servers aren’t really intended to be an archive for work that can’t be published in traditional outlets; instead, they are more like a parking spot for papers that are being shopped for publication in traditional outlets. Hence, the relatively small difference between effect sizes in gray literature and journals.

Preprints and p-hacking

Franco, Malhotra, and Simonovits (2014) suggest that papers in the social sciences don’t get written up and posted to a preprint server if the results don’t look publishable. Brodeur, Cook, and Heyes (2020) provide some complementary evidence that when authors engage in p-hacking, they do it in anticipation of the challenges of publication, not in response to pushback from peer reviewers.

Brodeur, Cook, and Heyes (2020) look for the statistical fingerprints of p-hacking in economics journals versus working papers. They’re able to detect this because p-hacking leaves a different statistical fingerprint than regular publication bias. Imagine a bunch of results are plotted in a scatterplot, like in the left-most figure below, with effect sizes on the horizontal axis and the imprecision of the estimate on the vertical axis. Everything in the red triangle is not statistically significantly distinguishable from zero (the red zone widens as you go up, because as imprecision gets larger, larger and larger estimates could actually be consistent with the true effect size being zero).

In the middle figure, we illustrate the statistical footprint of publication bias: all the results in the red triangle just disappear, because they don’t get published. (I talked about how to detect his here) In the right figure, we have the statistical footprint of p-hacking. Instead of disappearing, all the effects in the red triangle get perturbed again and again until they lie outside the red triangle. This leads to a suspicious pile of results that just (barely) happen to be statistically significant and therefore publishable.

Brodeur, Cook, and Heyes (2020) look for this kind of suspicious pileup right above the conventional thresholds for statistical significance. The figures below plot the distribution of something called a z-statistic, which divides a normalized version of the effect size by an estimate of precision. A big z-statistic is associated with a precisely estimated effect that is large - those are places where we can be most confident the true effect is not actually zero. A small z-statistic is a small and very imprecisely effect size; those are places where we worry a lot that the true effect is actually zero and we’re just observing noise. The figure on the left presents z-statistics from the main results of every empirical paper published in a top 25 economics journal in 2015 and 2018. The figure on the right presents the same thing for the preprint version of each of these papers.

Notice there’s a big hump in the middle; that’s the suspicious pileup I discussed above. Lots of papers are publishing results that are just barely statistically significant. Hmmmmmmm, how convenient. But in the current discussion, the main point is that this bump is already there in the working paper stage. If we interpret this as evidence of p-hacking, it’s telling us that researchers don’t do it when reviewers complain - they do it before they even submit to reviewers.

(Aside - It’s also notable that peer review doesn’t seem to do much to dampen this bump; we might therefore conclude that peer review fails to detect p-hacking. But hold on… it could also be that the cases where peer review sniffs out p-hacking just aren’t published!)

What’s the credit for a preprint?

Now, before we draw strong conclusions about the efficacy of preprint servers for reducing publication bias, we should pause to note that a journal may provide professional credit that a preprint does not. Franco, Malhotra, and Simonovits’ result about null results never even being written up suggests it might not be worth writing up a draft merely to post it on a preprint server. But it might be worth writing up a draft if it would result in a peer reviewed journal article. You can at least list that on your CV under published and peer reviewed papers, and maybe it helps you get tenure.

Maybe that’s enough to pull null results out of the file drawer and into the public domain. But in science, a successful publication is not just one that gets published - it’s one that is influential. One proxy we can use for that is the number of citations received. And when Fanelli, Costas, and Ioannidis look at the impact of citations on bias, they find the same kind of effect as they do for publication. More cited work tends to exhibit bigger effect sizes, though again, the overall effect is pretty small.

Change methodologies or publication outlet?

Taken all together, I suspect the opportunity to publish null results somewhere won’t make a huge difference to the prevalence of p-hacking, so long as researchers continue trying to make big and bold discoveries. That kind of ambition seems likely to always provide an incentive to abandon projects that seem unlikely to deliver those kinds of results, and also to draw researchers to methods that seem to deliver them. But that doesn’t mean it’s hopeless.

Brodeur, Cook, and Heyes’ results on p-hacking in economics also found something else quite interesting. It turns out, the extent of p-hacking varies a lot by methodology. Check out the following figure, which tries to measure how the extent of p-hacking in different methods.

Whereas there’s no significant difference between preprints and publications, there are quite large differences among economic methods. Brodeur, Cook, and Heyes argue something like 16% of statistically insignificant results in papers using instrumental variables are shifted into the significant range (upper right figure), but only 1.5% of insignificant results using randomized control trials are (lower left).

That suggests methodological changes might be able to make a big dent in some of these problems. We’ll turn to one such methodological innovation - preanalysis plans - in a future newsletter.

If you liked this, you might also like:

Why is publication bias worse in some fields than others?

Publication bias is a real thing

How bad is publish-or-perish for the quality of science?

Or, subscribe:

Subscribe at mattsclancy.substack.com

View Details

One of the most influential economics of innovation papers from the last decade is “Are Ideas Getting Harder to Find” by Bloom, Jones, Van Reenen, and Webb, ultimately published in 2020 but in earlier draft circulation for years. While the paper is ostensibly concerned with testing a prediction of some economic growth models, it’s broader fame is attributable to it’s documentation of a striking fact: across varied domains, the R&D efforts necessary to eke out technological improvement keep getting higher. Let’s take a look at their evidence, as well as some complementary evidence from other papers.

Moore’s Law

Bloom and coauthors start with Moore’s law, the observation that for the last half century, the number of transistors that can fit on an integrated circuit doubles every two years.

Each doubling of circuits is the fruit of human ingenuity. What Bloom and coauthors show is that the amount of human minds we have to throw at this problem to keep up this pace of doubling keeps rising:

In the figure above, the green line with the axis on the right is annual spending on semiconductor R&D divided by the wages of researchers; effectively, it’s the number of brains you can buy to throw at this problem if you expend your entire R&D budget on hiring (hence, “effective” number of researchers). While the rate of progress has been steady at 35% per year, the effective number of researchers has grown nearly 20-fold since 1971. It gets harder and harder to achieve a doubling of chips per integrated circuit on this schedule.

Crop Yields

Next, Bloom and coauthors turn to agriculture. Agriculture is a nice setting for studying innovation because agricultural products have been mostly unchanged for decades - maybe even centuries. While a phone today is not the same thing as a phone from fifty years ago, an ear of corn today is more or less the same as an ear of corn from fifty years ago. So it’s relatively easy to measure long-run technological change in agriculture. Like Moore’s law, the increase in annual yields across major crops has been remarkably consistent for half a century, on the whole (it bounces around more than semiconductors because of the weather).

Yet once again, those gains are costing us more and more. In the figure below we have the annual growth rate of yields of different crops, in blue, and again estimates of the number of effective researchers working on the problem of increasing yields in green. The top green line is a “narrow” measure of yield growth research, focusing on the growth of R&D effort devoted to breeding or engineering better crops. The lower green line is a “broad” measure of yield growth research, adding in research on things like chemical pesticides and fertilizers.

For corn and soybeans, the pattern is just like with Moore’s law. No matter how you slice it, the scale of R&D resources devoted to improving yields has increased 6-fold (if considered broadly) or more than 20-fold (if considered narrowly), with no concomitant increase in yield growth. However, for cotton and wheat, there are long periods where R&D resources remained roughly constant, and yet were able to maintain constant yield growth. We’ll see elsewhere that this does happen some of the time.

Before we move on, I can’t help pointing out a weird parallel between agricultural yield growth and Moore’s law. Just as Moore’s law is about packing transistors more densely onto circuit boards, the growth in agricultural yields is mostly about packing plants more densely onto farmland. At least, this is the case for corn (it may well be for other crops as well, but I haven’t seen data). The figure below plots changes in corn bushels per plant (in blue), plants per acre (in orange), and bushels per acre (in grey) in the state of Iowa (where I live). While a corn plant in 1963 yielded basically the same number of bushels as a plant in 2020, yield has more than doubled because we’re now packing more than twice as many of those plants onto each acre.

Health

Next, Bloom and coauthors look at health outcomes. In this case, rather than trying to count how many researchers or research dollars are being spent on any given disease (particularly challenging in the presence of spillovers), they measure the amount of research devoted to each disease by counting the number of publications or clinical trials related to a given disease. To measure progress against each disease, they calculate life years lost due to each disease (for a population of 100,000 people) and back out life-years saved. Below, they calculate the number of years of life saved by a given measure of R&D effort (clinical trials or journal publications).

Two observations. There are periods for which a constant supply of R&D effort does generate constant improvements, or even increasing improvements. These occur mostly before 1990. But in the long run, this seems to be another area where it takes more and more R&D effort - here measured in journal publications or clinical trials - to eke out constant improvements.

Machine Learning

In a domain not covered by Bloom and coauthors, Besiroglu (2020) is an interesting master’s thesis that compares progress in machine-learning to research effort. Besiroglu computes research effort in different machine learning domains with the number of unique authors publishing papers on these topics (in Web of Science and arXiv). Progress in machine learning can be measured on a wide variety of widely accepted benchmarks, such as accuracy classifying images in a standard dataset. Broadly speaking, increased research efforts have not yielded any visible increase in the growth rate of progress.

In the figure above, we have Besiroglu’s estimate of research effort related to computer vision at left, one example of a measure of progress on computer vision (there are actually 56 different measures for computer vision) in the middle, and a statistical model-based estimate of research productivity at right. The main take-away is clear: even though R&D resources have increased by an order of magnitude (note the figure at left has a log scale), the rate of progress has not sped up, because improvement per researcher has fallen. These results also hold for natural language processing and machine learning on graphs.

Firms At Large

This is pretty suggestive, but at the end of the day it’s quantitative case studies, and with case studies we might always be a bit worried that the cases selected are unusual. So Bloom and coauthors also extend their analysis to the much larger set of all US publicly traded firms.

It’s not that hard to compute R&D effort for all these firms - divide their R&D spending by the typical wage of scientist. The trouble here is coming up with a measure of “innovation” that is consistent across different kinds of companies doing different kinds of things. Bloom and coauthors resort to some crude measures that are plausibly linked to innovation: growth in sales, market capitalization, employment, and revenue per worker. The idea here is that a more innovative firm might create better products and services or find cost efficiencies that lead to growth along all these dimensions. These measures are crude though because many things besides innovation can affect them - a non-innovating firm that enters a new market, for example, could see growth in most of these metrics. But with thousands of observations, hopefully these omitted factors are randomly distributed among innovating and non-innovating firms, over time, so that if we just look at how our measures of “innovation” per R&D worker change over time, they won’t be misleading. If this was the only data Bloom and coauthors presented it wouldn’t be very convincing, but in concert with the other data we’ve seen, it’s more compelling I think.

Comparing growth in sales per R&D worker across two consecutive decades, they compute the change in “research productivity” for all these firms. The distribution of results is below (in blue):

The main message you should take from this is that most of the blue bars lie to the left of 1. That means growth in sales per R&D worker dropped from one decade to the next, for most firms. Note that it’s not universal though. As we saw for some crops, and for some of the time with health innovation, some firms observed an increase in research productivity over two decades.

So far all of these examples have been US-specific, so one objection might that this is actually something specific to the USA. Maybe this just reflects that fact that we’re a country that’s in decline for various reasons that are unique to us? Boeing and Hünermund (2020) replicates this part of Bloom et al. (2020) for Germany and China though and get the same flavor of results.

This data is once again computing growth in sales per effective R&D worker, for a broadly representative sample of German firms (that conduct R&D) and publicly traded Chinese firms. While the decline in research productivity is pretty similar in the USA and Germany, the decline is much higher in China.

As we’ll see, this isn’t the only evidence that innovation is getting harder in more places than just the USA.

Nations and Industries

Another crude but common measure of innovation is total factor productivity. Total factor productivity - TFP for short - is statistically estimated as the amount of quality-adjusted output that can be squeezed out of a mixed set of various inputs (e.g., capital, labor, land, energy, etc.). If you invent a new process to more efficiently make the same amount of output with less inputs, that shows up as an increase in TFP. Similarly, if you invent a new and more valuable kind of product that doesn’t take more inputs to build, that can also show up as an increase in TFP. Importantly, you can compute TFP for entire industries or economies, making it a favorite measure of innovation writ large. Note, however, that these measures can also be misleading: TFP can move for reasons unconnected to innovation. But again, in concert with the other evidence, it starts to look more compelling.

Miyagawa and Ishikawa (2019) have a working paper that uses TFP to look at how research productivity has changed over 1996-2015 for a set of Japanese manufacturing industries and for manufacturing and information services overall in Japan, France, Germany, the UK, and the US. Within Japanese industries, they find a mixed bag; some industries saw TFP growth per effective researcher rise over the period and others saw a fall. Overall, there was a decline in research productivity in Japanese industries, but not a statistically significant one.

Looking more broadly at how research productivity changed in the overall manufacturing sector of various countries, they do find declining research productivity in all five countries they study. But when they look at how research productivity in the information services sector they again find a mixed bag. In Germany, for example, TFP growth per effective R&D worker was higher in 2006-2015 than in 1996-2005.

Lastly, let’s return to Bloom and coauthors. In the USA, we can estimate total factor productivity for the entire country, as well as R&D effort, going all the way back to the 1930s. When we compare those two datasets, we see the same thing, going all the way back to the beginning. For as long as we have data, it’s taken increasing effort to sustain a constant rate of TFP growth.

So, looking at the rate of technological advance across a variety of sectors - computer chips, agricultural yields, health, and machine learning - we see a strong tendency for a constant rate of advance to be only sustainable by significantly increasing research efforts. Proxies for innovation in firms, industries, and countries find the same general tendency. The march of progress needs more and more effort to sustain it. This is not a universal rule. There are exceptions in certain fields that can sometimes go on for decades. But it does seem to be a general tendency. Innovation gets harder.

To close though, let’s consider a few potential objections to all this evidence.

Objection! Mismeasurement!

One common complaint about this exercise is that the case studies focus on the wrong things. For example, agricultural crop research is about a lot more than maximizing yields. GMO technology that makes a crop resistant to pests may not increase yield much, but it makes farming more profitable by reducing the need for some pesticides. Other agricultural crop research reduces the vulnerability of crops to extreme heat or drought; in a year without drought or heat stress, you won’t see any impact of this research. If an increasing share of research is devoted to these non-yield factors, than it might be that R&D is just as productive as ever, but we’re not measuring the correct outputs of R&D effort. You could make similar claims about the other case studies.

Totally fair. More correctly measuring the goals of research with R&D effort will probably make R&D effort look more productive. But it’s hard for me to believe the effect would be big enough to change our conclusion that innovation gets harder, as a general tendency. Looking at corn, research effort is up 6-20-fold depending on how you measure it. To have constant R&D productivity, that means the share of research that leads to better yields needs to fall by an offsetting 83-95%. It wouldn’t surprise me that the share of research devoted to yield growth has fallen over this period, but I don’t believe it’s cratered to a tiny fraction of overall effort, relative to the 1970s.

The numbers don’t look great for the other case studies either. Semiconductor research is about more than cramming transistors on circuits. But has the share of research devoted to that goal really fallen from 100% of R&D effort in 1970 to 5% in 2015? That’s the kind of change that would be needed to generate these numbers when research is actually just as productive as ever.

Objection! Growth is just linear!

Second, we might ask if it’s really appropriate to expect the rate of progress to be related to the number of scientists working on something. As a counter-example, imagine we’re talking about culinary scientists inventing new recipes. Suppose culinary scientists can come up with one new recipe per year. At one recipe per year, the growth rate of recipes is 10% per year when there are just 10 recipes, but 1% per year when there are 100. The rate of progress slows down, if the number of culinary scientists doesn’t change. Yet, in this thought experiment, we can’t really say innovation has gotten harder. It’s just that progress is constant and linear, and we’re incorrectly assuming it should be constant and exponential.

This is super reasonable. In fact it’s so reasonable that this is almost exactly what economic growth models do assume! Remember, at the beginning of this post, I said that Bloom et al. (2020) was ostensibly motivated as a test of one of the predictions of some economic growth models? This is closely related to the prediction the paper is trying to test.

The main difference between our thought experiment about culinary scientists and models of economic growth, is that the models assume the absolute increase in innovations (e.g., recipes per scientist per year) is related not to the number of scientists but the level of actual real R&D resources devoted to research. Think of this as a mix of labor and capital; scientists plus lab equipment, computers, arXiv, and everything else used to support inquiry. In terms of our example, these models assume a culinary scientist plus the same set of “culinary research tools” (a test kitchen and library of cookbooks?) would generate one new recipe per year.

There is a subtle distinction between R&D resources and R&D “effort”, as I’ve called it in this post. These models assume the same level of actual R&D resources always generate the same level of innovations. But the nature of R&D resources that a given level of R&D effort uses will change over time. The resources will improve, and come to embody technological advances themselves. Intuitively, the same level of “effort” will now go further, because it will be able to draw on faster computers, access deeper knowledge, and so on.

R&D effort in this post is measured as the effective number of researchers - that is, the number of researchers you could hire if you spent all your R&D on researchers. But that’s not what people actually do. In fact, they buy a bundle of research labor and research capital. The trick is that the capital side gets cheaper or “better” as an economy grows. Under some seemingly reasonable assumptions, all these effects balance out so that the same number of researchers, supported by capital embodying steadily better knowledge, can sustain a constant rate of technological progress. That is what is predicted by these models, and that is the prediction that is apparently falsified by all this evidence.

Let me say it one more time: under some seemingly reasonable assumptions, if you write down a model of economic growth where a constant level of real R&D resources generates a constant level of innovations, that will still lead to a constant growth rate of innovation, because R&D resources improve with the overall economy. This is the prediction Bloom et al. (2020) shows does not hold.

But I think there is another, even simpler way to respond to this critique (that we should never have expected innovation to be constant and exponential with the same number of researchers, but that it should be constant and linear). The rate of progress is what most people care about because that is what we’ve become accustomed to. It’s business as usual. We expect our computers to get twice as fast every few years because that’s how it’s been in our adult lifetimes. We expect crops to yield a couple more bushels per year, because that’s how it’s been in our lifetimes. We expect healthcare to save a few more years of life, machine learning benchmarks to be notched up, and to be a few percent richer as a society, every year, because that’s what we’re accustomed to. And what this line of work shows is that sustaining that business as usual requires steadily more effort.

If you liked this, you might like these other posts:

Are ideas getting harder to find because of the burden of knowledge?

Maybe there is no technological slowdown?

The slowdown we wanted

You might also want to subscribe!

Subscribe at mattsclancy.substack.com

View Details

Publication bias is real. In the social sciences, more than one study finds that statistically significant results seem to be about three times more likely to be published than insignificant ones. Some estimates from medicine aren’t so bad, but a persistent bias in favor of positive results remains. What about science more generally?

To answer that question, you need a way to measure bias across fields that might be very different in terms of their methodology. One way is to look for a correlation between the size of the standard error and the estimated size of an effect. In the absence of publication bias, there shouldn’t be a relationship between these two things. To see why, suppose your study has a lot of data points. In that case, you should be able to get a very precise estimate that is close to the actual population average of the thing you’re studying. On the other hand, if your study has very few data points, you’ll get a very imprecise estimate, including a high probability of getting something very much bigger than the actual population mean and a high probability of getting something very much smaller. But over lots of studies, if there’s no publication bias, you’ll get some abnormally high estimates, some abnormally low ones, but in the end they’ll cancel each other out. If, however, small estimates are systematically excluded from publication, then you’ll end up with a robust correlation between the size of your standard errors and the size of your effects. The extent of this correlation is a way to measure the extent of publication bias in a given literature.

(A downside of this approach is that it will only work in disciplines where this framework makes sense; where research is primarily about measuring effect sizes with noisy data. But enough disciplines do this that it’s a start.)

Fanelli, Costas, and Ioannidis (2017) obtain 1,910 meta-analyses drawn from all areas of science, and pull from these 33,355 datapoints from original underlying studies. For each meta-analysis, they compute the correlation between the standard error and the size of the estimated effect; they then do a weighted average across the different meta-analyses to generate a sort of average over the meta-analyses in fields they cover. In general, the more positive the estimate, the stronger the correlation between standard errors and effect size, implying stronger publication bias. Results below:

Note that the social sciences (up at the top) have pretty high measures of bias, estimated with a lot of precision, while many (but not all) of the biological fields also have fairly high bias. But also note the bottom two rows, which seem to exhibit no bias: computer science, chemistry, engineering, geosciences, and mathematics.

As noted already though, this method of measuring bias might not be appropriate for all fields, since it is rigidly defined in terms of sampling from noisy data. Fanelli (2010) uses a simpler, but more flexible measure of publication bias. Fanelli analyses a random sample of 2,434 papers from all disciplines that include some variation of the phrase “test the hypothesis.” For each paper, Fanelli determined if the authors of the paper argued they had found positive evidence for their hypothesis or not (that is, they either found no evidence in favor of the hypothesis, or actually found contrary results). As a rough and ready test of publication bias, he then looked at the share of hypotheses in each field for which positive support was found. He finds between 75 and 90% of hypotheses mentioned in published papters tend to be supported, across different disciplines. But there are some significant differences across disciplines.

Fanelli cuts papers into six categories: physical sciences, biological sciences, and social sciences, and for each one he further sub-divides papers into pure science and applied science. There are no major differences among applied papers in all three domains - bias seems to be quite high in every case. But in the pure science fields, physical sciences tended to find support for less than 80% of their hypotheses while social sciences tended to find support for nearly 90% of the hypotheses investigated. Biology is in the middle.

Taken together, these two studies suggest the social sciences have bigger problems with publication bias than do the biological sciences, which tend to have more problems than the hard sciences. Why?

What Drives Bias Across Fields?

Let me run through three possible explanations, before looking at one of the few studies that can provide evidence. Note that there are probably more possible explanations, and they’re not mutually exclusive. As far as I can tell, there is very little work on this question (please email me if you know of other relevant work - I would love to hear about it).

First, variation in publication bias could be related to the nature of publication in different fields. If it’s easier to draft and push an article through peer review in some fields than in others then some fields may end up getting more results out there (even if they’re not out there in a top-ranked journal). In the social sciences, we have some evidence that the biggest difference between null results and strong results is that most null results are never even written up and submitted for publication. Maybe that’s because it’s too much work for too little reward. In a field where writing up and publishing results from an experiment somewhere is easy, it might be worth doing, if only to add another line to the CV.

Second, variation in publication bias could be related to the nature of data in different fields. It may be easier in some fields to tightly control for noise in data, or to obtain many more observations, than in others. In economics, a big sample might be hundreds of thousands of observations. In physics, the Large Hadron Collider generates 30 petabytes of data per year. In fields where clean data is plentiful, it might not be the case that when you run an experiment sometimes you find support for a hypothesis and sometimes you don’t. You always find the same thing, or at least always come to the same conclusion about statistical significance. In that case, you won’t find much of a relationship between the size of standard errors and effect sizes: within the range of observed standard errors, everything is either significant or not.

Lastly, it may be that fields differ in their criteria for deciding what is worth publishing. The root cause of publication bias is that journals want to highlight notable research, in order to be relevant to their readership. But what counts as notable research?

Suppose that empirical research is most notable when it provides support for specific theories. In that case, a question in which multiple competing theories make different predictions might exhibit less publication bias. If there is a theory that predicts a null result, and another theory that predicts a statistically significant result, and we don’t have good evidence on which theory is correct, then either result is notable and helps us understand how the world works. Consequently, a journal should be more willing to publish either result.

It might be the case, for example, that the hard sciences have sufficiently established theories such that null results are quite surprising when they are found, and hence easier to publish. In the social sciences, in contrast, we’re just not there yet. Instead, we have an unstated assumption that most hypotheses are false. When we fail to find evidence for one of these hypotheses, it’s not surprising or notable, and so harder to publish.

A bit of evidence from economics

To shed a little light on these questions, let’s look at one more study of differential bias. We’ve seen some evidence that bias varies across major disciplines. But we also have some evidence that bias varies within a particular discipline.

Doucouliagos and Stanley (2013) looks at 87 different meta-analyses from empirical economics and measure the extent of publication bias in each of the literatures covered using the approach already covered, where standard errors are compared with effect sizes. In the figure below, they classify anything smaller than 1 as exhibiting little to modest selection bias, anything between 1 and 2 as exhibiting substantial selection bias, and anything over 2 as exhibiting severe selection bias. They find there are plenty of results in each category.

What drives different levels of bias in economics?

I think it is less likely that variation in publication bias within economics is driven by different publication standards within the different meta-analyses covered. In many cases, these literature are publishing in the exact same journals, but on different questions.

Doucouliagos and Stanley provide a bit more evidence that publication bias might be related to data though. My subjective read on the quality of data across economic fields is that macroeconomics has the toughest time with getting lots of clean data. And Doucouliagos and Stanley do find publication bias seems to more extreme in macroeconomics than other fields.

But Doucouliagos and Stanley (2013) is really set up to test the third explanation: that differences in the range of values permitted by theory explain a big chunk of the variation in publication bias across fields. How are you going to measure that though?

Doucouliagos and Stanley take a few different approaches. First, they just use their own judgement to code up each meta-analysis as pertaining to a question where theory predicts empirical results can go either way (i.e., positive or negative). Second, they use their own reading of the meta-analyses or draw on surveys (where they exist) to assess whether there is “considerable” debate around this area of research. Whereas they claim their first measure is non-controversial and that most economists would agree with how they code things, they acknowledge the second criteria is a subjective one.

By both of these measures, they find that when theory admits a wider array of results, there is less evidence of publication bias. And the effects are pretty large. A field whose theory they code as admitting positive and negative results has a lot less bias than one that doesn’t - the difference is large enough to drop from “severe” selection bias to “little or no” selection bias, for example.

But maybe we’re worried at this point that we have the direction of causality exactly backwards. Maybe it’s not that wider theory permits a wider array of results to be published. Maybe it’s that a wider array of published results leads theorists to come up with wider theories to accommodate this evidence. Doucouliagos and Stanley have two responses here. First, there is a difference between the breadth of results published and publication bias and they try to control for the former to really isolate the latter. After all, it is possible for a field to have both selection bias and a wide breadth of results published. Their methodology can separately identify both, at least in theory, and so they can check if there is more selection bias when there is more accommodating theory, even when two fields have an otherwise similarly large array of results to explain.

But in practice, I wonder if controlling for this is hard to do. So I am a fan of the second approach they take to address this issue. There are some theories in economics where there just really isn’t much wiggle room about which way the results are supposed to go. One of them is studies estimating demand. Except for some exotic cases, economists expect that if you hold all else constant, when prices go up, demand should go down, and vice-versa. We even permit ourselves to call this the “law” of demand. Economists almost uniformly will believe that apparent violations of this can be explained by a failure to control for confounding factors. They will strongly resist the temptation to derive new theories that predict demand and price go up or down together.

Moreover, it isn’t controversial to identify which meta-analyses are about estimating demand and which are not. So for their final measure, Doucouliagos and Stanley look at estimates of bias in studies that estimate demand and those that don’t. And they find studies that estimate demand exhibit much more selection bias than those that don’t (even more than in their measures about extent of debate or what theory permits). In other words, when economists get results that say there is no relationship between price and demand, or that demand goes up when prices go up, these results appear less likely to be published.

So, at least in this context, if your theory admits a wider array of “notable” findings, then you seem to have less trouble getting findings published. Of course, this is just one study, so I want to be cautious leaning too heavily on it. Indeed - who knows if others have looked for the same relationships elsewhere and gotten different results, but have been unable to publish them? (Joking! Mostly.)

If you liked this post, you might also like:

Publication bias is a real thing

One study, many results

What ails the social sciences?

Subscribe at mattsclancy.substack.com

View Details

Publication bias is when academic journals make publication of a paper contingent on the results obtained. Typically, though not always, it’s assumed that means results that indicate a statistically significant correlation are publishable, and those that don’t are not (where statistically significant typically means a result that would be expected to occur by random chance less than 5% of the time). A rationale for this kind of preference is that the audience of journals is looking for new information that leads them to revise their beliefs. If the default position is that most novel hypotheses are false, then a study that fails to find evidence in favor of a novel hypothesis isn’t going to lead anyone to revise their beliefs. Readers could have skipped it without it impacting their lives, and journals would do better by highlighting other findings. (Frankel and Kasy 2021 develops this idea more fully)

But that kind of policy can generate systematic biases. In another newsletter, I looked at a number of studies that show when you give different teams of researchers the same narrow research question and the same dataset, it is quite common to get very different results. For example, in Breznau et al. (2021), 73 different teams of researchers are asked to answer the question “does immigration lower public support for social policies?” Despite each team being given the same initial dataset, the results generated by these teams spanned the whole range of possible conclusions, as indicated in the figure below: in yellow, more immigration led to less support for social policies, in blue the opposite, and in gray there was no statistically significant relationship between the two.

Seeing all the results in one figure like this tells us, more or less, that there is no consistent relationship between immigration and public support for social policies (at least in this dataset). But what if this had not been a paper reporting on the coordinated efforts of 73 research teams, committed to publishing whatever they found? What if, instead, this paper was a literature review reporting on the published research of 73 different teams that tackled this question independently? And what if these teams could only get their results published if they found the “surprising” result that more immigration leads to more public support for social policies? In that case, a review of the literature would generate a figure like the following:

In this case, the academic enterprise would be misleading you. More generally, if we want academia to develop knowledge of how the world works that can then be spun into technological and policy applications, we need it to give us an accurate picture of the world.

On the other hand, maybe it’s not that bad. Maybe researchers who find different results can get their work published, though possibly in lower-ranked journals. If that’s the case, then a thorough review of the literature could still recover the full distribution of results that we originally highlighted. That’s not so hard in the era of google scholar and sci-hub.

So what’s the evidence? How much of an issue is publication bias?

Ideal Studies

Ideally, what we would like to do is follow a large number of roughly similar research projects from inception to completion, and observe whether the probability of publication depends on the results of the research. In fact, there are several research settings where something quite close to this is possible. This is because many kinds of research require researchers to submit proposals or pre-register studies before they can begin investigation. A large literature attempts to track the subsequent history of such projects to see what happens to them.

Dwan et al. (2008) reports on the results of 8 different healthcare meta-analyses of this type, published between 1991 and 2008. The studies variously draw on the pool of proposals submitted to different institutional review boards, or drug trial registries, focusing on randomized control trials where a control group is given some kind of placebo and a treatment group some kind of treatment. The studies then perform literature searches and correspond with the authors of these studies to see if the proposals or pre-registered trials end up in the publication record.

In every case, results are more likely to be published if they are positive than if the results are negative or inconclusive. In some cases, the biases are small. One study (Dickersin 1993) of trials approved for funding by the NIH found 98% of positive trial outcomes were published, but only 85% of negative results were (i.e., positive results were 15% more likely to be published - though, as a commenter below points out, you could just as easily say null results are not published 15% of the time, and that is 7.5 times as much as positive results are not published). In other cases, they’re quite large. In Decullier (2005)’s study of protocols submitted to the French Research Ethics Committees, 69% of positive results were published, but only 28% of negative or null results were (i.e., positive results were about 2.5 times more likely to be published).

These kinds of studies are in principle possible for any kind of research project that leaves a paper trail, but they are most common in medicine (which has these institutional review boards and drug trial registries). However, Franco, Malhotra, and Simonovits (2014) are able to perform a similar exercise for 228 studies in the social sciences that relied on the NSF’s Time-sharing Experiments in the Social Sciences program. Under this program, researchers can apply to have questions added to a nationally representative survey of US adults. Typically, the questions are manipulated so that different populations receive different versions of the question (i.e., different question framing or visual stimulus), allowing the researchers to conduct survey-based experiments. This paper approaches the ideal in several ways:

The proposals are vetted for quality via a peer review process, ensuring all proposed research question clear some minimum threshold

The surveys are administered in the same way by the same firm in most cases, ensuring some degree of similarity of methodology across studies

All the studies are vetted to ensure they have sufficient statistical power to detect effects of interest

That is, in this setting, what really differentiates research projects is the results they find, rather than the methods or sample size.

Franco, Malhotra, and Simonovits start with the set of all approved proposals from 2002-2012 and then search through the literature to find all published papers or working papers based on the results of this NSF program survey. For more than 100 results, they could find no such paper, and so they emailed the authors to ask what happened and also to summarize the results of the experiment.

The headline result is that 92 of 228 studies got strong results (i.e., statistically significant, consistent evidence in favor of the hypothesis). Of these, 62% were published. In contrast, 49 studies got null results (i.e., no statistically significant difference between the treatment and control groups). Of these, just 22% were published. Another 85 studies got mixed results, and half of these were published. In other words, studies that got strong results were 2.8 times as likely to be published as those that got null results.

The figure below plots the results from both studies, with the share of studies with positive results that were published on the vertical axis and the share of studies with negative results that were published on the horizontal axis. The blue dots correspond to 8 different healthcare meta-analyses from Dwan et al. (2008), and the orange ones to four different categories of social science research in Franco, Malhotra, and Simonovits (2014). For there to be no publication bias, we would expect dots to lie close to the 45-degree line, on either side. Instead, we see they all lie above it, indicating positive results are more likely to be published than negative ones. The worst offenders are in the social sciences. Of the 60 psychology experiments Franco, Malhotra, and Simonovits study, 76% of positive results resulted in a publication, compared to 14% of null results. In the 36 sociology experiments, 54% of positive results got a publication and none of the null results did.

Franco, Malhotra, and Simonovits (2014) has an additional interesting finding. As we have seen, most null results do not result in a publication for this type of social science. But in fact, of the 49 studies with null results, 31 were never submitted to any journal, and indeed were never even written up. In contrast, just 4 of 98 positive results were not written up. In fact, if you restrict attention only to results that get written up, 11 of 18 null results that are written up get published - 61% - and 57 of 88 strong results that get written up get published - 64%. Almost the same, once you make the decision to write them up.

So why don’t null results get written up and submitted for publication? Franco, Malhotra, and Simonovits got detailed explanations from 26 researchers. Most of them (15) said they felt the project had no prospect for publication given the null results. But in other cases, the results just took a backseat to other priorities, such as other projects. And in two cases, the authors eventually did manage to publish on the topic… by getting positive results from a smaller convenience sample(!).

Why is this interesting? It shows us that scientists actively change their behavior in response to their beliefs about publication bias. In some cases, they just abandon projects that they feel face a hopeless path to publication. In other cases, they might change their research practices to try and get a result they think will have a better shot at publication, for example by re-doing the study with a new sample. The latter is related to p-hacking, where researchers bend their methods to generate publishable results. More on that at another time.

Less than Ideal Studies

These kinds of ideal studies are great when they are an option, but in most cases we won’t be able to observe the complete set of published and unpublished (and even unwritten) research projects. However, a next-best option is when we have good information about what the distribution of results should look like, in the absence of publication bias. We can then compare the actual distribution of published results to what we know it should look like to get a good estimate of how publication bias is distorting things.

Among other things, Andrews and Kasy (2019) show how to do this using data on replications. To explain the method, let’s think of a simple model of research. When we do empirical research, in a lot of cases what’s actually going on is we get our data, clean it up, make various decisions about analysis, and end up with an estimate of some effect size, as well as an estimate of the precision of our estimate - the standard error. Andrews and Kasy ask us to ignore all that detail and just think of research as like drawing a random “result” (an estimated effect size and standard error) from an unknown distribution of possible results. The idea they’re trying to get at is just that if we did the study again and again, we might get different results each time, but if we did it enough we would see they cluster around some kind of typical result, which we could estimate with the average. In the figure below, results are clustered around an average effect size of -0.5 and an average standard error 0.4.

If we weren’t worried about publication bias, then we could just look at all these results and get a pretty good estimate of the “truth.” But because of publication bias we’re worried that we’re not seeing the whole picture. Maybe there is no publication bias at all and the diagram above gives us a good estimate. Or maybe results that are not statistically significant are only published 10% of the time. That would look something like this, with the light gray dots results that aren’t published in this example.

In this case, the unpublished results tend to be closer to zero for any given level of standard error. By including them, the average effect size is cut down by 40%.

What Andrews and Kasy do is show is that if you have a sample of results that isn’t polluted by publication bias, you can compare that to the actual set of published results to infer the extent of publication bias. So where do you get a sample of results that isn’t polluted by publication bias? One option is from systematic replications.

Two such projects are Camerer et al. (2016)’s replication of 18 laboratory experiments published in two top economics journals between 2011 and 2014, and Open Science Collaboration (2015)’s replication of 100 experiments published in three top psychology journals in 2008. In both cases, the results of the replication efforts are bundled together and published as one big article. We have confidence in each case that all replication results get published; there is no selection where successful or unsuccessful replications are excluded from publication.

We then compare the distribution of effect sizes and standard errors, as published in those journals, to an “unbiased” distribution that is derived from the replication projects. For example, if half of the replicated results are not statistically significant, and we think that’s an unbiased estimate, then we should expect half of the published results to be insignificant too in the absence of publication bias. If instead we find that just a quarter of the published results are not statistically significant, that tells us significant results are three times more likely to be published, since there are three times as many significant ones as insignificant, and we would expect them to have equal probability of being significant or insignificant. Using a more sophisticated version of this general idea, Andrews and Kasy estimate that null results (statistically insignificant at the standard 5% level) are just 2-3% as likely to be published in these journals as statistically significant results!

Now, note the important words “in these journals” there. We’re not saying null results are 2-3% as likely to be published as significant results; only that they are 2-3% as likely to be published in these top disciplinary journals. As we’ve seen, the studies from the previous section indicate publication bias exists but is usually not on the order of magnitude seen here. We would hope non-significant results rejected for publication in these top journals would still find a home somewhere else (but we don’t know that for sure).

Even Less Ideal Studies

This technique works when you have systematic replication data to generate an unbiased estimate of the “true” underlying probability of getting different effect sizes and standard errors. But replications remain rare. Fortunately, there is a large literature on ways to estimate the presence of publication bias even without any data on the distribution of unpublished results. I won’t talk about this whole literature (Christensen and Miguel 2018 is a good recent overview) but I’ll talk about Andrews and Kasy’s approach here, but plan to cover some of the others in a future newsletter.

Let’s go back to this example plot from the last section. Suppose we have this set of published results. How can we know if this is just what the data looks like, or if this is skewed by publication data?

As I noted before, statistically significant results are basically ones where the estimated effect size is “large” relative to imprecision of that estimate, here given by the standard error. For any point on this diagram, we can precisely calculate which results will be statistically significant and which won’t be. It looks sort of like this, with results in the red area not statistically significant.

Notice the red area of statistical insignificance is eerily unpopulated; an indication that we’re missing observations (these are called funnel plots, in the literature).

What Andrews and Kasy propose is something like this: for a given standard error, how many do you see in the zone of insignificant results versus outside it. For instance, look at the regions highlighted in blue in the figure below.

If we assume that without publication bias there is no systematic relationship between standard errors effect sizes (i.e., without publication bias, a big standard error is equally likely for a big effect as a small one), then we should expect to see a similar spread of effect sizes across each rectangle. In the bottom one, they range pretty evenly from -1 to 3, but in the top one they are all clustered on around 2. Also, in the top area, most of the area lies in the “red zone” of statistical insignificance and so we should expect to observe a lot of statistically insignificant results there - the majority of observations, in fact. But instead, we just observe 1 out of 6 total observations statistically insignificant. This kind of information about the difference between what we should expect to see versus what we actually see can also be used to infer the strength of publication bias. In this case, instead of using replications to get at the “true” distribution, it’s like we’re using the distribution of effect sizes for precise results to help tell us about what kinds of effect sizes we should see for less precise results. After all, they should be basically the same, since usually an imprecise result is just a precise result with fewer datapoints. There isn’t anything fundamentally different about them.

Andrews and Kasy apply this methodology to Wolfson and Belman (2015), a meta-analysis of studies on the impact of the minimum wage on employment. They estimate results that find a negative impact of the minimum wage on employment (at conventional levels of statistical significance) are a bit more than 3 times as likely to be published as papers that don’t. Unlike the replication based studies, this estimate is a closer match to the results we found for the ideal data studies, where social science papers that found strong positive results were around 2.8 times as likely to be published as those that didn’t.

So, to sum up; yes - publication bias is a real thing. You can see it’s statistical fingerprints in published data. And when you’re lucky enough to have much better data, such as on the distribution of results that would occur without publication bias, or when you can actually see what happens to unpublished research, you find the same thing: positive results have an easier time getting published.

So what do you do about it? Well, you can find ways to reduce the extent of bias, for example by getting journals to precommit to publishing papers based on the methods and the significance of the question asked, rather than the results. Or you can create avenues for the publication of non-significant findings, either in journals or as draft working papers. Recall however that Franco, Malhotra, and Simonovits found most null results weren’t even written up at all - it may be that if non-significant work can be published in some outlets but will still be ignored by other researchers, researchers might not bother to write them up and instead choose to allocate their time to pursuing research that might attract attention.

Alternatively, if you can identify publication bias, you can correct for it with statistical tools. Andrews and Kasy, as well as others, have developed ways to infer the “true” estimate of an effect by estimating the likely value of unpublished research. Indeed, Andrews and Kasy make such a tool freely available on the web.

But we can also look more deeply at the underlying causes of publication bias. It turns out that the extent of publication bias varies widely across disciplines and sub-disciplines. That’s a bit surprising. Why is that the case? The plan for next week is to dig into precisely that question.

If you liked this post, you might also like:

One study, many results

How a field fixes itself: the applied turn in economics

How bad is publish-or-perish for the quality of science?

Subscribe at mattsclancy.substack.com

View Details

In last week’s newsletter, we looked at a thought experiment by Jones and Summers that pretty convincingly argued the average return on a dollar of R&D was really high. That would seem to suggest we should be spending a lot more on R&D.

But the devil is in the details. How, exactly, should you increase your R&D spending? Today, let’s look at one kind of program that seems to work and would be an excellent candidate for more funds: the US’ Small Business Innovation Research (SBIR) program and the European Union’s SME instrument (which was modeled on the SBIR).

Grants for Small Business Innovation

The SBIR and SME instrument programs are competitive grant competitions, where small businesses submit proposals for R&D grants from the public sector. Each program consists of two phases, where the first phase involves significantly less money than the second. For the SBIR program, a phase 1 award is typically up to $150,000 and roughly $1mn in phase 2; for the SME instrument, phase 1 is just €50,000 and phase 2 is €0.5-2.5mn. In the US, firms apply for phase 1 first, and then phase 2, whereas in the EU firms can apply to either straightaway. Broadly speaking, the money is intended to be used for R&D type projects.

They’re pretty competitive. In the US, an application to the SBIR program run by the Department of Energy will typically take a full-time employee 1-2 months to complete and has about a 15% chance of winning a phase 1 award; conditional on winning a phase 1 award, firms have about a 50% chance of winning a phase 2 award (or overall chances about 8%). In the EU, the probability of winning in phase 1 is about 8%, and in phase 2, just 5%.

So, both programs involve the government attempting to pick winning ideas, and then giving the winners money to fund R&D. How well do they work? Does the money generate innovations? Does it get the good return on investment that Jones and Summers’ thought experiment implies should be possible?

Evaluating the Impact of Grants

Two recent papers look at this question using the same method. Howell (2017) looks at the Department of Energy’s SBIR program ($884mn disbursed over 30 years), while Santoleri et al. (2020) looks at the SME instrument program (€1.3bn disbursed over 2014-2017). Each paper has access to details on all applicants to the program, not merely the winners. That means they can follow the trajectory of businesses that apply and get a grant as well to those that apply but fail to get an R&D grant.

But to assess the impact of money on innovation, they can’t just compare the winners to the losers, because the government isn’t randomly allocating the money: it’s actively trying to sniff out the best ideas. That means the winning applicants would probably have done better than the losers, even if they hadn’t received any R&D funds, since someone thought they had a more promising idea. But the way these programs are administered have a few quirks that allow researchers to estimate the causal impact of getting money.

During the period studied, each program held lots of smaller competitions devoted to a specific sector or technology. For example, the Department of Energy might solicit proposals for projects related to Solar Powered Water Desalination. Within each of these competitions, proposals are scored by technical experts and then ranked from best to worst. After this ranking is made overall budgets are drawn up for the competitions (without reference to the applications received), and the best projects get funded until the money runs out. For example, the Department of Energy might receive 11 proposals for it’s solar powered water desalination topic, but (without looking at the quality of the proposals) decide it will only be able to fund 3. The top three each get their $150,000, and the fourth gets nothing.

The important thing is that, for applications right around the cut-off, although there is a big change in the amount of money received, there shouldn’t be a big change in the quality of the proposals (in our example, the difference between third and fourth place shouldn’t be abnormally large). That is, although we don’t have perfect randomisation, we have something pretty close. Proposals on either side of the cutoff for funding differ only a bit in the quality of their proposals, but experience a huge difference in their ability to execute on those proposals because some of them get money and some don’t. It’s a pretty close estimate of the causal impact of getting the money.

What’s the Impact

Each paper looks at a couple measures of impact. A natural place to start when evaluating the impact of small business innovation is patents. In each case, patents are weighted by how many citations they end up receiving (while citations might be a problematic measure of knowledge flows, they seem to be quite good as a way of measuring the value of patents: better patents seem to get more citations). In the text, I’ll just call them patents, but you should think of them as “patents adjusted for quality.” The papers produce some nice figures (Howell left, Santoleri et al. right).

These figures nicely illustrate the way the impact of cash is assessed in these papers. Especially in the left figure, you can see that, as we worried, proposals that are ranked more highly by the SBIR program do tend to get more patents: the SBIR program does have the ability to judge which projects are more likely to succeed. Even looking at projects that don’t get funding (the ones to the left of the red line), Howell’s measure of patenting rises by about 0.2 log points at each rank, as we move from -3 to -2 to -1. If that trend were to continue, we would expect the +1 ranked proposal to jump another 0.2, to something a bit less than 0.8. But instead, it jumps more than twice as much. That’s pretty suggestive that it was the funding that mattered. Estimated more precisely, Howell finds getting one of the SBIR’s phase 1 grants increases patenting by about 30%. Santoleri et al. (2020) get similar results, estimating a 30-40% increase in patenting from getting a phase 2 grant (though note that the phase 2 grants in the EU tend to be a lot larger than the phase 1 US grants).

To the extent we’re happy with patents as a measure of innovation, we’ve already shown that the program successfully manages to buy innovations. But the papers actually document a voluminous set of additional indicators all associated with a healthy and flourishing innovative business (in all cases below, SBIR refers to phase 1 and SME refers to phase 2):

Winning an SBIR grant doubles the probability of getting venture capital funding; winning an SME instrument grant triples the probability of getting private equity funding

Winning an SBIR grant increases annual revenue by $1.3-1.7mn, compared to an average of $2mn

Winning an SME instrument grant increases the growth rate of company assets by 50-100%, the growth rate of employment by 20-30%, and significantly decreases the chances of the firm failing.

Benchmarking Value for Money

OK, so winning money helps firms. Is that surprising? Do we really need scientists to tell us that? In fact, it’s not guaranteed. When Wang, Li, and Furman (2017) apply this methodology to a similar program in China they don’t find the money makes a statistically significant difference. That could be for a lot of reasons (discussed in the paper), but the main point is simply that we can’t take for granted seemingly obvious results like “giving firms money helps them.”

But still, even if we find R&D grants help firms, that doesn’t necessarily imply it’s a good use of funds. We want to know the return on this R&D investment. That’s challenging because although we know the cost of these programs, it’s hard to put a solid monetary value on the benefits that arise from them, which is what we would need to do to calculate a benefits cost ratio.

So let’s take a different tack. One thing we can measure reasonably well is whether firms get a patent. So let’s just see how many patents these programs generate per R&D dollar and compare that to the number of patents per R&D dollar that the private sector generates. If we assume the private sector knows what it’s doing in terms of getting a decent return on R&D investment, then that gives us a benchmark against which we can assess the performance of these government programs.

So how many patents per dollar does the private sector get? If you divide US patent grants (from domestic companies) over 2010-2017 by R&D funded by US businesses in the same year, you pretty consistently get a ratio of around 0.5 patents per million dollars of R&D (details here). That’s about the same ratio as this post finds, looking only at 7 top tech companies.

To be clear, the point isn’t that each patent costs 2 million dollars of R&D. R&D doesn’t just go into patents. This report found in 2008 that only about 20% of companies that did R&D reported a patent. Taking that as a benchmark, suppose that only 20% of inventions get patented; in that case, we could think of this as telling us that every $2mn in R&D generates 5 “innovations” of which one gets patented. As long as SBIR/SME grant recipients have a similar ratio between innovation and patenting as other US R&D performing firms, then looking at patents per R&D for them is an OK benchmark for the productivity of R&D spending.

If that sounds good enough to you, read on! If not, I say a bit more about this in an extra note at the end of this post. Feel free to check that out and then come back here if you’re feeling skeptical.

So do these programs generate patents at a similar rate of 0.5 per million? Yes!

Value for money in the SBIR Program

This isn’t something that Howell (2017) or Santoleri et al. (2020) calculate directly, but you can back out estimates from their results using a method described in the appendix of our next paper, Myers and Lanahan (2021). Myers and Lanahan estimate Howell’s results imply the DOE SBIR program gets about 0.8-1.3 patents per million dollars. Applying their method to the range of estimates in Santoleri et al. (2020) and converting into dollars, you get something in the ballpark of 0.7 patents per million dollars in the SME instrument program (see the extra notes section at the bottom for more on where that number comes from). In either case, that compares pretty favorably with a rough estimate of 0.5 patents per million R&D dollars for the US private sector.

That’s reassuring, but it’s not exactly what we’re interested in. As stated at the outset, Jones and Summers’ thought experiment implies that R&D is a really good investment once you take into account all the social benefits. What we have here is evidence that the SBIR and SME instrument programs can probably match the private sector in terms of figuring out how to wisely spend R&D dollars to purchase innovations. Frankly, that seems plausible to me. It just means governments, working with outside technical experts (that even the private sector might need to turn to) could do about as well as the private sector. But they don’t tell us much about the benefits that accrue from these R&D investments that aren’t captured by the patents the grant recipients get.

But that’s what Myers and Lanahan (2021) is about. What they would like to see is how giving R&D money to different technology sectors leads to more patents in that sector by grant recipients, as well as other impacts on patenting in general. For example, if we give a million dollars to a couple firms working on solar powered water desalination, how many new solar water patents do we get from those grant recipients? What about solar water patents from other people? What about patents that aren’t about solar powered water desalination at all?

Like Howell, they’re going to look at the Department of Energy’s SBIR program. They need to use a different quirk of the SBIR though, because they’re not comparing firms that get funds to firms that don’t; they’re comparing entire technology fields that get more money to fields that get less money.

Instead, they rely on the fact that some US states have programs to match SBIR funding with local funds. Importantly, DOE doesn’t take that into account when deciding how to dole out funds. For example, in 2006, North Carolina began partially matching the funds received by SBIR winners in the state. If a bunch of winning applicants in solar technology happen to reside in, say, North Carolina in 2008 instead of South Carolina in 2008 or North Carolina in 2005, then those recipients get their funds partially matched by the state and solar technology research, as a field, gets an unexpected windfall of R&D dollars. What Myers and Lanahan end up with is something close to random R&D money drops for different kinds of technologies.

Myers and Lanahan use variation in this unexpected “windfall” money to generate estimates of the return on R&D dollars. For this to work, you have to believe there are no systematic differences between SBIR recipients that reside in states with matching programs and those that don’t, and they present some evidence that this is the case.

One more hurdle though. You can crudely measure innovation by counting patents. And, with some difficulty, you can come up with estimates of more-or-less random R&D allocations to different technologies. But if you want to see how the one affects the other, you have to link patents to SBIR technology areas. Myers and Lanahan accomplish this with natural language processing. For every SBIR grant competition, they analyze the text of the competition description and identify the patent technology categories whose patents are textually most similar to this description. When, say, solar technology gets a big windfall of R&D money, they can see what happens to the number of patents in the patent technology categories that historically have been textually closest to the DOE’s description of what it was looking to fund. And this is also how they measure the broader impact of SBIR money on other technologies. When solar gets a big windfall, what happens to the number of patents in technology categories that are not solar, but are kind of “close” to solar technology (as measured by text similarity)?

OK! So that’s what they do. What do they find? More money for a technology means more patents!

In the above figure, each dot is a patent technology group and compares the funding received by that technology to subsequent patenting. The figure looks at the patents of DOE SBIR recipients only and nicely illustrates the importance of doing the extra work of trying to estimate “windfall” funding. The steep dotted green line is what you get if you just tally up all the funding the SBIR program gives to different technologies - it looks like a little more funding gives you a lot more patents. But this is biased by the fact that the technologies that get the most money were already promising (that’s why they got money!). The flatter dark blue line is the relationship between quasi-random windfall money and patenting. It’s still the case that more money gets you more patents, but the relationship isn’t as strong as the green one. But this is the more informative estimate on the actual impact of cash.

Using estimates based on windfall funding, an extra million dollars is associated with SBIR recipients being granted about 0.5 additional patents. Which is pretty typical (or so I’ve argued here). But the more important finding is that’s only a fraction of the overall benefit. When there’s more R&D in a given technology sector, we typically think that creates new opportunities for R&D from other firms, because they can learn from the discoveries made by the grant recipient. Indeed, other papers have found spillovers are often just as important, or even more important, than the direct benefits to the R&D performer.

Myers and Lanahan get at this in two ways. First, they look for an impact of R&D funding not only on the patent technology classes that are closest to SBIR’s description of the funding competition, but also ones that are more textually distant. Typically, a bigger share of extra patent activity comes from classes that are not the closest fields, but still closer than a random patent (consistent with other work). Second, they look at patents held by people who are not SBIR recipients themselves, but who live closer or farther away from SBIR recipients.

So, looking only at SBIR recipients, an extra million tends to produce an extra 0.5 patents. Looking at patents belonging to anyone in the same county as an SBIR recipient - a group for whom we might assume is likely to contain people with similar technical expertise and possibly overlapping social and professional networks - an extra million tends to produce an extra 1.4 patents (across a wide range of technology fields). And looking at all US patents (from inventors residing anywhere in the world), an extra million tends to produce an extra 3 patents.

If all those patents are equally valuable, that would imply when the SBIR gives out money, the innovation outputs created by recipients are only a small part of the overall effect (0.5 of 3 total patents). Of course, all those patents are not, in fact, equally valuable. The ones created by the grant recipients tend to be more highly cited than the ones that we’re attributing to knowledge spillovers. Still, Myers and Lanahan estimate that after adjusting for the quality of patents, half the value generated by an SBIR grant is reflected in the patents of non-recipients working on different (but not too different) technology.

Prospects for Scaling Up

Whew! That’s a lot. To sum up: we’ve got some good theoretical reasons to think the return on R&D is very high, on average. If we look at a specific R&D program that gives R&D grants to small firms, the grants are effective at funding innovation at about the same level as the private sector could manage. And if we try to assess the broader impact of that funding, we find including all the social benefits gives us a return at least twice as high as the ones we got by focusing just on the grant recipients; and those were already decent! All together, more evidence that we ought to be spending more on R&D.

Lastly, we have good reason to think these effects can also be maintained if we scale up these programs. The design of Howell (2017) and Santoleri et al. (2020) is premised on estimating the impact of R&D funding on firms right around the cut-off. For the purposes of scaling up, that’s great news, because if we increased funding the firms that would get extra money would be ones that are closest to the cut-off.

If you liked this post, you might also like:

More science leads to more innovation (the link between the production of science and technological innovation)

Free knowledge and innovation (another innovation policy that works: disseminating knowledge freely)

Importing knowledge (and another innovation policy that works: immigration)

Extra credit

One reason patents per R&D dollar might be a bad benchmark in this case is if we think small firms like the ones getting R&D grants are more likely to seek patents than the typical R&D performing firm. There are some good reasons to think that’s the case: basically, they’re small but they aim to grow on the back of their technologies and so they need all the protection they can get. But looking at Howell (2017) and Santoleri et al. (2020) finds the median firm in these programs still has zero patents (even after winning an award). If just 20% of firms that do R&D also have a patent, then these grant recipients can’t be much more than twice as likely to get patents as everyone else. Alternatively, if you compare patents per dollar of small firms to large ones for the USA, they don’t look that different in aggregate. Nonetheless, I fully concede R&D is certainly a noisey predictor of patents; but I’m out of other ideas.

Santoleri et al. (2020) find the mean patents per firm is 4 among phase II applicants, and that getting a phase II grant increases cite-weighted patenting by 15-40%. That implies between 0.15x4 = 0.6 and 0.4 x 4 = 1.6 patents for every €0.5-2.5mn, or 0.2-3.2 patents per million euros or 0.2-2.9 patents per million dollars. Using the midpoint of each you get 0.7 patents per million dollars

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

Jones and Summers (2021) is a new working paper that attempts to calculate the social return on R&D - that is, how much value does a dollar of R&D create? The paper is like something out of another time; the argument is so simple and straight-forward that it could have been made at any point in the last 60 years. It requires no new math or theoretical insights; just basic accounting and some simple data. The main insight is simply in how to frame the problem.

What I want to do in this post is walk through Jones and Summers’ simple “thought experiment.” At the end, we’ll have a new argument that the returns on R&D are quite high and that we should probably be spending much more on R&D. Next week we’ll look at some empirical data to see if it matches the intuition of the thought experiment. (Spoiler: It does)

Taking an R&D Break

Let’s start with a model of long-run changes in material living standards that is so simple it’s hard to argue with:

R&D is an activity that consumes some of the economy’s resources

R&D is the only way new technologies come into existence

Growth in GDP per capita comes entirely from new technologies (at least, in the long run)

We’ll re-examine all of these points later, but for now let’s accept them and move on.

This model helps clarify what it means to compute the returns to R&D. If we do more R&D, we have to use more of the economy’s resources, but in return we’ll get more GDP per capita. So computing the returns to R&D is really about computing how much does growth change when we spend a bit more on R&D. Specifically, if we increase R&D by, say, 1%, what will the expected impact be on GDP per capita?

That’s actually a really hard question to answer! And the clever thing Jones and Summers do is they don’t ask it. Instead, they ask a different question which is much easier to answer: what would happen if we took a break from R&D for a year?

Why is this easier to answer? Because in our simple model, if we stop all R&D, we stop all growth! We no longer have to estimate “how much” growth we get for an extra dollar of R&D. We know that if we stop all R&D, we stop all growth. Simple as that!

Let’s get more specific. Suppose under normal circumstances, we spend a constant share of GDP on R&D. Let’s label the share s. In return, the economy grows by a long-run average that we’ll call g. In the USA, between 1953 and 2019, the annual share of GDP spent on R&D was about 2.5%, so s = 0.025. Over the same time period, GDP per capita (adjusted for inflation) grew by about 1.8% per year, so g = 0.018. If we hit “pause” on R&D for one year, then in that year we save 2.5% of GDP (since we don’t have to spend it on R&D), but GDP per capita stays stuck at its current level for one year, instead of growing by 1.8%.

But that’s not a full accounting of the benefits or the costs of doing R&D. In the next year, our R&D break will end and we’ll start spending 2.5% of GDP on R&D again. But because we took that break, we didn’t grow in the previous year, GDP will be smaller than it would otherwise have been. Since we always spend 2.5% of GDP on R&D, we’ll be devoting a bit less money to R&D than we otherwise would (since it will be 2.5% of a smaller GDP). And because we didn’t grow in the previous year, we’ll also be growing from a lower level than we would have been if we hadn’t taken our R&D break. And that will be true in the next period, and the next, and the next: in every year until the end of time, GDP per capita will be 1.8% lower than it would have been if we had not taken that R&D break.

Adding up all these costs and benefits over time requires us to do some calculations using the interest rate r, which is how economists value dollars at different points in time. In the USA, a common interest rate to use might be 5%, so that r = 0.05. Jones and Summers show the math shakes out so that the ratio from here to infinity of benefits from R&D to costs of R&D is:

Benefits-to-Cost Ratio = g/(sr)

In other words, on average the return on a dollar spent on R&D is equal to the long-run average growth rate, divided by the share of GDP spent on R&D and the interest rate. With g = 0.018, s = 0.025, and r = 0.05, this gives us a benefits to cost ratio of 14.4. Every dollar spent on R&D gets transformed into $14.40!

One thing I really like about this result is that you do not need any advanced math to derive it. It’s just a consequence of algebra and the proposed model of how growth and R&D are linked. In the video below, I show how to get this result without using any math more advanced than algebra.

Can that really be all there is to it? Well, no. If we look more critically at the assumptions that went into generating this number, we can get different benefit-cost ratios. But the core result of Jones and Summers is not any exact number. It’s that whatever number you believe is most accurate, it’s much more than 1. R&D is a reliable money-printing machine: you put in a dollar, and you get back out much more than a dollar.

But let’s turn now to some objections to the simple argument I’ve made so far.

Is there really no growth without R&D?

Starting at the beginning, we might question the assumption that R&D resources are really the only way to get improvements in per capita living standards. If that’s wrong, and growth can happen without R&D, then our thought experiment would be over-estimating the returns to R&D, since growth wouldn’t actually go to zero if we (hypothetically) stopped all R&D for that year.

There are two ways we could get growth without doing R&D. First, it may be that we can get new technologies without spending resources on R&D. Second, we could get growth without new technologies.

The latter case is basically excluded by assumption in economics, at least for countries operating at the technological frontier. In 1956, Robert Solow and Trevor Swann argued that countries cannot indefinitely increase their material living standards by investing in more and more capital. That’s because the returns to investment drop as you run out of useful things to build, until you reach where the returns to investment are offset by the cost of upkeep. To keep growth going, you need to discover new useful things to build. You need new technology.

On the other hand, the first objection - that we may be able to get new technologies without spending resources on R&D - has more going for it. For instance, a common understanding of innovation is that it’s about flashes of insight, serendipity, and ideas that come to you in the shower. Good ideas sometimes just come to us without being sought.

The trouble with this notion of innovation is that in almost all cases, the free idea is only part of the story. It might provide a roadmap, but there is still a long journey from the idea to the execution, and that journey typically requires resources to be expended. In terms of our thought experiment, if it still takes R&D to translate an unplanned inspiration into growth, then we are actually measuring the returns to R&D correctly. If we turned off R&D, those insights wouldn’t get realized, and so growth would freeze until we began R&D again.

But maybe that’s not always the case. In The Secret of Our Success, Joseph Henrich gives a (fictional) example of how a package of hunting techniques could evolve over several generations without diverting any economic resources to innovation. In the example, proto-humans use sticks to fish termites out of a nest to eat, but one of them mistakenly believes the stick must be sharpened (their mother taught them the technique with a stick that happened to be sharp). One day, they accidentally plunge their sharp stick into an abandoned termite mound and impale a rodent - he has “invented” a spear. The proto-humans start using the sticks to impale prey. A generation later, another proto-human sees rabbits leaving tracks in the mud and going back into their hole; he realizes he can follow tracks to the hole and use the spear, instead of just hoping he sees an animal. Bit by bit, cumulative cultural evolution can happen, leading to a steadily more technologically sophisticated society.

These kinds of processes still happen today. In learning-by-doing models of innovation, firms get more productive as they gain experience in a production process. The process by which this happens is likely another form of evolution, with workers and managers tinkering with their process and selectively retaining the changes that improve productivity. We could call this kind of tinkering R&D if we wanted, but it’s almost certainly not part of the national statistics.

But here’s the rub. With modern learning-by-doing, we typically think of firms and workers finding efficiencies and productivity hacks in production processes that are novel. And where do new and unfamiliar production processes come from? In the modern world, typically they are the result of purposeful R&D. If that’s the case, then in the long run we are once again accounting correctly for the costs and benefits of R&D. In this case, if we turned off R&D for one year would delay by one year the creation of new production processes that would then experience rapid learning-by-doing gains in subsequent years.

Of course, there still might be learning-by-doing with older technologies. But learning-by-doing models typically assume progress is very, very slow in mature technologies because there are not many beneficial tweaks left to discover. The process has already been optimized.

That’s pretty consistent with what we know about growth in the era before much purposeful R&D. Tinkering and cumulative cultural evolution is probably the right model for innovation before the industrial revolution, and as best as we can tell, growth during that era was painfully slow. Nearly zero, compared to today’s standards.

All that said, if you still believe growth can happen without R&D, then you can still use Jones’ and Summers’ approach to compute the benefits of R&D and adjust the estimate to take all this into account. It’s just now you need to use take only the fraction of growth that comes from R&D as your benefit. I have argued that almost all long-run growth comes from R&D. But if you think it’s just 50%, then that would cut the benefit cost ratio of R&D in half - to a still very high 7.2.

What about other costs?

A second objection to our initial estimate of the returns to R&D makes the opposite point: R&D is not costlessly translated into growth. New ideas must be laboriously spun out into new products and infrastructure that are then disseminated across the economy, before growth benefits are realized. Focusing exclusively on the R&D costs overstates the returns to R&D by understating the full costs of getting growth.

Take the covid-19 vaccine as an example. Pfizer has said the R&D costs of developing the vaccine were nearly $1bn. But once Pfizer had an FDA-approved vaccine, the benefits were not instantly realized by society. Instead, the shots needed to be manufactured and put into arms, and the cost of building that manufacturing capacity ought to be accounted as part of the cost of deriving a benefit from the R&D.

We don’t know exactly how much the US spends on “embodying” newly discovered ideas in physical form so that they can affect growth. But we do know that since 1960, the total US private sector investment in new capital (not merely upkeep or replacement of existing capital) has been about 4.0% of GDP per year. Not all of that is the upgrading of capital to incorporate new ideas. Some of it is just extending existing forms of capital over a growing population (think building new houses). But it’s a plausible upper bound on how much we spend turning ideas into tangible things.

If we add the 4.0% of GDP spent annually on net investment to the 2.5% spent explicitly on R&D, we get a revised estimate that the US spends 6.5% of GDP per year on creating and building new technologies. If we return to our original estimate for the benefit-cost ratio of R&D, but use s = 0.065 instead of s = 0.025, we get that the benefit-cost ratio is 5.5. Every dollar spent on R&D still generates $5.50 in value!

Does R&D Instantly Impact Growth?

OK, so it’s important to count costs correctly. By the same token, we may believe the benefits of R&D are overstated. The simple framework I laid out above assumed if you pause R&D, you pause growth at the same time. Clearly that’s incorrect.

In reality, R&D is not instantly translated into growth. About 17% of US R&D is spent on basic research - that is, science that is not necessarily directed towards any specific technological application. As I’ve argued before, this kind of investment does eventually lead to technological innovation, but it takes time: twenty years is not a bad estimate of how long it takes to go from science to technology.

Invested at 5% annually, $1 today is worth $2.65 in twenty years. Alternatively, $1 received in twenty years is only worth $0.38 today (since you can invest the $0.38 at 5% per year and end up back with $1 in twenty years). The implication is that benefits that arrive in the more distant future should be more discounted in our accounting framework. For example, if we believed spending R&D resources today only had an impact on growth in 20 years, then we would want to discount our estimate of the benefits to 38% of the levels we came up with when we naively assumed the benefits of R&D arrived instantly. That would imply a benefit-costs ratio of 5.5, as compared the 14.4 we initially computed.

But that’s surely a big over-estimate, since only 17% of R&D is spent on basic science. The other 83% is spent on applied science and development, both of which have much shorter time horizons. Just to illustrate, let me assume 17% of R&D has a 20-year time horizon (38% discount), 33% of R&D has a 10-year horizon (61% discount), and the remaining 50% of R&D has a 5-year horizon (78% discount). In that case, the average discount we should apply, due to the fact that R&D is not instantly translated into growth, is 66%. That implies a benefits-cost ratio of 9.5, as compared the original 14.4. Again - the point is not any specific number. Just that under a lot of sensible assumptions, the return is a lot more than 1!

What about other benefits?

So far we have looked at some ways in which the benefits-cost ratio is over-estimated. Of these, I think the argument that we should include investment as part of the cost of getting a benefit from R&D is a good one, as well as the argument that we should discount the benefits by time since they don’t arrive instantly. Combining those estimates gives us a benefits cost-ratio on the order of 3.6 (i.e., 0.66 * 0.018 / (0.065 * 0.05). Every dollar spent on R&D + investment gets us at least $3.60 in value!

But we also have plenty of reasons why we could argue it is inappropriate to simply use GDP per capita as our measure of the benefits of R&D. There are many benefits from R&D that may not show up in the GDP numbers: reduced carbon emissions from alternative energy sources and greater fuel efficiency; the reduction in work hours that more productive technology has allowed us to realize over the last century; the increased value of leisure time due to the internet; the years of life saved by the covid-19 vaccine; indeed, the years of life saved by biomedical innovation overall.

Jones and Summers take a stab at an estimate for the benefits of biomedical innovation that do not show up in GDP. Biomedical innovation is probably the single largest sector of our innovation system: probably 20-30% of total R&D spending. One way to try and get at the non-GDP benefits of this biomedical innovation is to estimate the value people place on longer lives using things like their spending to reduce their risk of death. I’m not sure how much confidence we want to put in those numbers, but Summers and Jones estimate that a range of reasonable estimates would lead us to increase the estimated benefits of R&D by 20-140%. Taking my tentatively favored benefit cost ratio of 3.6 as our starting line, scaling up the benefits by 20-140% gets us a range of 4.3-8.6.

Estimating the general non-GDP benefits of innovation beyond biomedical innovation is probably an inherently subjective task. But here’s one attempt at a thought experiment to get a sense of how much value you get out of innovations that isn’t reflected in GDP. Suppose there was a magical genie (such things happen in thought experiments) who offered to set you on one of two parallel timelines.

The first is our own timeline, where innovation will happen the same as it has been for a century, and GDP per capita growth will continue to be 1.8% per year. The second timeline is a weird one where technology is frozen at our current level, but (magically) everyone gets richer at a rate of 2.25% per year (as long as you do R&D) - 25% faster than in our current timeline. That is, in the second timeline, you get a bit more money, but you don’t have access to new products and services that innovation would bring. If we were to compute the benefits to cost ratio of R&D in that second world it would be 25% higher than in our timeline, since growth is 25% faster.

If GDP per capita is a good metric of the value of innovation, you’ll clearly choose the second timeline. But if you pick the first one, it means you value access to the newly invented technologies at a level that is at least 25% above their measured impact on GDP per capita.

It’s kind of hard to think what choice you would actually make in this scenario, since choosing between different growth rates is a very foreign decision to most of us. So consider an alternative formulation where the genie offers you the following choice:

A cash payment (right now) equal to 20% of your current income, plus the opportunity to purchase products and services developed between now and 2031

A cash payment (right now) equal to 25% of your current income

Which do you choose?

The first choice is basically where you will expect to be in the year 2031; 1.8% growth compounded over 10 years means you’ll have 20% more income. And in the year 2031, you’ll also have access to all technologies invented between now and then. The second choice gives you a growth rate that is about 25% higher than the other, but no access to the non-monetary benefits of innovation - just the cash. Again, if you pick the first choice, you are saying GDP per capita undervalues the benefits of innovation over the next decade by at least 25%. And so you should scale up your assessment of the benefits to cost ratio of R&D by 25%.

What if option #2 was a payment equal to 30% of your current income? If you would still prefer option #1 in that case, then you think GDP per capita undervalues the benefits of innovation by 50%. And so you should scale up your assessment of the benefits to cost ratio of R&D by 50%. And so on.

What if the average doesn’t matter?

All told, Jones and Summers’ thought experiment essentially argues that R&D is a money-printing machine. Ignoring benefits that don’t accrue to GDP, every dollar you put into your R&D machine gets you back more than $3.60 in value. Possibly much more. So why don’t we use this money printing machine much more? Why are we only spending $2.50 out of every $100 on R&D?

There are two main reasons. First, the value created by R&D is distributed widely throughout society and does not does not primarily accrue to the R&D funder. If I put $1 of my own cash into the R&D machine, I’m not getting back $3.60. Very likely I might get back less than the dollar I put in. The private sector funds about 70% of US R&D and for them the average social return on R&D doesn’t really matter. What matters is the private return that the firm will receive.

But that doesn’t account for why the US government doesn’t spend more on R&D. Presumably, it should care about the social return. One obvious possibility is that decision-makers in government face incentives that don’t reward R&D spending. Maybe election cycles are too short for any politician to get credit for funding more R&D maybe the R&D funded by government gets implemented by businesses who get all the credit; maybe government is just skeptical of academic theory. I don’t know!

But another possibility is that it’s a problem related to knowledge. The average return to R&D must be quite high if we buy the argument just made. But that doesn’t mean the next dollar we spend will earn the average return. Maybe we funded the best R&D ideas first, and every additional dollar is spent on a successively less promising R&D project. Maybe the supply of talented scientists and inventors is already maxed out. If we want to argue that R&D should be increased, we want to know the marginal return to R&D that is, how much extra GDP will get if we spend another dollar on R&D, given what we’re already spending. As I said at the outset of this post, that’s a much harder question to answer. But there are some attempts to answer it, and they also find a quite high rate of return. We’ll look at those next week.

If you liked this post, you might also enjoy the following:

More science leads to more innovation

How useful are learning curves, really?

How long does it take to go from science to technology?

How important are spillovers?

Subscribe at mattsclancy.substack.com

View Details

Science is commonly understood as being a lot more certain than it is. In popular science books and articles, an extremely common approach is to pair a deep dive into one study with an illustrative anecdote. The implication is that’s enough: the study discovered something deep, and the anecdote made the discovery accessible. Or take the coverage of science in the popular press (and even the academic press): most coverage of science revolves around highlighting the results of a single new (cool) study. Again, the implication is that one study is enough to know something new. This isn’t universal, and I think coverage has become more cautious and nuanced in some outlets during the era of covid-19, but it’s common enough that for many people “believe science” is a sincere mantra, as if science made pronouncements in the same way religions do.

But that’s not the way it works. Single studies - especially in the social sciences - are not certain. In the 2010s, it has become clear that a lot of studies (maybe the majority) do not replicate. The failure of studies to replicate is often blamed (not without evidence) on a bias towards publishing new and exciting results. Consciously or subconsciously, that leads scientists to employ shaky methods that get them the results they want, but which don’t deliver reliable results.

But perhaps it’s worse than that. Suppose you could erase publication bias and just let scientists choose whatever method they thought was the best way to answer a question. Freed from the need to find a cool new result, scientists would pick the best method to answer a question and then, well, answer it.

The many-analysts literature shows us that’s not the case though. The truth is, the state of our “methodological technology” just isn’t there yet. There remains a core of unresolvable uncertainty and randomness in the best of circumstances. Science isn’t certain.

Crowdsourcing Science

In many-analyst studies, multiple teams of researchers test the same previously specified hypothesis, using the exact same dataset. In all the cases we’re going to talk about today, publication is not contingent on results, so we don’t have scientists cherry-picking the results that make their results look most interesting; nor do we have replicators cherry-picking results to overturn prior results. Instead, we just have researchers applying judgment to data in the hopes of answering a question. Even still results can be all over the map.

Let’s start with a really recent paper in economics: Huntington-Klein et al. (2021). In this paper, seven different teams of researchers tackle two research questions that had been previously published in top economics journals (but which were not so well known that the replicators knew about them). In each case, the papers were based on publicly accessible data, and part of the point of the exercise was to see how different decisions about building a dataset from the same public sources lead to different outcomes. In the first case, researchers used variation across US states in compulsory schooling laws to assess the impact of compulsory schooling on teenage pregnancy rates.

Researchers were given a dataset of schooling laws across states and times, but to assess the impact of these laws on teen pregnancy, they had to construct a dataset on individuals from publicly available IPUMS data. In building the data, researchers diverged in how they handled different judgement calls. For examples:

One team dropped data on women living in group homes; others kept them.

Some teams counted teenage pregnancy as pregnancy after the age of 14, but one counted pregnancy at the age of 13 as well

One team dropped data on women who never had any children

In Ohio, schooling was compulsory until the age of 18 in every year except 1944, when the compulsory schooling age was 8. Was this a genuine policy change? Or a typo? One team dropped this observation, but the others retained it.

Between this and other judgement calls, no team assembled exactly the same dataset. Next, the teams needed to decide how, exactly, to perform the test. Again, each team differed a bit in terms of what variables it chose to control for and which it didn’t. Race? Age? Birth year? Pregnancy year?

It’s not immediately obvious which decisions are the right ones. Unfortunately, they matter a lot! Here were the seven teams’ different results.

Depending on your dataset construction choices and exact specification, you can find either that compulsory schooling lowers or increases teenage pregnancy, or has no impact at all! (There was a second study as well - we will come back to that at the end)

This isn’t the first paper to take this approach. An early paper in this vein is Silberzahn et al. (2018). In this paper, 29 research teams composed of 61 analysts sought to answer the question “are soccer players with dark skin tone more likely to receive red cards from referees?” This time, teams were given the same data but still had to make decisions about what to include and exclude from analysis. The data consisted of information on all 1,586 soccer players who played in the first male divisions of England, Germany, France and Spain in the 2012-2013 season, and for whom a photograph was available (to code skin tone). There was also data on player interactions with all referees throughout their professional careers, including how many of these interactions ended in a red card and a bunch of additional variables.

As in Huntington-Klein et al. (2021), the teams adopted a host of different statistical techniques, data cleaning methods, and exact specifications. While everyone included “number of games” as one variable, just one other variable was included in more than half of the teams regression models. Unlike Huntington-Klein et al. (2021), in this study, there was also a much larger set of different statistical estimation techniques. The resulting estimates (with 95% confidence intervals) are below.

Is this good news or bad news? On the one hand, most of the estimates lie between 1 and 1.5. On the other hand, about a third of the teams cannot rule out zero impact of skin tone on red cards; the other two thirds find a positive effect that is statistically significant at standard levels. In other words, if we picked two of these teams’ results at random and called one the “first result” and the other a “replication,” they would only agree whether the result is statistically significant or not about 55% of the time!

Let’s look at another. Breznau et al. (2021) get 73 teams, comprising 162 researchers to answer the question “does immigration lower public support for social policies?” Again, each team was given the same data. This time, that consisted of responses to surveys about support for government social policies (example: “On the whole, do you think it should or should not be the government’s responsibility to provide a job for everyone who wants one?”), measures of immigration (at the country level), and various country-level explanatory variables such as GDP per capita and the Gini coefficient. The results spanned the spectrum of possible conclusions.

Slightly more than half of the results found no statistically significant link between immigration levels and support for policies - but a quarter found more immigration reduced support, and more than a sixth found more immigration increased support. If you picked two results at random, they would agree on the direction and statistical significance of the results less than half the time!

We could do morestudies, but the general consensus is the same: when many teams answer the same question, beginning with the same dataset, it is quite common to find a wide spread of conclusions (even when you remove motivations related to beating publication bias).

At this point, it’s tempting to hope the different results stem from differing levels of expertise, or differing quality of analysis. “OK,” we might say, “different scientists will reach different conclusions, but maybe that’s because some scientists are bad at research. Good scientists will agree.” But as best as these papers can tell, that’s not a very big factor.

The study on soccer players tried to answer this in a few ways. First, the teams were split into two groups based on various measures of expertise (teaching classes on statistics, publishing on methodology, etc). The half with greater expertise was more likely to find a positive and statistically significant effect (78% of teams, instead of 68%), but the variability of their estimates was the same across the groups (just shifted in one direction or another). Second, the teams graded each other on the quality of their analysis plans (without seeing the results). But in this case, the quality of the analysis plan was unrelated to the outcome. This was the case even when they only looked at the grades given by experts in the statistical technique being used.

The last study also split its research teams into groups based on methodological expertise or topical expertise. In neither case did it have much of an impact on the kind of results discovered.

So; don’t assume the results of a given study are definitive to the question. It’s quite likely that a different set of researchers, tackling the exact same question and starting with the exact same data would have obtained a different result. Even if they had the same level of expertise!

Resist Science Nihilism!

But while most people probably overrate the degree of certainty in science, there also seems to be a sizable online contingent that has embraced the opposite conclusion. They know about the replication crisis and the unreliability of research, and have concluded the whole scientific operation is a scam. This goes too far in the opposite direction.

For example, a science nihilist might conclude that if expertise doesn’t drive the results above, then it must be that scientists simply find whatever they want to find, and that their results are designed to fabricate evidence for whatever they happen to believe already. But that doesn’t seem to be the case, at least in these multi-analyst studies. In both the study of soccer players and the one on immigration, participating researchers reported their beliefs before doing their analysis. In both cases there wasn’t a statistically significant correlation between prior beliefs and reported results.

If it’s not expertise and it’s not preconceived beliefs that drive results, what is it? I think it really is simply that research is hard and different defensible decisions can lead to different outcomes. Huntington-Klein et al. (2021) perform an interesting exercise where they apply the same analysis to different teams data, or alternatively, apply different analysis plans to the same dataset. That exercise suggests roughly half of the divergence in the teams conclusions stems from different decisions made in the database construction stage and half from different decisions made about analysis. There’s no silver bullet - just a lot of little decisions that add up.

More importantly, while it’s true that any scientific study should not be viewed as the last word on anything, studies still do give us signals about what might be true. And the signals add up.

Looking at the above results, while I am not certain of anything, I come away thinking it’s slightly more likely that compulsory schooling reduces teenage pregnancy, pretty likely that dark skinned soccer players get more red cards, and that there is no simple meaningful relationship between immigration and views on government social policy. Given that most of the decisions are defensible, I go with the results that show up more often than not.

And sometimes, the results are pretty compelling. Earlier, I mentioned that Huntington-Klein et al. (2021) actually investigated two hypotheses. In the second, Huntington-Klein et al. (2021) ask researchers to look at the effect of employer-provided healthcare on entrepreneurship. The key identifying assumption is that in the US, people become eligible for publicly provided health insurance (Medicare) at age 65. But people’s personalities and opportunities tend to change more slowly and idiosyncratically - they also don’t suddenly change on your 65th birthday. So the study looks at how rates of entrepreneurship compare between groups just older than the 65 threshold and those just under it. Again, researchers have to build a dataset from publicly available data. Again every team made different decisions, such that none of the data sets are exactly alike. Again, researchers must decide exactly how to test the hypothesis, and again they choose slight variations in how to test it. But this time, at least the estimated effects line up reasonably well.

I think this is pretty compelling evidence that there’s something really going on here - at least for the time and place under study.

And it isn’t necessary to have teams of researchers generate the above kinds of figures. “Multiverse analysis” asks researchers to explicitly consider how their results change under all plausible changes to the data and analysis; essentially, it asks individual teams to try and behave like a set of teams. In economics (and I’m sure in many other fields - I’m just writing about what I know here), something like this is supposedly done in the “robustness checks” section of a paper. In this part of a study, the researchers show how their results are or are not robust to alternative data and analysis decisions. The trouble has long been that robustness checks have been selective rather than systematic; the fear is that researchers highlight only the robustness checks that make their core conclusion look good and bury the rest.

But I wonder if this is changing. The robustness checks section of economics papers has been steadily ballooning over time, contributing to the novella-like length of many modern economics papers (the average length rose from 15 pages to 45 pages between 1970 and 2012). Some papers are now beginning to include figures like the following, which show how the core results change when assumptions change and which closely mirror the results generated by multiple-analyst papers. Notably, this figure includes many sets of assumptions that show results that are not statistically different from zero (the authors aren’t hiding everything).

Economists complain about how difficult these requirements make the publication process (and how unpleasant they make it to read papers), but the multiple-analyst work suggests it’s probably still a good idea, at least until our “methodological technology” catches up so that you don’t have a big spread of results when you make different defensible decisions.

More broadly, I take away three things from this literature:

Failures to replicate are to be expected, given the state of our methodological technology, even in the best circumstances, even if there’s no publication bias

Form your ideas based on suites of papers, or entire literatures, not primarily on individual studies

There is plenty of randomness in the research process for publication bias to exploit. More on that in the future.

If you liked this post, you might also enjoy the following:

How bad is publish-or-perish for the quality of science?

How a field fixes itself: the applied turn in economics

What ails the social sciences?

Subscribe at mattsclancy.substack.com

View Details

As a source of data for studying innovation, patents are really seductive. There’s nothing else quite like them:

detailed descriptions of millions of inventions…

dating back over a century...

from all corners of the economy...

validated by an expert as novel, useful, and non-obvious.

At the same time, they have real biases. Not every invention gets patented (in fact, it may be more common that inventions don’t get patented; a topic for another day), and not every patent represents a useful invention. Worse, the set of things that get patented isn’t just a random sample. There are systematic differences in the kinds of things that get patented and the kinds of things that don’t. It’s a constant source of tension for me, as someone who writes about social science research on innovation, the majority of which is based on patent data: does this cool result hold up when we use alternatives to patent data?

What we’re going to talk about today though is just one aspect of the patent problem: what are patent citations really telling us?

Patent citations are the citations patents make to “prior art”: those include other patents, as well as academic journal articles, patent applications, legal documents, and other stuff. We’re just going to focus on citations to patents today - many of the issues highlighted here don’t necessarily apply to the other kinds of documents.

Why do we care about patent citations? Because one of the most important things about innovation - as compared to other activities - is that knowledge spills over to new applications. Understanding exactly how and why that happens is a big part of the problem of understanding how innovation happens, and citations promise a very clean way to identify these spillovers. Or at least, they seem to. (See a list of posts I’ve made about papers using patent citations to measure knowledge flows at the end of this post)

I think most researchers who begin to study patents (including myself, when I began to use them as a dataset) start out by thinking of patent citations in the same way that we think of the citations we make in our academic work. For an academic, citations are simply addresses for ideas we are referencing, whether to build on them, critique them, or nod to them. In short, they provide a list of the ideas we are engaging with when we make our own intellectual contribution. If a patent is basically like the invention equivalent of a research project, and a patent citation is basically like the citation in a research paper, then they are a great way to measure the flows and uses of knowledge.

And sometimes, that’s basically what a citation in a patent is! But just as often, it’s not. In this post I’m going to walk through some of the issues we have in using citations as a proxy for knowledge flows, but then ultimately argue they’re still useful in some contexts, and especially when used in concert with different forms of evidence that have different strengths and weaknesses.

Patent Citations Have Many Authors

A first important distinction between citations in academia and citations in patents (to patents) is that whereas the former tend to be added exclusively by the authors, many parties besides the inventor add citations to a patent. This gets to the differing rationales for citations in academic papers and patents.

Citations in patents are mainly about establishing an invention is eligible for a patent, in the sense that the invention is novel and makes a non-obvious improvement on any existing work. In the USA, inventors have the duty to disclose all relevant information they are aware of (or risk having their patent invalidated upon challenge), which could imply inventors will cite any patents for inventions whose ideas they improved on. But many other citations may be added to show the invention is legally patentable (i.e., citing a famous patent for a GMO crop to establish GMO crops are patentable), or that it’s improvements are not obvious (i.e., perhaps by citing improvements made by other patents, only to argue they are distinct). Importantly, these citations can be added by patent attorneys or the patent examiner evaluating the application for a patent. Or citations may be added by the inventor, after they have “completed” their invention and begin to do the research needed to secure a patent.

So, for a variety of reasons, it may be that a patent is cited even though the inventor was completely unaware of it while doing the R&D that resulted in the invention. To get a sense of how common this is, we can look at a survey by Jaffe, Trajtenberg, and Fogarty, conducted in the 1990s. They simply used the addresses that inventors listed on their patents to mail a bunch of them surveys asking about a citation they had made. Only 38% of 166 respondents knew about the cited patent before or during the invention process; that is, the majority of citations do not represent “knowledge flows” at all, since the inventor only became aware of the cited patent after the invention was complete.

Moreover, in 2002 the US Patent and Trademark Office began to report when a citation was added by a patent examiner, the applicant, or “other parties.” For citations made between 2005 and 2014, around a quarter of citations were added by the patent examiner. Again, that means a large share of citations don’t seem to be measuring knowledge flows, in the sense that they’re not even added by the inventor.

What’s the bottom line? A fraction of citations probably do correspond to genuine knowledge flows - but only a fraction. In Jaffe, Trajtenberg, and Fogarty’s survey, they ask respondents what they learned from the cited patent, and many of the answers correspond to the kind of thing we want citations to be: inventors said they learned about a concept that could be improved, that the idea was technically feasible, or other information useful for development. These kinds of citations are probably in the minority, but they’re there.

Inventors May Play Games

So far we have two problems with citations. First, they are frequently not added by the inventors themselves. Second, even the citations an inventor adds do not necessarily serve as a simple record of the ideas that were useful in the invention process - that’s not what the citations are for. But a third problem is when patent applicants intentionally do not cite all relevant patents, but strategically cite documents as part of a strategy that they believe will help make their application more likely to be granted or their patent less likely to be invalidated.

As noted above, applicants have a legal duty to disclose all relevant information; failure to do so means a patent could - in theory - be invalidated in a subsequent court challenge. This creates an incentive not to withhold relevant information. On the other hand, if you draw the patent examiner’s attention to the existence of a patent covering some aspect of your invention, you may have to settle for a narrower set of claims about what’s protected under your patent. So there is a gamble in play - or at least, the perception of a gamble by the applicant: if you can get away with citing less, you may be able to get a patent covering a wider range of things, but at the risk of your patent not holding up in court. It looks like how patent applicants think about this gamble matters and has changed over time.

Take Lampe (2012). Lampe is looking for evidence that applicants are intentionally withholding relevant citations during their applications. To do that, he makes the assumption that applicants probably know about patents that they or their coauthors have previously cited in other patents. Then, he looks at the citations the patent examiner added; these are citations that the examiner has decided are relevant to the patent application. If the applicant knew about these patents but did not supply them, and the patent examiner thinks they are relevant, it’s possible the applicants were trying to sneak something by the examiner. (Of course, it doesn’t prove anything, but it’s suggestive.)

Does this happen much? Yeah!

To see how much it matters, let’s pick a patent and scrutinize its citations. Let’s focus on the set of citations it makes to patents that had also been previously cited by one of the co-inventors on this patent. We’re going to assume the co-inventors knew about these patents when they made their application. Even though the co-inventors knew about these patents, about 20% were not added by the applicants, but instead judged relevant and added by the examiner.

Well, maybe it was an honest mistake. But slightly more suspicious is the fact that the extent of this withholding is lowest for the patents typically considered most valuable (and therefore with the most to lose if the patent gets invalidated); that is patents for drugs and chemicals technologies, or for patents that get the most citations in the future (a common proxy for the value of a patent). On the other hand, larger firms are more likely to withhold potentially relevant citations, possibly because they have more resources to defend challenged patents or maybe because holding so many patents makes them less risk averse.

That all indicates that citations may be missing relevant work. We don’t know what applicants fail to cite and the examiners fail to catch. But there is also the opposite problem: citation of irrelevant work.

Why would a patent applicant cite irrelevant work? One potential rationale is that it could be another way to try and sneak something past the examiner (or, as importantly, the applicant might believe this, whether or not it’s true) by hiding an important citation in tons of meaningless citations, so that the examiner doesn’t have time to scrutinize it. But a more benign rationale could be that the applicant has an overly generous interpretation of the duty to disclose all relevant information. Since a patent can be invalidated for failing to cite relevant material, and since it’s costless for the applicant to cite more things, why not cite everything that is even remotely relevant? This problem can be especially salient for an applicant submitting many linked and interrelated patents; instead of trying to parcel out which citation should be rolled over from one patent to another, why not just copy them all?

Unlike omitted citations, there is some evidence that this problem has become much more severe in the last decade. Kuhn, Younge, and Marco (2020) documents the rise of super-citing patents - a relatively small share of patents who cite so many patents, that they skew the entire landscape of citations. It is most common for patents to cite less than 20 other patents; in 2014, 75% of patents fell into this category. Sometimes, however, it is appropriate for a patent to cite more patents, perhaps up to 100. In these cases, it becomes difficult for a patent examiner to carefully check every citation offered. Still, 95% of patents in 2014 made less than 100 citations, with the vast majority making less than 20.

That remaining 5% is causing problems. These patents - which were all but non-existent prior to the year 2000 - cite more than 100 patents each. In fact, this small number of super-citing patents now accounts for nearly half (46%) of all patent citations, even though they comprise under 5% of patents!

Moreover, the quality of the citations made by these super-citers is highly dubious. Kuhn, Younge, and Marco (2020) compute the similarity of the text of citing and cited patents (based on the degree to which they contain the same words that are otherwise uncommon). As indicated in the figure below, the similarity of citing and cited patents gets steadily worse, the more citations a patent makes.

What’s particularly worrying is Kuhn, Younge, and Marco (2020) show the rising share of low-quality citations is eroding the usefulness of citations for studying patents. The average textual similarity of citing and cited patents has been declining for decades, as the share of citations associated with super-citers rises.

They also replicate a few canonical results from the economics of innovation, that rely on patent citations, and show that these results are affected by the decline in the quality of citations.

For example, a famous result in the patent literature showed that firms whose patents receive more citations have higher stock market valuations than otherwise observationally similar firms whose patents receive fewer citations. But the magnitude of this correlation has halved between 2003 and 2008; patent citations just don’t seem to be “worth” as much as they used to be, as judged by the market.

Another study - which I’ve mentioned before - used patent citations in the 1980s and 1990s to measure local knowledge flows. Essentially, they showed patents were more likely to cite the patents of local inventors as compared to distant inventors of the same kind of technology. This has long been an important line of evidence about the importance of local knowledge, and why innovation tends to happen in cities. Kuhn, Marco, and Younge update this study and show that the results differ significantly if you try to control for the rise of low-quality patent citations. If you do not control for them, the propensity to cite local work has remained stable and consistent; if you adjust for the quality of citations, this propensity has fallen considerably.

Time to give up?

So, uh, citations have problems. But it’s important to remember that, at the end of the day, there is genuinely useful information in a subset of patent citations and that some information is better than none. To begin, we have that old survey evidence that the inventor knew about 38% of the citations on their patent before or during the inventive process. And in another survey from the 1990s, inventors (this time in the EU) rated the patent literature as one of the most important sources of knowledge used to develop innovations (though the survey did not ask if the debt to other patents was reflected in citations).

In the last decade, partially in response to the dawning recognition that patent citations are not nice analogues for academic citations and partially due to the growing sophistication of natural language processing tools, the research community has begun developing much better tools for analysing the raw text of patents. Increasingly scholars are tracking knowledge flows by looking at the similarity of the textual description of inventions in patents. Somewhat reassuringly, there is a lot of overlap between textual similarity and citations. Younge and Kuhn (2016) show that the similarity of text between patents is as good a predictor of citation as other methods based on the US patent classification system, while Feng (2020) shows the text of patents with a direct citation link between each other are 2.5x as similar to each other as a baseline, which is about the same as patents that share an inventor.

Given all that, is it time to give up on patent citations? I don’t think so. They’re an imperfect source of information - but that’s life in the social sciences. The best we can do is understand the strengths and weaknesses of the data, and try to find cases where different kinds of data tell a mutually confirming story. In this newsletter, wherever possible, I try to complement patent-based papers with others. It’s not always possible (patent data is one-of-a-kind), but when it’s not, I consider the results more provisional than they would otherwise be, especially if the citation data is of a more recent vintage.

Here are just some posts on work that uses patent citations as a proxy for knowledge flows:

Ripples in the River of Knowledge

Adjacent Knowledge is Useful

Proximity: More important for meeting than collaborating

Cities as a platform to mix up knowledge

Innovation in the city

Subscribe at mattsclancy.substack.com

View Details

In More Science Leads to More Innovation, we looked at four natural experiments where the “supply” of science was increased or decreased differently across scientific fields. When the supply of science increased, we saw more downstream technological innovation, and when the supply of science decreased, less. In this post, I’ll argue those studies underestimate the influence of science on innovation.

Direct dependence on science is uncommon

To begin, while it’s true that more science leads to more innovation, the majority of technological innovations probably do not directly depend on recent science:

Citing academic articles in patents has become more common over time, but even in 2018, 74% of patents did not cite any scientific journal articles.

In a survey from the 1990s, European inventors rated the importance of knowledge from scientific literature for developing innovations at 2.5 out of 5 (lower than they rated the importance of knowledge from customers/users and the patent literature). They rated the importance of universities and public research laboratories at just 1.4 out of 5.

In a 1994 survey, US R&D managers estimated only 20% of R&D projects relied on public research.

Instead, the dependence on science is unevenly distributed. The figure below illustrates the average number of citations to science per patent, by technical classification. Patents in chemistry/metallurgy, and human necessities (which include the biomedical and pharma sectors) cite science much more intensively than other fields. Fields like mechanical engineering barely cite the scientific literature at all.

This is broadly consistent with a 1994 survey of corporate R&D managers that found R&D projects in automobiles, general manufacturing, and electrical equipment relied on public research significantly less than in fields like biotechnology and pharmaceuticals.

Indirect Dependence on Science?

But just because an invention doesn’t rely directly on science doesn’t mean that science plays no role in it. Scientific knowledge and principles can become embodied in technologies that other technologies, in turn, use as components. For example, chemistry and metallurgy heavily cite the scientific literature. It may be that new chemistry production processes (which are patented) allow for the manufacture of new kinds of composites, that in turn allow for, say, the creation of new, more powerful engines. But then the patents for the composites and the engines may not cite science and the inventors of these things might not report any dependence on science, even though without the production processes enabled by the science they would be out of luck.

There is actually a way to measure this “distance from science” that we’ve discussed before. In the example above, although a new type of composite might not cite any scientific articles itself, it might cite the patents for the new chemistry production processes that do. And the engine patent might cite the composite patent, which in turn cites science-based production processes. Ahmadpoor and Jones (2017) use this basic idea to measure the “distance” from science of US patents by counting the smallest number of citation steps between a patent and a scientific article. A patent that cites a scientific article has a distance of 1. A patent that cites no science itself, but does cite a patent that cites a scientific article has a distance of 2. And so on. In Ahmadpoor and Jones’ sample of patents from 1976 to 2013, although only 16% of patents directly cite a scientific article, 61% of patents are “connected” to science via some kind of chain of citation (most often, a distance of 3).

There is some indirect evidence that this measure of distance from science is capturing something real. Patents that are “closer” to science as measured in this way have some of the characteristics of patents that directly cite science. For example, as discussed in other posts, patents that cite scientific articles are more valuable than those that do not, and also more likely to be traded. But it’s also true that, looking at patents that don’t directly cite science, those that are closer to science are still more valuable and more likely to be traded than those that are farther from science.

Most of the natural experiments discussed in More Science leads to More Innovation pertain to patents with a distance of 1 (the closest to science). Indeed, half of them explicitly measure the link between science and technology via a citation from a patent to a journal article (which means, by definition, they have a distance of 1). But if there is a knock-on effect for patents further from science (distance 2 or greater), they probably miss it.

Upstream Patenting Predicts Downstream Patenting

Another strand of literature gives us some good reasons to think there are significant knock-on effects. Technologies tend to be hierarchically composed of many sub-technologies, and to build on each other in ways that are relatively stable and predictable over multiple years. This means there is some degree of predictability about technological trends. If there is a flurry of breakthroughs in an upstream technology, downstream technologies that use it as an important component, or which adapt its principles and uses for new contexts, are likely to see a flurry of breakthroughs in subsequent years. (Think of how we might be reasonably confident that all sorts of new and improved engine designs will be enabled by a new and improved composite)

There are ways to observe this hierarchy and use these relationships to make predictions. US patents are classified as primarily belonging to one of several hundred technology classifications (examples range from “Class 012: Boot and shoe making” to “Class 706: Data processing – artificial intelligence”). The hierarchical relationship between these technology classes can be observed in the citations of the patents belonging to these classes. Acemoglu, Akcigit, and Kerr (2016) build a directed network between different technology classes, where the strength of a link between two classes is given by the probability a patent in one cites the other.

Acemoglu, Akcigit, and Kerr show a statistically significant relationship between patent activity in upstream classes and the patenting of downstream classes (that is, the ones that historically cite this class heavily). Pichler, Lafond, and Farmer (2020) perform a similar exercise. In the figure below, they plot the correlation between the growth rate of patents in a given class, and the growth rate of the weighted average growth rate of upstream technology classes.

It turns out these correlations are robust enough to be used for forecasting. In one application, Acemoglu, Akcigit and Kerr use data from 1975 to 1994 to fit their statistical model, and then they use it to predict the number of patents in the following ten years. After adjusting for the influence of technology classification (some classes always patent more than others) and time (in most classes there tend to be more patent applications per year), they find a 10% increase in predicted patenting (based on the growth of patenting in upstream classes) is associated with an actual out-of-sample 3-4% increase in patenting.

Pichler, Lafond, and Farmer predict the growth rate of patenting as a function of upstream patenting activity using methods derived from machine learning. They fit a number of alternative models based on data from 1945-1987 to model the correlation between the growth rate of patenting in each technology class and prior patenting growth in upstream technologies. They then choose the model that makes the best forecast for 1988-2002. Finally, they use all the data from 1945-2002 to refit this model and predict out-of-sample patent growth rates, in each class, over 2003-2017.

How well does it do? To understand it’s performance, they need a benchmark. They replicate the whole process above, but for models that exclude data on upstream patenting. Instead, the benchmark predicts patenting in class x by the historical patenting activity of just class x (is patenting in this class rising or falling over time? Does it tend to move in booms and busts? And so on). Again, they find the model that best predicts out-of-sample from 1988-2002, re-estimate is with data up to 2002, and then forecast out-of-sample through 2017.

The results are in the figure below, with the relative performance of the models using upstream patent data in blue (the green is a model not discussed in this post). At it’s peak, models using data on upstream patenting gain nearly 40% in predictability relative to a benchmark.

It all boils down to this: historical patent citations allow us to identify the technology classes that lie “upstream” of any other class; and upstream patenting predicts downstream patenting in the future, out of sample, in two different papers.

Upstream = Closer to Science?

So we have two related but different ways of measuring indirect knowledge flows among patented technologies. Some papers have measured the distance from science, via the shortest citation chain to a scientific paper. Others have defined upstream and downstream relationships among technologies based on the total share of citations that flow from one technology class to another. A natural question is the extent to which the two line up. That can tell us something about how science indirectly impacts technology.

For example, we know that science tends to lead to more innovation in technology classes that directly depend on science. Are these technology classes, in turn, directly upstream of many other classes? If so, the results from Acemoglu-Akcigit-Kerr and Pichler-Lafond-Farmer would predict an increase in science-based innovation would lead to a second round of innovation in the technologies that lie immediately downstream. And if these classes are themselves upstream of many other classes, there would be a second reverberation, and so on.

In the figure below, I computed the average distance to science for US patents over 1976-2018, based on the Ahmadpoor and Jones method and using Marx and Fuegi’s dataset. On the horizontal axis, we have the average distance to science of the patents belonging to each of 307 different technology classes. Since I’m interested here in the indirect impact of science on technology, I limited my attention to technology classes with distance of 2 or greater: that is, technologies that do not typically cite scientific papers directly. Classes to the left are closer to science, classes to the right are farther away. On the vertical axis, we have the average distance to science of their upstream technology classes, weighted by citation share. The lower the dot, the closer to science are the classes cited.

Let’s look at an example. In the upper right corner, we have a red dot corresponding to Class 81: Tools. On the horizontal axis, we see the average patent in this class has a distance to technology of 4.5. That means the shortest distance to science for tool patents often involves citing a patent that cites a patent that cites a patent that cites a patent that cites a science article. On the vertical axis, we see the typical distance to science for technology classes that are heavily cited by tools patent is just 3.6. These upstream classes include classes like Class 29: Metal working (average distance to science 3.0) and Class 30: cutlery (average distance to science 4.1). The main take-away is that classes directly upstream of tool patents tend to be closer to science than tool patents.

The black line cutting through the middle of this figure demarcates the split between technologies that mostly cite technologies closer to science (dots below the line) and those that mostly cite technologies farther from science (dots above the line). For the classes displayed, 78% of them lie below this line – that is, the technologies that lie upstream also tend to lie closer to science. This difference isn’t uniform though. Looking only at technologies far from science, with a distance of 3 or greater, 95% lie downstream of technologies closer to science then they are! But looking at classes in the 2-3 interval, it’s basically 50/50. Essentially, that means technologies that are 1-2 citation steps removed from science largely cite each other; they don’t primarily build on technologies that are closer to science. But technologies further out do.

Indirect Impact of Science

So where do we stand? In More Science leads to More Innovation, we looked at pretty compelling evidence that increasing the supply of science tends to lead to more innovation. But that direct effect is concentrated in a relatively small share of technologies; only a quarter of patents directly cite scientific work. However, any innovation has spillover effects. As discussed elsewhere, the magnitude of the unintended benefits of science tend to be at least comparable to the intended benefits. In this post we focused on specific kind of spillover: an increase in innovation tends to lead to further innovation in “downstream” technologies (which we can identify based on citation patterns).

We have no reason to think that spillover effect wouldn’t hold if innovation increased because of science. Any field that sees an increase in innovation due to science will probably have some downstream fields that will also benefit. Initially, these downstream fields might be relatively “close” to science themselves, so they might also directly benefit from an increased supply of science. But eventually, the passing of technological concepts and improved components from upstream to downstream becomes a channel through which the fruits of science might also measurably flow.

If you liked this post you might also like:

More science leads to more innovation (it’s basically a prequel to this post)

How important are spillovers? (very)

How useful are learning curves really? (casting a slightly critical eye on a favorite method of forecasting technological progress)

Subscribe at mattsclancy.substack.com

View Details

Like the rest of New Things Under the Sun, this article will be updated as the state of the academic literature evolves; you can read the latest version here.

Here’s another obvious idea that turns out to work: if you want more technological innovation, more science helps.

That probably shouldn’t come as a surprise: lots of inventors (not all!) say in surveys that science is an important input to their inventions and about a quarter of patents in 2018 directly cite scientific papers. And we can think of countless examples of technologies with their roots in fundamental scientific advances. But none of that necessarily means that investing in more science will lead to more innovation. For example, maybe some inventors intensely rely on a tiny sliver of useful science, but the bulk of scientific work never leaves the ivory tower. And maybe we already do science that’s most likely to be useful, so that any additional science will be of the latter variety.

To see if that’s the case, we can turn to a few examples where scientific production in some fields either increased or decreased unexpectedly, and then see what happens to technology that tends to rely on that science, relative to other technologies. As we’ll see, when you increase or decrease scientific output in a field, that tends to respectively increase or decrease the use of science in related technological domains.

Geopolitical Shocks to Science

Let’s start with a geopolitical shock that decreased science, rather than increased it: World War I. Leading up to World War I, science had increasingly become an international endeavour, with scientists reading articles published all over the world and meeting frequently at international conferences. This all changed, however, when the world split into the Allied powers, the Central powers, and the Neutral states in World War I. Iaria, Schwarz, and Waldinger (2018) show that the onset of war led to huge delays in the receipt of new scientific journals from countries belonging to a scientist’s enemies. For example, here is how the gap between a journal’s publication and arrival in the Harvard library evolved over 1910-1930 for journals published in Central power countries and Allied countries (which the USA was a part of).

As the figure above indicates, delivery of articles from Central power scientists to the USA got delayed for nearly two years during the war, while articles from allied nations saw barely any delay at all. The Central powers saw similar delays of receipt of Allied power journals. Or consider the following image of attendees to the famed Solvay conferences for top international physicists. We can note two things. First, the absence of any conference at all between 1913 and 1921. Second, even after hostilities ceased, things did not return to their pre-war norm. The circled faces belonged to Central power scientists: note their absence between 1921 and 1927.

Iaria, Schwarz, and Waldinger also show a substantial drop in citations to articles published by scientists from the “other side” during this period. And lastly they show, using text similarity analysis, that the titles of journal articles (after being translated into English) began to become less and less similar between scientists on opposite sides during the war.

Taken together, it’s clear the war cleaved the large international scientific community into two smaller ones. But the impact was different across fields and countries. In some cases, for example, US biology, there wasn’t much of a disruption at all, because the field tended to rely on domestic or allied research. But in others, like US biochemistry, scientists heavily cited and used work produced by Central power scientists, to which they no longer had ready access once the war began. Iaria, Schwarz, and Waldinger compare the effect of this disruption for science fields and countries that relied heavily on outside science to those that didn’t.

They show that scientists in fields that most regularly cited top articles by scientists from the other side produced fewer publications than scientists who were not so reliant on foreign science. Moreover, the quality of the work they did publish was also reduced. For example, scientists in these fields saw a reduction in the probability they would produce work good enough to be nominated for a Nobel prize.

Most importantly for the purposes of this post, Iaria, Schwarz, and Waldinger have an interesting way to track linkages from science to technology by looking for new and novel words that first appear in the titles of scientific articles, and later appear in US patents. These words included things like “magnetron” or “electroencephalogram.” As a general rule, on average every year a scientist in their dataset produces new words that pop up in the text of 0.43 patents down the road. However, scientists most reliant on top foreign science tended to produce fewer of these technological/scientific words once they were unable to access science from abroad.

In short, it seems a reduction in science caused by wartime disruption to knowledge flows reduced the production of new scientific ideas that would normally have been picked up and used in patents in later years.

That was a long time ago and World War I was a very unique event; but the results are pretty similar to a more recent geopolitical shock that affected research differently across different fields.

Arora, Belenzon, and Suh (2021) is not primarily concerned with how science leads to technology, but instead documents myriad ways that science facilitates trade in patented technology. But for the purposes of this post, we’ll focus on a part of their study where they identify another scenario where funding unexpectedly changes for different fields. In their case, they use the unexpected collapse of the USSR and the end of the Cold War as a surprise shock to US science funding. The end of the Cold War resulted in a significant reallocation of US federal R&D dollars away from areas like defense and defense-related sciences such as physics, chemistry, and electrical engineering, and towards new scientific areas like computer science, oceanography and biology. Arora, Belenzon, and Suh show that these shifts in funding had the kind of effect you would anticipate - more funding led to more scientific papers in some fields and less funding led to fewer papers in others, relative to the counterfactual. Since the collapse of the USSR and subsequent reshuffling of research dollars was an unexpected event, if you’re a private sector research firm in the early 1990s, from your perspective there was a sudden windfall of new research in some fields and a deficit in others, relative to your expectations.

Arora, Belenzon, and Suh show that technologies that typically relied on research from the sciences that benefited from the end of the Cold War responded by citing that research more often. The share of patents citing scientific papers was correlated with the surge (or decline) of scientific papers brought about by reallocation of research dollars. And in other parts of their paper, they confirm previouswork that patents citing scientific research tend to be more valuable by a variety of metrics. Patents citing scientific research:

Receive more citations from other patents

Are more likely to be bought and sold

Lead to a larger increase in the stock market valuation of the firms owning them

In short, the end of the Cold War disrupted different science fields in different ways; some fields got more funding and generated more research, others the opposite; technological fields associated with the “winning” fields cited more research when more became available; and the patents that cite science look to be more valuable.

The Benefits of Science Funding Windfalls

Both of the above papers use massive geopolitical shocks to document how big changes to the production of science get reflected in technology down the road. But while they are very suggestive, they don’t exactly attempt to show that more science necessarily leads to more technology. Instead, they show that a decline in science leads to a reduction in the use of science by technology, and an increase in science leads to an increase in the use of science by technology (and also that more science-based technology seems to be more valuable). But two other papers look at much smaller scale changes in science production to show more explicitly that increasing the production of science tends to increase technological innovation as well.

To start, let’s look at a clever 2019 paper by Tabakovic and Wollman about football and scientific research. One peculiarity of the US university system is that for some schools a non-negligible share of university revenues come from college football. For universities in the NCAA, when their team does unexpectedly well, the university receives windfall funding from licensing broadcast rights, merchandise, and alumni fundraising. The amounts can be quite large - football revenues are equal to more than a third of tuition revenue at Louisiana State University. Much of this funding finds its way into supporting university research in the subsequent year.

Tabakovic and Wollman exploit this idiosyncracy in US funding mechanisms to identify universities that receive unexpected research dollar windfalls. Specifically, they look at how university support for research changes when teams outperform or underperform preseason expectations as measured by votes in the NCAA top 25 AP poll. As the left-most panel in the figure below indicates, universities that outperform expectations (receive more votes in the poll at the end of the season than at the beginning) get more institutional funding for research in the following year. Meanwhile, the other two panels serve as a bit of a sanity check; they indicate football performance has no impact on the receipt of federal grants or other forms of funding, which is what we would expect.

What happens when researchers get more money? Tabakovic and Wollman estimate that a 10% increase in funding is associated with 3% more publications in the following year. And it turns out the funded science also leads to patents: they find about $2.6mn in funding for university research is associated with about one more patent. Lastly, they can use the revenue the university earns by licensing out these patents to get a rough estimate on whether the patents are valuable enough for a private sector firm to pay for access to them. And they are - a 10% increase in funding is associated with at least a 10% increase in licensing revenues.

Those estimates are noisy, but they’re more-or-less consistent with one of my favorite recent studies, which also exploits a similar special case when funding is more-or-less randomly handed out to different groups of scientists.

Azoulay et al. (2019) want to know how NIH funding for basic biological research eventually leads to private biomedical innovation. To measure biomedical innovation, Azoulay and coauthors identify a large set of biomedical patents. The pharmaceutical sector relies on patent protection to an unusually large degree, and so for this sector patents tend to be a better measure of actual innovation than might be the case for other sectors. They then go on to link these biomedical patents to journal articles via citations. They then go a step further and link journal articles to grants from the NIH. This lets them see in unusual detail the entire pipeline from funding for basic research to technological innovation. They can see if more NIH funding is associated with more journal articles, and more journal articles is associated with more patents, with each link in the chain observable with direct citations. And it is! In general, an extra $10mn for a given disease-science area is associated with 2.6 more related patents. That’s $3.8mn per patent, which is not too different from what Tabakovic and Wollman find.

This finding holds even when you include lots of control variables, such as scientific field-specific variation over time. For example, maybe gene sequencing technology means genetic studies get a lot more bang for their buck than in the past, and the federal government responds by giving more grants to scientists proposing genetic research. In principle, we might worry that inflates the apparent value of funding, since the fields that get more money were more likely to generate valuable knowledge even if they got less money. But in that case, Azoulay et al. (2019) show their results still hold when you compare within this scientific area at the same point in time. That is, even comparing two different proposals, relying on the same science at the same time, the one that gets funding (for one disease) is associated with more patents then the one that doesn’t (and which is associated with a different disease).

But the paper goes further than that. Azoulay and co-authors exploit idiosyncrasies in funding to get plausibly random variation in which grants are funded and which are not. To simplify, it’s kind of like this: all grant proposals belong to both a scientific area and a disease area. Proposals are scored by scientific committees. These scientific committees review proposals associated with several different diseases, as long as they are in the same scientific area. But then proposals are handed off to be funded or not based on how well they rank compared to other proposals for the same disease - not the same science.

That means you can have situations like the following. Suppose one proposal scores very well (say, an 8/10), but because the other proposals in the same scientific area are even better, it gets a low ranking from it’s science committee (say, ranked 5 out of 5). Meanwhile, in another scientific committee, another proposal faces the opposite situation: it is poor quality and gets a poor score (say, 4/10), but since the other proposals are even worse, it gets a high ranking from its science group (say, ranked 1 out of 5). If both these proposals belong to the same disease group, you can end up with a situation where the weaker proposal might get funded simply because it had a higher ranking. But that ranking is less due to the proposal’s inherent quality than it is to the quality of the other proposals in its same scientific area.

Azoulay and coauthors look at the 10 proposals that straddle the funding cut-off for a given disease group; 5 are above the cut-off and get funded, 5 are below and don't. Down towards this cutoff, the uncertainties discussed above loom large and whether a given proposal lies above or below this threshold is largely a matter of luck. The authors create measures of "windfall" funding based on whether a disease-science area has more than 50% of it's proposals lying above or below this cut-off. In that way, they can see what happens to innovation in scientific fields that got a couple extra million dollars of grant money, compared to those that just missed out. Using this more complicated approach they find basically the same thing - a random extra $10mn results in about 2.3 more related patents. But in this case, we can more confidently assert the relationship is causal: (biological) science leads to (biomedical) innovation down the road.

Beyond Patents?

So, to sum up:

Scientific fields more disrupted by war produced fewer new scientific words that go on to show up in patents than less disrupted fields

Scientific fields that benefited from the restructuring of R&D during the end of the cold war produced more papers, which got cited by more technologies that typically rely on science from those fields, and in general patents citing science tended to be higher quality

Universities that got windfall funding on the unexpected strength of their football teams produced more papers, patents, and technology licensing revenue

Science-disease areas that received more funding due to idiosyncrasies in NIH funding rules produced more papers (which explicitly acknowledge grants support), and more patents (which explicitly cited these papers)

And there’s one more line of supporting evidence. One potential downside of all these papers is that they rely on patents to measure technological innovation. The strength of patents in this case is they provide such detailed data in the form of citations, text, and licensing fees that we can really nail down the link between science and technology. But at the same time, patents are a highly imperfect measure of technological innovation. But if we’re willing to look at messier correlational data, we also have some correlational data on R&D and measures of industry-wide total factor productivity. Higher levels of basic R&D tends to be associated with higher productivities in related industries with a lag of about 20 years (a lag that happens to match reasonably well the lags between patents and the research they cite).

Taken all together, in addition to our prior beliefs that modern technology relies to a large degree on better science, it makes a pretty compelling case. More science tends to lead to more innovation. It’s not the only thing that leads to more innovation, and it doesn’t lead to innovation in all fields, but it does work.

If you liked this post you might also like:

How useful is science? (answer: pretty useful!)

Does chasing citations lead to bad science? (answer: not for the most part)

How long does it take to go from science to technology? (answer: 20 years is a good rule of thumb)

Subscribe at mattsclancy.substack.com

View Details

Sometimes obvious ideas work. If you want to encourage more innovation, give people better access to knowledge.

Let's start in the 1800s. Over 1883-1919 (but mostly after 1899) Andrew Carnegie provided the funds for the construction of ~1700 free public libraries, scattered across the USA. There was even one in my hometown.

Berkes and Nencka (2020) sets out to measure the impact of this library building spree on innovation. They want to compare the patenting rate of cities and towns that received libraries to otherwise identical ones that did not. The challenge is finding cities that did not receive libraries, but otherwise serve as a good control group. What Berkes and Nencka use is a set of 200 towns that applied for library funding and were approved for funding, but then changed their mind and rejected Carnegie’s money.

Why would they do that? There are a variety of possible reasons, but a big one is distaste for Carnegie himself after he hired a militia to violently put down a mining strike (many people died). Berkes and Nencka show that these rejecting cities were, on average, no different than the acceptors on various observable metrics such as education levels, racial makeup, job mix, age, population. For Berkes and Nencka’s comparison to work, we have to believe that whatever the reason these cities rejected funding, it’s largely uncorrelated with a tendency to change patenting behavior after an application is accepted. Berkes and Nencka do a couple other things to establish this pretty convincingly in my view.

Because it turns out patenting does change pretty noticeably for the towns that get libraries, compared to the ones that apply and do not. The main story is visible in the figure below, which shows the average number of patents per city-year on a log scale: towns that ended up with libraries looked pretty similar to towns without ones up until they both applied for libraries, but then afterwards the towns that got libraries tended to have 8-12% more patents than the ones that didn’t. In the figure below, the solid line is when funds for a library are granted, and the dashed line is when libraries were typically opened (three years later).

(Aside: why does the figure above have this inverted U-shape? That has to do with larger trends of patent activity shifting away from towns and into big cities during the period under study - having a library did not stop that trend)

Moreover, Berkes and Nencka also provide some supportive evidence that the increase in patenting really does come from library access: patents from cities with libraries are more likely to contain words associated with citing a book (e.g., "vol.", "his book", "pp.", "pages", etc).

Let’s fast forward to 1975. In that year, the US Patent and Trademark Office began a program to dramatically increase the number of patent depository libraries around the country, with a goal of having at least one library in every state. Patents provide inventors with the right to exclude others from using their invention for a period of time, in exchange for disclosing how the invention works. In theory, that should let other people build on the underlying principles expressed in new inventions. But prior to the internet, to easily read published patent documents, you needed to go to one of these libraries and access to them was very unequally distributed.

Furman, Watzinger, and Nagler (2018) measure the impact of getting a patent library by following a similar strategy as Berkes and Nencka (2020): they compare regions that got a patent depository library to otherwise similar regions that did not. In their case, they use the fact that federal depository libraries serve as a potentially obvious control group. Federal depository libraries are libraries that provide access to federal regulation and laws, and they were also the most common library sites where patent depositories were set up. They compare patent rates in the 15 miles around federal depositories that get a patent depository library to the patent rates in the 15 miles around federal depository libraries that did not, but which are close to ones that did. (Why did some federal depository libraries get patents and others didn’t? The patent office basically followed a principle of first come, first served, so reasons were often idiosyncratic, like the person running a federal depository library wanted to travel to DC for annual patent training).

As with Berkes and Nencka, this paper finds the patent rate of regions that get a patent library diverged from those that didn’t in the subsequent years, as illustrated in the figure below. All told, getting a library seems to boost patent rates by about 17%.

Furman, Nagler, and Watzinger are also able to provide a lot of supporting evidence that this increase in patenting is driven by improved access to patents:

The effect is strongest for young firms and small firms, which we might assume are less likely to have alternative ways of accessing patents

The effect is strongest for technologies that disclose the most information in patents (chemistry patents)

Patents from near patent libraries cite more geographically distant patents and a wider variety of technology classes. That suggests the inventors are learning about what’s relevant from the library, rather than their social network (which is more likely to be local and to work on similar technologies).

Of course, today, anyone can read any patent online. We shouldn’t really expect it to matter if you are close to a patent library anymore. And this actually provides confirmatory evidence that those patent depository libraries really mattered. The first internet searchable patent databases became available in 1995. And it turns out the positive impact of having a physical patent depository library disappeared in that year!

That also suggests improving online access to knowledge might be another way to boost innovation. Enter wikipedia. Anyone with access to the internet can now read a free encyclopedia that has 6.3mn articles; for comparison, the encyclopedia Britannica never had more than 100,000 articles, even in it's digital incarnations. Wikipedia also has extremely detailed scientific articles: Thompson and Hanley (2020) find wikipedia covered 93% of the topics in upper level undergraduate chemistry classes and nearly half of the topics in masters' level graduate school. Does access to wikipedia have a similar effect on innovation as access to libraries?

To test this, Thompson and Hanley perform an experiment. They commission 43 new chemistry articles, written by PhD students, and then they randomly post half to wikipedia. They want to see if access to these articles (as compared to the unposted ones) exerts an influence on science.

Only 0.01% of academic papers directly cite wikipedia (I guess it’s embarrassing), but Thompson and Hanley provide a lot of evidence that scientists read and are influenced by these wikipedia articles. First off, these are quite specialized chemistry topics - for their experiment, they focus on material from chemistry grad school that wasn't already on wikipedia. But even though the topic is quite niche, the readership is huge: 4,400 views per month, 2 million total views as of February 2017. Moreover, while people may be shy to cite wikipedia, they are not shy about citing the scholarly literature that is listed in the wikipedia reference list. Thompson and Hanley show references in articles they published to wikipedia got 91% more citations on average than reference in their control group of unpublished articles. People are reading these articles and citing the referenced literature, rather than wikipedia itself.

Finally, the bulk of Thompson and Hanley's paper uses a textual similarity metric to identify the influence of wikipedia. For each of their chemistry articles, they compute the similarity of the wiki text to the text of published academic work in Elsevier. Basically, they have a method of checking things like the extent to which both articles use the same unusual words. They compare the similarity of elsevier and wiki articles published 6 months before a wiki article is published to the similarity of elsevier and wiki articles published 6 months after the wiki article. The figure below shows how the distribution of similarity changed over that period. The blue line corresponds to similarity with wiki articles that were posted and the green to similarity with the control wiki articles that were not posted (until after the experiment ended). How you can interpret this is that after wiki articles get posted, you see more elsevier articles that have high similarity to wiki articles (the ones in the top 10% of similarity). In the control, you see the opposite; more elsevier articles dissimilar to the (unposted) wiki articles are published over time.

So, Thompson and Hanley provide lots of evidence that these wikipedia articles shape the direction of science. They get read a bunch; the things they cite get cited a bunch; and after they are published, you see more peer-reviewed articles using similar words and phrases as the wiki article. Another thing the paper does is generalize this approach to all of chemistry wikipedia. Looking at 27,000 chemistry articles published on wikipedia and 326,000 chemistry articles published on elsevier, they ask how does the distribution of similarity change for elsevier articles published before and after wikipedia articles? What they find is quite similar to their much smaller experiment comprised of 43 new wikipedia articles. There is an increase in the number of elsevier articles using similar language as a given wikipedia article, after it gets published, as compared to the elsevier articles published before the wikipedia article.

(As the author of a free newsletter about academic research that strives to be accessible to lay readers, I have chosen to believe I too exert an inescapable gravitational pull on the direction of research)

Public libraries, patent depository libraries (in the pre-internet days), and chemistry articles in wikipedia; in all three cases, new access seems to have had a measurable impact on innovation.

Subscribe at mattsclancy.substack.com

View Details

Innovation is hard because you have to step into the unknown and it’s never certain what you’ll find there. Most of the time, nothing useful. We use knowledge - models, regularities, analogies from similar cases and so on - to reduce that uncertainty; but it’s always there.

When an individual or organization learns new knowledge, with a bit of luck the knowledge opens up new possibilities: it provides a map (more or less high resolution, depending on the state of things) of the unknown territory. So people whose business is innovation and discovery spend a lot of time searching for new and useful knowledge by reading, attending conferences, and talking with people. But the universe of knowledge is vast. Is there any rhyme or reason to searching through it? What kind of knowledge is most likely to be useful?

This is a big literature, but today I want to look at three papers that use different metrics to suggest knowledge which is distinct but close to your existing knowledge tends to be most useful.

To start, let’s look at a really clean experiment by Lane, Ganguli, Gaule, Guinan and Lakhani. Lane and coauthors ran an experiment back in 2011 where they invited all the life sciences faculty and researchers affiliated with a large US medical school to a symposium on medical research. As part of the symposium, participants were randomly assigned to different rooms, where they could look at research posters and talk with other researchers. Attendees were also fitted with a “sociometric badge” (see image below) that the authors used to determine who talked with whom: basically, if two badges were within 1 meter of each other and facing each other for one minute, the paper codes the attendees as having a face-to-face interaction.

Taken from the appendix of Kim et al. (2012)

Lane and coauthors also measure how “similar” each attendee’s knowledge is to their conversation partners. The primary measure they focus on uses the keywords attached to articles they’ve published. All these researchers are in the life sciences, which has a standardized vocabulary of descriptive keywords called the MeSH lexicon. They rate two people as having a low overlap if their published work has only two or fewer MeSH keywords in common, medium overlap if they have 3-11 words in common, and high overlap if they have 12 or more in common. They then keep track of these attendees for the next six years to see the long-run fallout from these accidental encounters with new knowledge (or, more precisely, with people whose brains have new knowledge).

So; we randomize people into different rooms; we see who they talk to; and we have very rough proxies for the kinds of things everyone knows. The first thing the authors show is that coauthorships are most likely to be born out of these conversations when researchers have an intermediate level of knowledge similarity. Perhaps surprisingly, people are slightly less likely to collaborate if they have a very high level of overlap and they meet each other at this conference.

Of course, there are other ways knowledge can be useful besides collaboration. Lane and coauthors try to get at that in two ways. First, they look to see if people are more likely to cite each other’s work when they meet. Again - it’s that intermediate level of knowledge overlap that most benefits from the face-to-face encounter:

What about knowledge that you can’t directly trace back to a citation? To get at that, Lane and coauthors look at those MeSH keywords again. In the following six years, are you more likely to start working on topics that match the keywords of your conversation partner? It turns out that again the answer is yes - if your conversation partner has that intermediate distance in terms of knowledge similarity.

Pretty impressive that you can see anything at all from a 90 minute opportunity to mingle! That said, since all these researchers were affiliated with the same institution, it’s quite likely that they had the opportunity to expand on these initial encounters if they found them fruitful.

Now let’s look at a very different and non-experimental context: what kind of ideas are found useful by new agricultural technologies?

Clancy, Heisey, Ji, and Moschini (2020) (that’s me, buyer beware) identified slightly more than 50,000 US patents, granted between 1976-2016, for agricultural technologies (think veterinary medicine, GMO plants, fertilizer, pesticides, tractors, etc). We then tried to identify the sources of the ideas that the patented technologies built on. There are no perfect measures for this, so we went about it in a few different ways. In each case, we found it to be pretty common that the majority of “knowledge” originated outside of agriculture… but not too far outside agriculture. This will be clearer with some examples.

For example, we looked at the citations agricultural patents make to academic journal articles. We then sorted the cited academic journals into different categories: agricultural science, biology, chemistry, and everything else. Most of the time, patents cite journals that don’t belong to the agricultural science category. But they still mostly cite chemistry and biology journals, even though the “everything else” category is much larger. Knowledge “close” to agriculture just tended to be more useful, or at least that’s our interpretation.

We also tried to identify the sources of knowledge that were borrowed from other technologies, instead of academia. For example, we looked at the citations agricultural patents make to other patents. We also looked at the actual text of these agricultural patents. Specifically, for each agricultural sub-field we identified at least 100 phrases (1-3 words) that corresponded to new technological concepts. These were phrases which were absent from the agricultural patent record before 1996, but relatively common thereafter (we think of them as new and important concepts). One example is the word “pyrimethamine”, which became common in veterinary medicine patents after 1996 but was absent beforehand. We then looked to see if these phrases popped up in other non-agricultural patents before 1996. Most of the time, they did. That means agriculture wasn’t the first patent to use these phrases. For example, pyrimethamine was pretty common in patents for human medicine before it began to appear in veterinary medicine patents after 1996.

So what kind of patents produce knowledge that’s useful to agriculture? Most of the time it wasn’t other agricultural patents (with one important exception - patents for plants tend to heavily cite patents for plants and other agricultural research patents). Most of the time, non-agricultural patents were the ones we linked to agricultural patents, via citation or shared phrases. But still, even though this indicates agricultural patents are borrowing knowledge from non-agricultural patents, most of these non-agricultural patents belonged to firms that also already had agricultural patents. This is despite the fact that these firms are in the minority. That is, most of the borrowed ideas came from firms who already had some pre-existing connection to agriculture, even though it was not their main area of patenting.

Let’s look at one more example. Cornelius, Gokpinar, and Sting (2020) study a completely different context: temporary assignments of automobile workers to different factories. One thing that’s really nice about this paper is it has a completely distinct measure of innovation. For a large unnamed European auto manufacturer, employees are encouraged to submit ideas for improving the efficiency and quality of auto production. These ideas are submitted to a database where they are evaluated by accountants to see if they will save the company money if implemented. So this is the dataset Cornelius, Gokpinar, and Sting are working with - employee ideas with dollar values in savings attached! (It’s super rare to find a dataset that provides such a good measure of the “value” of an idea, even though ideas vary enormously in their value)

They then look to see what happens to the value of ideas submitted by employees after they visit another plant. Plants differ (more on this in a minute), so visiting another plant is a way to learn about different approaches for assembling auto parts. After a visit, the average estimated savings of an employee’s ideas increase by about $25,000 per idea basically forever (and several times as high in the short term). The notion here is that being exposed to new ways of doing things allows workers to think of new and better improvements.

The challenge a paper like this has to confront is that employees aren’t just randomly assigned to visit another plant. The kinds of employees that are sent out for a visit tend to be very good employees and they already generate more than the average number and value of ideas. So if we just compare the value of ideas for employees who have gone on trips to those that haven’t, we’ll get a biased result, since the people making visits probably would have had more valuable ideas whether they went on a trip or not.

The authors try to account for this in two ways: first, they look only at how site visits change the value of an individual’s ideas over time (comparing the same person’s ideas before and after a visit), and also taking into account typical trends in how idea values evolve over time (more experienced employees usually have better ideas). In econ jargon, they have individual fixed effects and time-varying measures of idea quality. Second, they compare employees who go on site visits to those who come from the same plant and are similar in terms of their ability to generate valuable ideas, but where one is assigned to visit and the other is not. Both methods deliver the above result: visits increase the value of subsequent ideas.

But not all site visits are alike. Some plants have significantly more overlap in terms of the products produced and the machinery used than others. The authors find it’s visits to these similar plants that generate by far the most value.

Bottom line: learning something new by visiting a new plant increases the value of subsequent ideas for an employee; but that impact falls off if the plant is too different from the one the employee normally works at.

So, in these three cases - producing papers in the life sciences, descriptive work on the sources of ideas in agricultural technology patents, and submitted ideas for improving the efficiency of automobile manufacturing - it looks like inventors find most useful knowledge that is not exactly where they already are, but adjacent.

This isn’t the last word though. This is a big literature and we’ll be back here sometime again.

Subscribe at mattsclancy.substack.com

View Details

Innovation disproportionately happens in cities. What is it about packing people together that makes them so innovative?

Last week we looked at a few papers that showed denser neighborhoods with lots of restaurants, cafes, and bars facilitated more innovation, and that the kind of innovation that happens in cities tends to reflect the social milieu of the surroundings. In particular, the residents of dense parts of cities tend to work on a more diverse collection of technologies, and that this is reflected in the kinds of patents they create.

This week I want to look at some evidence that one of the most important functions of cities is to introduce us to new people. I’ll go on to argue being close seems to be very important for initiating and consolidating new relationships, but that once they’re formed it’s no longer so important that you stay physically close - at least from the perspective of facilitating innovation.

Consider a 2018 paper by Christian Catalini. Catalini exploits a natural experiment related to a French university’s 17-year quest to rid its campus of asbestos. Asbestos removal is a disruptive process and whenever a university lab’s turn for renovation came up, it required relocating them to whatever space was available, with little ability for the lab to petition about its location. This meant that labs across campus suddenly found themselves with new neighbors, or separated from old neighbors.

Catalini finds that labs are more likely to collaborate after they are moved into the same building. In the diagram below, year 0 represents the time when two previously separated labs come to reside in the same building.

However, when labs that used to be in the same building are separated, it doesn’t seem to have any impact on their probability of collaborating.

That suggests being neighbors is important for meeting new people, but close proximity isn’t really important after you know about each other. Another line of evidence from Catalini comes from the likelihood two labs would know about each other’s work regardless of their proximity. Catalini uses the publications the labs put out to measure the degree to which different labs work on different topics. It turns out being located in the same building has a big impact on the probability two labs collaborate, when they normally work on different things. For labs that work on similar scientific topics, it doesn’t seem to matter much if they are near or far. In this case, it seems likely the labs don’t need to be close to meet; they were probably always going to meet since they go to the same conferences, seminar series, and so on.

But in Catalini’s study the distances between labs are never that great. The worst case scenario is they get moved across campus; it’s still probably pretty easy to collaborate at those distances. But other work finds similar results for much larger moves.

Agrawal, Cockburn and McHale (2006) look at the citations patents receive when an inventor moves. For example, suppose Ada is an inventor in Ames, Iowa who moves to New York City. Once in New York Ada comes up with a new patented invention. Agrawal, Cockburn, and McHale show inventors in Ames are more likely to cite Ada’s new patent then are inventors of technologically similar patents from cities equally far from New York. It’s like the Ames inventors have kept in touch with Ada and know what she’s working on. Indeed, 80% of the increased citations Ada receives from Ames can be attributed to people who either (1) worked on a patent with Ada back when she was in Ames or (2) worked at the same organization as Ada back when she was in Ames. These are precisely the kind of people we would expect to know Ada personally.

Agrawal, Cockburn, and McHale also show this effect is stronger for citations across different technologies. For example, suppose Ada works on lithium battery technology. When she moves to New York, she receives about 135% as many citations from other lithium battery inventors who are located in Ames, than she receives from lithium battery inventors who are located from some other equally distant city (for example, Birmingham Alabama). But she receives nearly 175% as many citations from inventors who don’t work on batteries but reside in Ames, than she does from inventors who don’t work on batteries but reside in somewhere like Birmingham. Again; this is consistent with the idea that being located together in the same city was especially important for Ada to meet people who she wouldn’t normally meet - in this case, people who do not work on the same kind of thing as her.

We can go farther. Miguelez and Noumedem Temgoua (2020) look at citations between patents in different countries, when inventors migrate. They find patents from country A are more likely to cite patents from country B when more inventors have migrated from A to B (this is possible thanks to a nice new dataset that actually tracks migrant inventors). As with Agrawal, Cockburn, and McHale, this effect is actually stronger for countries that are otherwise technologically less similar. This is again consistent with the notion that proximity - here, merely residing in the same country - is especially helpful for forming social ties with people who work on different technologies than is typical for the country.

Digging into Academia

But we should always be cautious leaning too heavily on patent citation data. They are a pretty imperfect measure. Another major source of “paper trails” for knowledge are academic papers.

Head, Li, and Minondo (2019) look at citations between mathematics papers as evidence of how knowledge moves through a community of researchers. Specifically, they want to know what kinds of things predict whether paper x cites paper y. For our purposes, the variable of interest is their measure of social ties between mathematicians. They measure this in a lot of different ways: advisor-advisee relationships, whether two mathematicians worked in the same place at the same time, or went to the same graduate school around the same time, etc. Note - by definition - most of these relationships are defined in terms of physical proximity at some point in time. Head, Li, and Minondo find a couple of things that are consistent with what we’ve talked about so far.

First, if you do not include any data on social ties, mathematicians are less likely to cite each others’ work if they live far away from each other. But, when you do include social ties, the strength of this relationship gets cut in half. And, looking only at data from the early 2000s onward, the impact of distance disappears completely once you account for social ties. In layman’s terms, what’s going on is something like this: mathematicians are more likely to cite mathematicians they know, and more likely to know mathematicians who live nearby. But, since the year 2000, the only thing that matters (in the data) is the existence of a social tie. If two mathematicians work in the same department and then one moves away this doesn’t really impact the likelihood that they cite each other’s work. It’s the same kind of finding as we had for patents.

Moreover, as with the patents, the importance of social ties is stronger for mathematicians who work in different fields. Again - proximity helps forge relationships, especially relationships that would not normally form in the course of keeping up to date on the field. And those relationships remain pretty durable to moves.

Freeman, Ganguli, and Murciano-Goroff (2015) have some descriptive data on distance and academic collaboration that’s also consistent with this. Looking at 126,000 papers in the fields of particle and field physics, nanoscience and nanotechnology, and biotechnology and applied microbiology, they don’t find any consistent evidence about the impact of having geographically distant coauthors. When all the authors are based in the USA, the citations received by papers authored by geographically distant coauthors are no different than those received by geographically proximate ones in two of the three fields (international collaboration did tend to reduce citations in all cases). But even if it’s possible to productively collaborate at a distance, a strong majority of coauthors first met while they were geographically close (either as colleagues or advisors and advisees).

Tying it together

So, across a lot of contexts we find evidence consistent with this story: innovators meet other innovators who live nearby, whether they work in the same field or not. Once a relationship is formed, it remains pretty productive even after you get subsequently separated, in the sense that you can still collaborate well or at least learn from each other.

Next week’s post will be about what kind of knowledge tends to be useful for an innovator to know: near, far, or something in between?

If you liked this post, you might also like Cities as a platform to mix up knowledge.

Subscribe at mattsclancy.substack.com

View Details

Innovation disproportionately happens in cities. Carlino, Chatterjee, and Hunt (2007) find, all else equal, that doubling the number of jobs per square mile is associated with 20% more patents per capita. Why? What is it about packing people together that makes them so innovative? I’ll argue in the next few weeks that a big part of the answer is that it’s about meeting new people.

This week, let’s have a look at three suggestive recent papers. Berkes and Gaetani (2020) have data on population density and patenting at the level of the county sub-division (CSD), which is a smaller geographic area than an entire city. How do really dense parts of the country differ from less dense parts? For one, they tend to house inventors working on lots of different kinds of technology.

Berkes and Gaetani come up with a measure of the technological diversity of patents in a particular CSD using the technology classifications assigned to patents by the USPTO. When their measure is low, it means that if you pick two patents at random from a given CSD, the technology classes the patents belong to are not usually found together in the same place. As you can see below, the more densely populated a CSD (horizontal axis), the more unusual combinations of technologies you get if you pick two patents at random from that CSD.

So, we can imagine that means there are a lot of people working on different technologies, all milling around in a tightly packed corner of some city. Maybe they bump into each other, talk about their work, and sometimes this sparks a new idea. The inventions and patents that emerge from those spontaneous encounters are more likely to reflect unusual combinations of knowledge. And in fact, patents in more dense parts of the country tend to cite more unusual combinations of technology.

As in the first figure, this is a measure of how unlikely it is that the technology classes of two cited references in a patent are found together. The lower the number, the weirder it is for the same patent to cite these two classes. So, if you live in a place where lots of people are packed into a small area, you likely have a lot of neighbors working on different technologies, and that gets reflected in the kinds of technologies that end up being built there - they draw on diverse sources of knowledge.

Berkes and Gaetani’s paper has a lot more interesting findings, but let’s keep our eyes on evidence about meeting people and move on.

Next, let’s look at a 2020 paper by Maria Roche. Roche is not focused so much on the people who reside in a particular corner of a city, but rather the structure of the city itself. Specifically, she has a measure of the density of walkable streets at different US census blocks. The notion here is that a higher density of streets (importantly, with sidewalks) facilitates more serendipitous meetings. And, indeed, after controlling for many possible confounding variables she finds places with a higher density of streets have more patents, though the effects are tiny (10% denser streets leads to 0.04% more patents). She also finds the denser the streets, the more likely patents are to cite other patents from the same census block. Again, it’s consistent with people meeting each other at a higher rate, exchanging ideas, and then inventing based on what they learn.

Are people really just running into each other on the street, stopping to chat, and then coming up with an invention? Probably not. The streets are a proxy for the existence of other structures that facilitate socializing and mixing with other people. One such place is a restaurant or bar. In fact, Roche finds the impact of density is significantly stronger for places that have 5 or more restaurants, cafes and bars. Increasing the density of walkable streets by 10% in a census block with at least 5 restaurants, cafes or bars is correlated with 0.4% more patents (10x as strong!).

Another paper looks specifically at the impact of bars on innovation.

Why should bars affect innovation? It’s not so much that inventive work happens in bars. But bars are a key social institution in America. When you are just meeting someone for the first time and wish to know them better, it’s easier to ask them to grab a drink at the bar then to come over for dinner. And unlike workplaces and homes, bars are a “third place” where people commonly come into contact with people they wouldn’t normally see. What would happen to all those social interactions if you shut bars down?

Well, America did just that in the early 20th century, with prohibition. Andrews (2019) looks at the change in patenting in counties affected by prohibition (i.e., counties that had previously been wet ones) as compared to counties unaffected by prohibition (i.e., because they were always dry counties). You can see a big drop in patenting per capita in the wet counties that is not matched in the dry ones right around year 0, which is how Andrews labels the onset of prohibition.

Notice also that patenting rebounds after 3 years. People remain social creatures, and though it took a few years, maybe they eventually built up replacements for the social functions previously served by bars: speakeasies, parties, and church picnics?

When I first read the abstract of this paper, frankly I didn’t believe it. It seemed too cute, and the effect too large. But the paper slowly convinced me it was on to something by drawing together a battery of different lines of evidence. For example:

If this is really about access to bars, then people who don’t frequent bars should be unaffected by the policy change. In the early 20th century, it was uncommon for women to frequent bars and Andrews finds no impact of prohibition on patenting by women. Similarly for ethnic groups that were not as associated with bar culture at the time.

Collaborations between people who are less likely to meet in a bar should be less impacted by the policy. Andrews shows collaborations between inventors who reside in different cities were not impacted by the policy change.

If new social structures had to be built to replace the functions served by bars, it’s unlikely these replacements would perfectly replicate the network of contacts that bars supported. Compared to persistently dry counties, Andrews shows the wet counties showed signs of changes to the social structure undergirding innovation. In formerly wet counties there were fewer patents by people who had worked together prior to prohibition, and the technology class of patents exhibited noticeable shifts, as compared to counties that were unaffected by prohibition.

Taken together, it adds up to a lot of suggestive evidence that interactions with the local social milieu is really important for the amount and type of innovation that happens.

We’ll have another newsletter next week (February 8), which will look at whether the physical proximity of cities is important because it lets you meet new people, or because it makes it easier to work with people you already know.

If you liked this post you might also like this set of three thematically linked posts:

Local knowledge spillovers are weakening, because of cheaper travel and the internet.

Subscribe at mattsclancy.substack.com

View Details

Getting an academic field to change its ways is hard. Independent scientists can’t just go rogue - they need research funding, they need to publish, and they need (or want) tenure. All those things require convincing a community of your peers to accept the merit of your research agenda. And yet, humans are humans. They can be biased. And if they are biased towards the status quo, it can be hard for change to take root.

But it does happen. And I think changes in the field of economics are a good illustration of some of the dynamics that make that possible.

The Applied Turn in Economics

In 1983, economist Edward Leamer, writing about economics, said:

Hardly anyone takes data analysis seriously. Or perhaps more accurately, hardly anyone takes anyone else’s data analysis seriously.

But in subsequent decades, economics made a famous “applied turn.” The share of empirical papers in three top journals rose from a low of 38% in the year of Leamer’s paper to 72% by 2011! The field also began giving its top awards - the economics Nobel and the John Bates Clark award for best economist under 40 - to empirical economists. These new empirical papers were strongly characterized by careful attention to credibly detecting causation (rather than mere correlation).

This applied turn seems to have worked out pretty well. In 2010, Joshua Angrist and Jorn-Steffen Pischke wrote a popular follow-up to Leamer’s article titled The Credibility Revolution in Empirical Economics: How Better Research Design is Taking the Con out of Econometrics. The title is an apt summary of the argument. More quantitatively, a 2020 paper by Angrist et al. tried to assess how this empirical work is received in and outside of economics, by looking at citations to different kinds of articles. Consistent with Leamer’s complaints, when looking at total citations received by papers, Agrist et al. find empirical papers faced a significant citation penalty prior to the mid 1990s.

When they restrict their attention to citations of economics papers by non-econ social science papers, they find empirical papers now benefit from a citation premium, especially since the 2000s. This suggests the applied turn has been viewed favorably, even by academics who are outside of economics.

Elsewhere, Alvaro de Menard, a participant in an experiment to see if betting markets could predict whether papers will replicate, produced this figure summarizing the market’s views on the likelihood of replication for different fields (higher is more likely to replicate).

Among the social sciences, economics is perceived to be most likely to replicate (though note the mean likelihood of replication is still below 70%!). I should also add at this point, that if you are a skeptic about the validity of modern economics research, I think you can still get a lot out of the rest of this post, since it’s about how a field comes to embrace new methods, not so much whether those new methods “really” are better.

So. How did this remarkable turn-around happen?

Changing a Field is Hard

Changing a field is hard. Advancing in academia is all about convincing your peers that you do good work. The people you need to convince are humans and humans are biased. In particular, they may be more or less biased towards the methods and frameworks that are already prevalent in the field.

Bias might reflect the cynical attitudes of researchers who have given up on advancing truth and knowledge, and just want to protect their turf. But bias can also emerge from other motives. Maybe biased researchers simply haven’t been trained in the new method and don’t appreciate its value. Or maybe they choose to value different aspects of methods, like interpretability over rigor. Or maybe there really is disagreement over the value of a new method - even today there is lots of debate about the proper role of experimental methods in economics. Lastly, it may just be that people subconsciously evaluate their own arguments more favorably than those advanced by other people.

Akerlof and Michaillat have a nice little 2018 paper on these dynamics that shows how even a bit of bias can keep a field stuck in a bad paradigm. Suppose there are two paradigms, an old one and a new (and better) one. How does this new paradigm take root and spread? Assume scientists are trained in one paradigm or the other and then have to convince their peers that their work is good enough to get tenure. If they get tenure, they go on to train the next generation of scientists in their paradigm.

Research is risky and whatever paradigm they are from, it’s uncertain whether they’ll get tenure. The probability they get tenure depends on the paradigm of the person evaluating them (for simplicity, let’s just assume every untenured scientist is evaluated by just one evaluator).

The probability a candidate gets tenure is...

In this example, the new paradigm is better, in the sense that an unbiased evaluator would give one of its adherents tenure more often than someone trained in the old paradigm (50% of the time, versus 30%). But in this model, people are biased. Not 100% biased, in the sense that they will only accept work done in their own paradigm. Instead, they are 30% biased: they give a 30% penalty to anyone from the other paradigm and a 30% bonus to anyone from their own paradigm. This means the old paradigm evaluators still favor candidates from their own field, but not by much (39% versus 35%). On the other hand, the new paradigm people are also biased, and since their paradigm really is better, the two effects compound and they are much more likely to grant tenure to people in their own paradigm (65% vs 21%).

What’s the fate of the new paradigm? It depends.

Suppose the new paradigm has only been embraced by 5% of the old guard, and 5% of the untenured scientists. If scientists cannot choose who evaluates them, and instead get matched with a random older scientist to serve as an evaluator then in the first generation:

95% of the new paradigm scientists are evaluated by old paradigm scientists, and only 35% of them are granted tenure

5% of the new paradigm scientists are evaluated by new paradigm scientists and 65% of them get tenure

95% of the old paradigm scientists are evaluated by old paradigm-ers, and 39% of them get tenure

5% of the old paradigm scientists are evaluated by new paradigm-ers and only 21% of these get tenure.

Thus, of the original 5% of untenured scientists in the new paradigm, only 1.8% get tenure (5% x (0.95 x 0.35 + 0.05 x 0.65)) and of the original 95% of untenured scientists in the old paradigm, 36.2% get tenure (95% x (0.95 x 0.39 + 0.05 x 0.21)). All told, 38% of the untenured scientists get tenure. They become the new “old guard” who will train and then evaluate the next generation of scientists.

In this second generation of tenured scientists, 4.8% belong to the new paradigm and 95.2% belong to the old paradigm. That is, the share of scientists in the new paradigm has shrunk. If you repeat this exercise starting with 4.8% of scientists in the new paradigm, the new paradigm fares even worse, because they are less likely to receive favorable evaluators than the previous generation. Their share shrinks to 4.6% in the third generation; then 4.4% in the fourth; and so on, down to less than 0.1% of the population by the 25th generation.

Escaping the Trap

In contrast, if the share of scientists adopting the new paradigm had been a bit bigger - 10% instead of 5%, for example - then in every generation the share of new paradigm scientists would get a bit bigger. That would make it progressively easier to get tenure, if you are a new paradigm scientist since it gets increasingly likely you’ll be evaluated by someone favorable to your methods.

Akerlof and Michaillat show that if a new paradigm is going to take root, it needs to get itself above a certain threshold. In the example I just gave, the threshold is 8.3%. In general, this threshold is determined by just two factors: the superiority of the new paradigm and the extent of bias. The better the paradigm and less bias, the lower is this threshold and therefore the lower is the population of scientists that needs to be “converted” to the new paradigm for it to take root.

The applied turn in economics looks reasonably good on these two factors.

Let’s talk about the new paradigm’s superiority first. Here, we are talking about the extent to which the new paradigm really is better than the old one, as judged by an unbiased evaluator. In general, the harder it is to deny the new paradigm is better, the better the outlook of the new paradigm. In my example above, if the new paradigm earns tenure from an unbiased evaluator 55% of the time instead of 50%, this is enough to outweigh the 30% bias penalty. The new paradigm will grow in each generation.

One way to establish a paradigm’s superiority is if it is able to answer questions or anomalies that the prior paradigm failed to answer (this is closely related to Kuhn’s classic model of paradigms in science). Arguably, this was the case with the new quasiexperimental and experimental methods in economics. As retold in Angrist and Pischke’s “The Credibility Revolution” a 1986 paper by Lalond was quite influential.

Lalond had data on an actual economic experiment: in the mid-1970s, a national job training program was randomly given to some qualified applicants and not to others (who were left to fend for themselves). Comparing outcomes in the treated and control groups indicated the program raised incomes by a bit under $900 in 1979. In contrast, using economics’ older methods produced highly variable results with some estimates off by more than $1,000 in either direction.

The second factor that matters in Akerlof and Michaillat is the extent of bias. If the bias in this example was 20% instead of 30%, the new paradigm would gain more converts in each generation, further raising the tenure prospects for new paradigm-ers.

I suspect certain elements of quasiexperimental methods also helped reduce the bias these methods faced when they were taking root in economics. It’s probably important here that economics has a long history of thinking of experiments as an ideal but regrettably unattainable method. When some people showed that these methods were indeed usable, even in an economic context, it was easier to accept them than if the field had adopted an attitude of not valuing experiments even if they were valuable. Moreover, some quasi-experimental ideas could easily be understood in terms of long-standing empirical approaches in economics (like differentiating between exogenous and endogenous variables).

So, given all that, the applied turn in economics probably had a relatively low bar to clear. But is that the whole story? I don’t think so, for two reasons.

First, this version of the story is really flattering to economists. Basically, it says the applied turn happened because economists were not that biased against this kind of paradigm shift, and because good work convinced them of it’s value. But we should be skeptical towards explanations that are flattering to our self-image.

Second, even supposing the threshold economics needed to clear was a small one, supposing there was still a bar to clear at all, how did economics get above it? In any new paradigm, there will tend to be a very small number of people who initially adopt it. How does this microscopic group get above the threshold?

One possibility is simple luck. This is a bit obscured in my presentation; the dynamics I describe are only the averages. The new paradigm lot could get lucky and be assigned to a greater share of new paradigm tenured evaluators than the average, or they might catch a few breaks from old paradigm evaluators. If a bit of luck pushes them over a crucial threshold (in this case, 8.3%), then they can expect to keep growing in each generation. Luck is especially important when fields are small.

In economics, the important role of a single economist, Orley Ashenfelter, is often highlighted. Panhas and Singleton’s wrote a short history of the rise of experimental methods in economics, noting:

...Ashenfelter eventually returned to a faculty position at Princeton. In addition to his promoting the use of quasi-experimental methods in economics through both his research and his role as editor of the American Economic Review beginning in 1985, he also advised some of the most influential practitioners who followed in this tradition, including 1995 John Bates Clark Medal winner David Card and, another of the foremost promoters of quasi-experimental methods, Joshua Angrist.

We will return to Ashenfelter in a bit. But if he had not been in this position at this time (editor of the flagship journal in economics and an enthusiastic proponent of the new methods), it’s possible that things might have turned out differently for the applied turn.

All of these factors are probably part of the story of the applied turn in economics. But a 2020 preprint by O’Conner and Smaldino suggests a fourth factor that I think turned out to also be very important.

O’Connor and Smaldino suggest interdisciplinarity can also be an avenue for new paradigms to take root in a field. The intuition is simple: start with a model like Akerlof and Michaillat’s, but assume there is also some probability that you will be evaluated by someone from another field. This doesn’t have to be a tenure decision - it could also represent peer review from outside your discipline that helps build a portfolio of published research.

If that other field has embraced the new paradigm, then this introduces a new penalty for those using the old paradigm, since they will get dinged anytime they are reviewed by an outsider, and it raises the payoff to adopting the new paradigm since they benefit anytime an outsider reviews them. If we add a 10% probability that you are evaluated by an evaluator from another field that has completely embraced the new paradigm, this is also enough to ensure the share of new paradigm scientists grows in each period.

Outside Reviewers in Economics

At first, this wouldn’t seem to be relevant. Economics has a reputation for being highly insular, with publication in one of five top economics journals a prerequisite for tenure in many departments. When Leamer was writing in 1983, well under 5% of citations made in economics articles were made to other social science journals (which compares unfavorably to other social sciences). Even if there was another field using the same quasi-experimental methods as would later become popular in economics, they weren’t going to be serving as reviewers for many economics articles.

But economics is unusual in having a relatively large non-academic audience for its work: the policy world. Economists testify before congress at twice the rate of all other social scientists combined and this was especially true in the period Leamer was writing. I think policy-makers of various stripes played a role analogous to the one O’Connor and Smaldino argue can be played by other academic disciplines.

The quasi-experimental methods that came to prominence in the applied turn began in applied microeconomics, more specifically in the world of policy evaluation. Policy-makers had a need to evaluate different policies, and experimental methods were viewed as more credible by this group than existing alternatives (such as assuming a model and then estimating its parameters). In 2017, Panhas and Singleton wrote a short history of the rise of experimental methods in economics, highlighting these roots in the world of policy evaluation:

The late 1960s and early 1970s saw field experiments gain a new role in social policy, with a series of income maintenance (or negative income tax) experiments conducted by the US federal government. The New Jersey experiment from 1968 to 1972 “was the first large-scale attempt to test a policy initiative by randomly assigning individuals to alternative programs” (Munnell 1986). Another touchstone is the RAND Health Insurance Experiment, started in 1971, which lasted for fifteen years.

This takes us back to Orley Ashenfelter. Upon graduating with a PhD from Princeton in 1970, Ashenfelter was appointed director of the Office of Evaluation at the US Department of Labor. As Panhas and Singleton write:

[Ashenfelter] recalled about his time on the project using difference-in-differences to evaluate job training that “a key reason why this procedure was so attractive to a bureaucrat in Washington, D.C., was that it was a transparent method that did not require elaborate explanation and was therefore an extremely credible way to report the results of what, in fact, was a complicated and difficult study” (Ashenfelter 2014, 576). He continued: “It was meant, in short, not to be a method, but instead a way to display the results of a complex data analysis in a transparent and credible fashion.” Thus, as policymakers demanded evaluations of government programs, the quasi-experimental toolkit became an appealing (and low cost) way to provide simple yet satisfying answers to pressing questions.

This wasn’t the only place that the economics profession turned to quasi-experimental methods to satisfy skeptical policy-makers. In another overview of the applied turn in economics, Backhouse and Cherrier write:

...in 1981 Reagan made plans to slash the Social Sciences NSF budget by 75 percent, forcing economists to spell out the social and policy benefits of their work more clearly. Lobbying was intense and difficult. Kyu Sang Lee (2016) relates how the market organization working group, led by Stanley Reiter, singled out a recent experiment involving the Walker mechanism for allocation of a public good as the most promising example of policy-relevant economic research.

So it seems at least part of the success of the applied turn in economics was the existence of a group outside of the academic field of economics who favored work in the new “paradigm” and that this allowed the methods to get a critical toehold in the academic realm.

Through the 1990s and 2000s, the share of articles using quasi-experimental terms rose, led by applied microeconomics. But by the 2010s, the experimental method also began to rise rapidly in another field: economic development. This too was a story about economics’ contact with the policy world, albeit a different set of policy-makers.

Economic Development and the Rise of the RCT

de Souza Leão and Eyal (2019) also argue the rise of a key new method in development economics - the randomized control trial (RCT) - was not the inevitable result of the method’s inherent superiority. They point out that the current enthusiasm for RCTs in international development is actually the second such wave of enthusiasm, after an earlier wave dissipated in the early 1980s.

The first wave of RCTs in international development occurred in the 1960s-1980s, and was largely led by large government funders engaging in multi-year evaluations of big projects or agencies. In this first wave, experimental work was not in fashion in academic economics, and experiments were instead run primarily by public health, medicine, and other social scientists, as well as non-academics. Enthusiasm for this approach waned for a variety of reasons. The length and scale of the projects heightened concerns about the ethics of giving some groups access to a potentially beneficial program and not others. Moreover, this criticism was particularly acute when directed at government funders, who are necessarily responsive to political considerations, and who it could be argued have a special duty to provide universal access to potentially beneficial programs. The upshot is a lot of experiments were not run as initially planned (after political interference), and then a lot of time and money had been spent on evaluations that weren’t that informative.

The second wave, on the other hand, shows no sign of slowing down. The second wave of RCTs in international development began after the international development community fragmented after the breakdown of the Washington consensus. International NGOs and philanthropists were a new source of funding for experiments. Compared to governments, international NGOs and philanthropists were more insulated from the arguments about the ethics of offering an intervention to only part of a population. They did not claim to represent the general will of the population and budget constraints usually meant that universal access was infeasible anyway. Moreover, the nature of these interventions tended to be shorter and smaller, which also tended to blunt the argument that experiments were unethical. (Though critiques of the ethics of RCTs in economic remain quite common)

On the economist’s side, appetite for using RCT methods had become somewhat established by this time, thanks to its start in applied microeconomics. Moreover, economists were content to study small or even micro interventions, because of a belief that economic theory provided a framework that would allow a large number of small experiments to add up to more than the sum of their parts. Whereas the first wave of RCTs was conducted by a wide variety of different academic disciplines, this second wave is dominated by economists. This also creates a critical mass in the field, where economists using RCTs can be confident their work will find a sympathetic audience from their peers.

How to Change a Field

So, to sum up, it’s hard to change a field if people are biased against that change since any change necessarily has a small number of adherents at the outset. One way change can happen though, is if an outside group sympathetic to the innovation provides space for its adherents to grow. Above a threshold, the change can be self-perpetuating. In economics, the rise of quasi-experimental methods in the field probably arises partially from the presence of the policy world, which liked these methods and allowed them to take root. It was also important, however, that these methods credibly established their utility in a few key circumstances, and also that they could be framed in such a way that was consistent with earlier work.

The next newsletter (February 2) will be about one reason cities are so innovative: because they create encounters between people who would not normally meet.

If you enjoyed this post, you might also enjoy:

Does science self-correct? (evidence from retractions)

Does chasing citations lead to bad science? (sometimes yes, in general no)

How bad is publish-or-perish for the quality of science? (it’s not great)

Subscribe at mattsclancy.substack.com

View Details

In economic models of “learning-by-doing,” technological progress is an incidental outcome of production: the more a firm or worker does something, the better they get at it. In its stronger form, the idea is formalized as a “learning curve” (also sometimes called an experience curve or Wright’s Law), which asserts that every doubling of total experience leads to a consistent decline in per-unit production costs. For example, every doubling of the cumulative number of solar panels installed is associated with a 20% decline in their cost, as illustrated in the striking figure below (note both axes are in logs).

Learning curves have an important implication: if we want to lower the cost of a new technology, we should increase production. The implications for combating climate change are particularly important: learning curves imply we can make renewable energy cheap by scaling up currently existing technologies.

But are learning curves true? The linear relationship between log costs and log experience seems to be compelling evidence in their favor - it is exactly what learning-by-doing predicts. And similar log-linear relationships are observed in dozens of industries, suggesting learning-by-doing is a universal characteristic of the innovation process.

But what if we’re wrong?

But let’s suppose, for the sake of argument, the idea is completely wrong and there is actually no relationship between cost reductions and cumulative experience. Instead, let’s assume there is simply a steady exponential decline in the unit costs of solar panels: 20% every two years. This decline is driven by some other factor that has nothing to do with cumulative experience. It could be R&D conducted by the firms; it could be advances in basic science; it could be knowledge spillovers from other industries, etc. Whatever it is, let’s assume it leads to a steady 20% cost reduction every two years, no matter how much experience the industry has.

Let’s assume it’s 1976 and this industry is producing 0.2 MW every two years, and that total cumulative experience is 0.4 MW. This industry faces a demand curve - the lower the price, the higher the demand. Specifically, let’s assume every 20% reduction in the price leads to a doubling of demand. Lastly, let’s assume cost reductions are proportionally passed through to prices.

How does this industry evolve over time?

In 1978, cost and prices drop 20%, as they do every every two years. The decline in price leads demand to double to 0.4 MW over the next two years. Cumulative experience has doubled from 0.4 to 0.8 MW.

In 1980, cost and prices drop 20% again. The decline in price leads demand to double to 0.8 MW per decade. Cumulative experience has doubled from 0.8 MW to 1.6 MW.

In 1982… you get the point. Every two years, costs decline 20% and cumulative experience doubles. If we were to graph the results, we end up with the following:

In this industry, every time cumulative output doubles, costs fall 20%. The result is the same kind of log-linear relationship between cumulative experience and cost as would be predicted by a learning curve.

But in this case, the causality is reversed - it is price reductions that lead to increases in demand and production, not the other way around. Importantly, that means the policy implications from learning curves do not hold. If we want to lower the costs of renewable energy, scaling up production of existing technologies will not work.

This point goes well beyond the specific example I just devised. In any demand curve with a constant elasticity of demand, it can be shown constant exponential progress yields the same log-linear relationship predicted by a learning curve. And even when demand doesn’t have a constant elasticity of demand, you can frequently get something that looks pretty close to a log-linear relationship, especially if there is a bit of noise.

But ok; progress is probably not completely unrelated to experience. What if progress is actually a mix of learning curve effects and constant annual progress? Nordhaus (2014) models this situation, and also throws in growth of demand over time (which we might expect if the population and income are both rising). He shows you’ll still get a constant log-linear relationship between cost and cumulative experience, but now the slope of the line in such a figure is a mix of all these different factors.

In principle, there is a way to solve this problem. If progress happens both due to cumulative experience and due to the passage of time, then you can just run a regression where you include both time and total experience as explanatory variables for cost. To the extent experience varies differently from time, you can separately identify the relative contribution of each effect in a regression model. Voila!

But the trouble is precisely that, in actual data, experience does not tend to differ from time. Most markets tend to grow at a steady exponential rate, and even if they don’t, their cumulative output does. This point is made pretty starkly by Lafond et al. (2018), who analyze real data on 51 different products and technologies. For each case, they use a subset of the data to estimate one of two forecasting models: one based on learning curves, one based on constant annual progress. They then use the estimated model to forecast cost for the rest of the data and compare the accuracy of the methods. In the majority of cases, they tend to perform extremely similarly.

To take one illustrative example, the figure below forecasts solar panel costs out to 2024. Beginning in 2015 or so, the dashed line is their forecast and confidence interval for a model assuming constant technological progress (which they call Moore’s law). The red lines are their forecasts and confidence intervals for a model assuming learning-by-doing (which they call Wright’s law). The two forecasts are nearly identical.

So the main point, so far, is that consistent declines in cost whenever cumulative output doubles is not particularly strong evidence for learning curves. Progress could be 100% due to learning, 100% due to other factors, or any mix of the two, and you will tend to get a result that looks the same.

But that doesn’t mean learning curves are not true - only that we need to look for different evidence.

A theoretical case for learning curves

One reason I think learning curves are sort-of true is that they just match our intuitions about technology. We have a sense that young technologies make rapid advances and mature ones do not. This is well “explained” by learning curves. By definition, firms do not have much experience with young technologies; therefore it is relatively easy to double your experience. Progress is rapid. For mature technologies, firms have extensive experience, and therefore achieving a doubling of total historical output takes a long time. Progress is slow.

There is a bit of an issue of survivor bias here. Young technologies that do not succeed in lowering their costs never become mature technologies. They just become forgotten. So when we look around at the technologies in widespread use today, they tend to be ones that successfully reduced cost until they were cheap enough to find a mass market, at which point progress plateaued. All along the way, production also increased since demand rises when prices fall. (Even here, it’s possible to think of counter-examples: we’ve been growing corn for hundreds of years, yet yields go up pretty consistently decade on decade)

But even acknowledging survivor bias, I think learning-by-doing remains intuitive. Young technologies have a lot of things that can be improved. If there’s a bit of experimentation that goes alongside production, then many of those experiments will yield improvements simply because we haven’t tried many of them before. Mature technologies, on the other hand, have been optimized for so long that experimentation is rarely going to find an improvement that’s been missed all this time.

There’s even a theory paper that formalizes this idea. Auerswald, Kauffman, Lobo, and Shell (2000) apply models drawn from biological evolution to technological progress. In their model, production processes are represented abstractly as a list of subprocesses. Every one of these sub-processes has a productivity attached to it, drawn from a random distribution. The productivity of the entire technology (i.e., how much output the technology generates per worker) is just the sum of the productivities of all the sub-processes. For instance, in their baseline model, a technology has 100 sub-processes, each sub-process has a productivity ranging from 0 to 0.01, so that the productivity of the entire technology when you add them up ranges from 0 to 1.

In their paper, firms use these technologies to produce a fixed amount of output every period. This bypasses the problem highlighted in the previous section, where lower costs lead to increased production - here production is always the same each period, and is therefore unrelated to cost. As firms produce, they also do a bit of experimentation, changing one or more of their sub-processes at a constant rate. When changes result in an increase in overall productivity, the updated technology gets rolled out to the entire production process next period, and experimentation begins starting from this new point.

In this way, production “evolves” towards ever higher productivity and ever lower costs. What’s actually happening is that when a production process is “born” the productivity of all of its sub-processes are just drawn at random so they are all over the map: some high, some low, most average. If you choose a sub-process at random, in expectation it’s productivity will just be the mean of the random distribution, and so if you change it there is a 50:50 shot that the change will be for the better. So progress is fairly rapid at first.

But since you only keep changes that result in net improvements, the productivity of all the sub-processes gets pulled up as production proceeds. As the technology improves, it gets rarer and rarer that a change to a sub-process leads to an improvement. So tinkering with the production process yields an improvement less and less often. Eventually, you discover the best way to do every sub-process, and then there’s no more scope to improve.

But even though this model give you progress that gets harder over time, it actually does not generate a learning curve, where a doubling of cumulative output generates a constant proportional increase in productivity. Instead, you get something like the following figure:

To get a figure that has a linear relationship between the log of cumulative output and the log of costs, the authors instead assume (realistically, in my view) that production is complex and sub-processes are interrelated. In their baseline model, every time you change one subprocess, the productivity of four other sub-processes is also redrawn.

In this kind of model, you do observe something like a learning curve. This seems to be because interdependence changes the rate of progress such that it speeds up in early stages and slows down in later ones. The rate of progress is faster at the outset, because every time you change one subprocess, you actually change the productivity of multiple subprocesses that interact with it. Since these changes are more likely to be improvements at the outset, that leads to faster progress when the technology is young, because you can change lots of things at once for the better.

But when a technology matures, the rate of progress slows. Suppose you have a fairly good production process, where most of the sub-processes have high productivity, but there are still some with low productivity. If you were to tinker with one of the low-productivity sub-processes, it’s pretty likely you’ll discover an improvement. But, you can’t just tinker with that one. If you make a change to the one, it will also lead to a change in several other sub-processes. And most of those are likely to be high-productivity. Which means any gains you make on the low-productivity sub-process will probably be offset by declines in the productivity of other ones with which it interacts.

When you add in these interdependencies between sub-processes, their model generates figures like the following. For much of their life, they look quite a lot like learning curves. (And remember, this is generated with constant demand every period, regardless of cost)

What’s encouraging is the story Auerswald, Kauffman, Lobo, and Shell are telling is one that sounds quite applicable to many real technologies. In lots of technologies there are many sub-components or sub-processes that can be changed, changes may result in improvements or deterioration, and changing one thing frequently affects other sub-components. If you go about this kind of randomly, you can get something that looks like a learning curve.

Evidence from an auto plant

Another paper by Levitt, List, and Syverson (2013) use a wealth of data from an automobile assembly plant to document exactly the kind of learning from experience and experimentation that undergirds learning curves. The paper follows the first year of operation for an auto assembly plant at an unnamed major automaker. Their observations begin after several major changes: the plant went through a major redesign, the firm introduced a new team-based production process, and the vehicle model platform had its first major redesign in six years. Rather than focus on the cost of assembling a car, the paper measures the decline in production defects as production ramps up.

Levitt, List, and Syverson observe a rapid reduction in the number of defects at first, when production is still in early days, followed by a slower rate of decline as production ramps up. Consistent with the learning curve model, the relationship between the log of the defect rate and the log of cumulative production is linear.

Learning-by-doing really makes sense in this context. Levitt, List and Syverson provide some concrete examples of what exactly is being learned in the production process. In one instance, two adjacent components occasionally did not fit together well. As workers and auditors noticed this, they tracked the problem to high variance in the shape of one molded component. By slightly adjusting the chemical composition of the plastic fed into the mold, this variability was eliminated and the problem solved. In another instance, an interior part was sometimes not bolted down sufficiently. In this case, the problem was solved by modifying the assembly procedure for those particular line workers, and adding an additional step for other workers to check the bolt. It seems reasonable to think of these changes as being analogous to changing subprocesses, each of which can be potentially improved and where changes in one process may affect the efficacy of others.

Levitt, List, and Syverson also show that this learning becomes embodied in the firm’s procedures, rather than the skill sets of the individual workers. Midyear, the plant began running a second line and on the second line’s first full day (after a week of training), their defect rate was identical to the first shift workers.

This is a particularly nice context to study because there were no major changes to the plant’s production technology during the period under review. There were not newly designed industrial robots installed midway through the year, or scientific breakthroughs that allowed the workers to be more efficient. It really does seem like, what changed over the year was the plant learned to optimize a fixed production technology.

An Experiment

So we have a bit of theory that shows how learning curves can arise, and we’ve got one detailed case study that seems to match up with the theory pretty well. But it would be really nice if we had experimental data. If we were going to test the learning curve model and we had unlimited resources, the ideal experiment would be to pick a bunch of technologies at random and massively ramp up their production, and then to compare the evolution of their costs to a control set. Better yet, we would ramp up production at different rates, and in a way uncorrelated with time, for example, raising production by a lot but then shutting it down to a trickle in later years. This would give us the variation between time and experience that we need to separately identify the contribution to progress from learning and from other stuff that is correlated with the passage of time. We don’t have that experiment, unfortunately. But we do have World War II.

The US experience in World War II is not a bad approximation of this ideal experiment. The changes in production induced by the war were enormous: the US went from fielding an army of under 200,000 in 1939 to more than 8 million in 1945, and also equipping the allied nations more generally. The production needs were driven by military exigencies more than the price and cost of production, which should minimize reverse causality, where cost declines lead to production increases. Production was also highly variable, so that it is possible to separately identify cost reductions associated with time and cumulative experience. The following figure, for example, illustrates monthly production of Ford’s Armored Car M-20 GBK.

A working paper by Lafond, Greenwald, and Farmer (2020) uses this context to separately identify wartime cost reductions associated with production experience and those associated with time. They use three main datasets:

Man hours per unit over the course of the war for 152 different vehicles (mostly aircraft, but also some ships and motor vehicles)

Total unit costs per product for 523 different products (though with only two observations per product: “early” cost and “later” costs)

Indices of contract prices aggregated up the level of 10 different war sub departments

So, in this unique context, we should be able to accurately separate out the effect of learning-by-doing from other things that reduce cost and are correlated with the passage of time. When Lafond, Greenwald, and Farmer do this, they find that cost reductions associated specifically with experience account for 67% of the reduction in man hours, 40% of the reduction in total unit costs, and 46% of the reduction in their index of contract prices. Learning by doing, at least in World War II, was indeed a significant contributor to cost reductions.

Are learning curves useful?

So where do stand, after all that? I think we have good reason to believe that learning-by-doing is a real phenomenon, roughly corresponding to a kind of evolutionary process. At the same time, it almost certainly accounts for only part of the cost reductions we see in any given project, especially over the long-term when there are large changes to production processes. In particular, the historical relationship between cost reductions and cumulative output that we observe in “normal” circumstances is so hopelessly confounded that we really can’t figure out what share accrues to learning-by-doing and what share to other factors.

That means that if we want to lower the costs of renewable energy (or any other new technology), we can probably be confident they will fall to some degree if we just scale up production of the current technology. But we don’t really know how much - historic relationships don’t tell us much. In World War II, at best, we would have gotten about two-thirds of the rate of progress implied by the headline relationship between cost reduction and cumulative output. Other datasets imply something closer to two fifths. Moreover, the evidence reviewed here applies best to situations where we have a standard production process around which we can tinker and iterate to a higher efficiency. If we need to completely change the method of manufacture or the structure of the technology - well, I don’t think we should count on learning by doing to deliver that.

The next newsletter (January 19) will be about the challenges in getting any academic field to embrace new (better) methodologies, and how the field of economics overcame them.

If you enjoyed this post, you might also enjoy:

More people = more ideas?

Subscribe at mattsclancy.substack.com

View Details

Heads up: this week’s post is a bit different, in that it’s not built primarily from academic papers. We’ll be back to the normal format in 2021 with a post about learning-by-doing and learning curves. Happy Holidays!

GPT-3 is OpenAI's new generative language model. It’s built by feeding a gigantic neural network hundreds of gigabytes of text, pulled from books, wikipedia, and the internet at large. The neural network identifies statistical patterns in text, which are translated into the structure and weights of the neural network. The upshot is that you end up with a predictive text generator. Prompt GPT-3 with a bit of text and it will complete the text based on its underlying model of statistical regularities in text. The results range from spooky good to funny bad.

GPT-3 could end up having a big impact on innovation and research, but that’s not what this post is about. Instead, I’m going to talk about how writing stories with GPT-3 is kind of a neat metaphor for innovation in general.

Let's look at an example. The Eye of Thuban is a short incomplete science fiction-fantasy story about a woman named Vega who lives on an alien world, longs to be a space pilot, and encounters a mysterious artifact that seems to give her powers. It's a joint effort between GPT-3 and the human Arram Sabati, and while it's not particularly good, it is coherent and recognizably a science-fiction story.

Sabeti generated the story by prompting GPT-3 with the following:

This novel is a science fiction thriller that can be thought of as a strange mix of the fantastic and whimsical worlds of Hayo Miyazaki and Ian M. Banks Culture novels. They’re set in a post singularity world where humanity and its descendants span thousands of worlds, and sentient super intelligent ships with billions of people living on them wander the galaxy.

Chapter 1.

Starting with that, GPT-3 wrote a few sentences of story by predicting what kind of text would be most likely to continue on from this prompt. Sabeti accepted or rejected these sentences based on his own judgment: was it coherent? interesting? If not, he would prompt GPT-3 to try again. If he liked it, Sabeti would add the GPT-3 generated text to what he had and use the story so far as the next prompt for GPT-3. It would generate a few more sentences, based on what came before, which Sabeti would accept or reject. Repeat, until you have a story.

Who "wrote" the story that emerged? After the initial prompt, GPT-3 wrote all the text. But Sabeti played a crucial role in selecting text that he liked most, in the process shaping future prompts and pulling the story in one direction or another.

More mysteriously, we have very little understanding of what GPT-3 was "thinking" or "drawing on" when it composed it's contributions. We don't know what words or phrases in the prompts caught GPT-3's attention, and where the patterns it used to generate text originated in it's training data. To Sabeti, it's a black box. Indeed, even to the team at OpenAI that created it, GPT-3 is largely a black box. The way neural nets encode statistical patterns in data is useful, but hard to translate into the kind of explanations that people understand, at least at the present.

So, in essence, through a process Sabeti doesn't understand, text is generated. Sabeti evaluates it critically, and then chooses whether to leave it or try again. At some point, he is satisfied with the results and a story is published. This isn't writing as we know it.

The Mystery of Writing

Or is it? In fact, the kind of creation described above is remarkably similar to how some writers describe their process. George Saunders writes:

we often discuss art this way: the artist had something he “wanted to express”, and then he just, you know … expressed it. We buy into some version of the intentional fallacy: the notion that art is about having a clear-cut intention and then confidently executing same...

The actual process, in my experience, is much more mysterious and more of a pain in the ass to discuss truthfully.

Authors often do not exactly understand where their own ideas come from. Stephen King, in his memoir On Writing writes:

Good story ideas seem to come quite literally from nowhere, sailing at you right out of the empty sky. Two previously unrelated ideas come together and make something new under the sun. Your job isn't to find these ideas but to recognize them when they show up.

In the same memoir, King describes his process of organic and intuitive writing. He advises writers to come up with a scenario and discover what happens to the characters, rather than setting up the plot in advance and maneuvering them through it. After finishing the first draft, he tells writers to reread it to discover what they were "really" writing about. Rather than sculpture, King likens the whole process to unearthing a fossil. The story is discovered, rather than planned.

King and Saunders don't know where they are going; but they still get somewhere interesting. How? Part of the answer is taste. They may not be able to create a satisfying plot twist or turn of phrase on command, but when they see one, they recognize it. Notice how in the previous quote King says:

Your job isn't to find these ideas but to recognize them when they show up.

If you are capable of creating work and then evaluating it, you are capable of using an evolutionary algorithm to write great fiction. Saunders is most explicit about how this works:

My method is: I imagine a meter mounted in my forehead, with “P” on this side (“Positive”) and “N” on this side (“Negative”). I try to read what I’ve written uninflectedly, the way a first-time reader might (“without hope and without despair”). Where’s the needle? Accept the result without whining. Then edit, so as to move the needle into the “P” zone. Enact a repetitive, obsessive, iterative application of preference: watch the needle, adjust the prose, watch the needle, adjust the prose (rinse, lather, repeat), through (sometimes) hundreds of drafts. Like a cruise ship slowly turning, the story will start to alter course via those thousands of incremental adjustments.

More generally, in Old Masters and Young Geniuses: The Two Lifecycles of Artistic Creativity David Galenson describes two broad approaches to artistic creation: conceptual and experimental. The experimental creator is unsure of where they are going and proceeds by creating, evaluating the results, tweaking them, and repeating over and over again. Galenson provides a wealth of anecdotes supporting this creative style for many famous writers: Charles Dickens, Virginia Woolf, Mark Twain, etc.

Just as the blind forces of evolution are capable of creating highly complex creatures via a process of mutation, selection, and retention, this process of myopic, evolutionary writing is capable of generating work that surprises it's own authors. They didn't know they would end up here, but they recognized it was a good place to be, once they did. Again, Saunders:

The interesting thing, in my experience, is that the result of this laborious and slightly obsessive process is a story that is better than I am in “real life” – funnier, kinder, less full of crap, more empathetic, with a clearer sense of virtue, both wiser and more entertaining.

This method of writing sounds remarkably like collaborative writing with GPT-3 to me. Through a lifetime of reading and writing, great writers develop their own intuitive subconscious map of the regularities in good writing. Like GPT-3, they have a kind of prompt - where are they starting from - and like GPT-3, they have some kind of internalized model of what writing looks like, but for which they might struggle to explain the exact sources and reasons. Then, once their intuition or subconscious tosses out some text, they can evaluate whether it’s any good. If so, they keep it and “add it” to the prompt. If not, they try again.

Indeed, François Chollet, makes this analogy pretty explicitly:

Not so different after all

That said, there are many differences too. The extent to which human brains are “like” digital neural networks is not clear, and it may be that the differences are really important. But for the purposes of this essay, that’s not a big problem. What matters is that both processes involve the generation of text via methods that are hard to describe but effective, and then a second step of conscious evaluation of this output.

A more important difference is that the unedited text generated by Saunders or King might be much better than the text generated by GPT-3. Certainly the Eye of Thuban is not a particularly compelling story. In writing The Eye of Thuban, Sabeti notes that sometimes GPT-3 adds text that is logically inconsistent with what has come earlier - we can assume King and Saunders usually don’t do that.

The quality of the text GPT-3 generates isn’t actually so important though, at least for the end result. Instead, the quality of the unedited text generated by Saunders, King, or GPT-3 mostly has an effect on the amount of time that must be spent curating the text. Indeed, if one was willing to put in a lot of time (like, longer than the lifespan of the universe) monkeys banging on keyboards could eventually replicate Shakespeare. GPT-3 is much, much better than that. It’s also clearly an improvement over what’s come before. The pseudonymous Gwern has extensively played with GPT-3 and it’s predecessor GPT-2, writing:

With GPT-2-117M poetry, I’d typically read through a few hundred samples to get a good one… But for GPT-3, once the prompt is dialed in, the ratio appears to have dropped to closer to 1:5—maybe even as low as 1:3! I frequently find myself shrugging at the first completion I generate, “not bad!” (Certainly, the quality of GPT-3’s average prompted poem appears to exceed that of almost all teenage poets.) I would have to read GPT-2 outputs for months and probably surreptitiously edit samples together to get a dataset of samples like this page.

Sabeti claims to have written the Eye of Thuban in the course of a few hours, whereas Saunders describes potentially hundreds of iterations on each draft. It may well be that collaborative writing with GPT-3 could come up with something really good, if one was willing to put in a lot more time.

What GPT-3 does is reduce the time spent iterating to good work, by exploiting regularities in language to avoid wasting the curators time selecting obviously bad passages.

Saunders and King do the same thing - a lifetime of reading, writing, and thinking critically about what they read and write, has led them to internalize “good writing” such that the raw text they come up with is not bad, even before they selectively curate their own writing. Stretching the analogy, we might view the texts Saunders and King read as the training data they use to make their inner models of good writing. Unlike GPT-3, some of the curation that Sabeti performs on GPT-3 probably occurs in a writer's head, rather than through the process of actually writing text and evaluating it (think of a writer mentally searching through a series of adjectives to find just the right one before putting pen to paper). To the extent the uncurated text is already good, this speeds up their iterative process.

There is one respect, however, in which time may not be sufficient to offset any weaknesses in GPT-3’s generated text. It may be that there are certain turns of phrase and sentences that GPT-3 will never produce, simply because they lie so much at odds with its internalized model of the regularities in writing. If this is true, then no amount of time would suffice to generate good writing from GPT-3. In this regard, GPT-3 could potentially be worse than monkeys pounding on keyboards, since they are at least capable of generating any text. The tradeoff seems unavoidable; when you exploit regularities in language to weed out certain passages and save the collaborator time, you also might weed out good passages that do not exhibit these regularities. We just hope these cases are rare.

Beyond Writing

What we have then, in both collaborative writing with GPT-3 and a certain method of professional writing, is an iterative process of generation and evaluation. Good writing does not emerge fully formed - instead, many texts are generated, and good ones are retained. If we were to generate our texts randomly, we could still generate great writing - but it would take a very long time, because most random text is gibberish. GPT-2 and GPT-3 greatly reduce the time necessary to generate great writing by restricting the generation of texts to those more likely to be good (they are coherent, and obey certain regularities in language). Great writers do both sides of the process - they have good subconscious models of writing, so that their raw text is reasonably good, and they have great taste, so they can prune their output and so direct their writing towards interesting ends.

Collaborative writing with GPT-3 is more than an analogy for good writing though; it’s an analogy for innovation in general.

Let’s pause briefly to consider what “innovation” is. To me, innovation is the emergence of interesting, reproducible novelty. Novel, because it must be something that has not been done before. Reproducible, because the innovation cannot be one-time miracle, but a new class of thing which can serve (in principle) as a blueprint for more of its kind. And interesting, because otherwise, who cares? The challenge with innovation is that most novel things are not interesting. How to find the few that are?

In practice, what we do is use some kind of simplified model of reality to guide our efforts, so that we don’t just try things at random. That “model” could literally be a scientific model of the physical world; these are, after all, ways of representing observed regularities in data. But they could also be much more prosaic: rules of thumb, analogies to other examples, or expert intuition (built up from long study of relevant precedents). When these models are good, they allow us to develop ideas and technologies that would take an eternity to arrive at by evolution or random chance. When they are bad, they restrict us and prevent us from trying things that would have worked if we could only take off our blinders.

In the analogy of collaborative writing with GPT-3 to innovation, our models of the world are analogous to GPT-3, but it is the world itself that is analogous to the human collaborator. The world is messy and complex, and even the best models may miss things. It is only when innovations are brought into reality - whether as clinical trials, prototypes, product launches, or startups - that we see if they are, in fact, “interesting.” Those that are, are retained. Like GPT-3’s text in The Eye of Thuban, they get added to the “prompt” and further work builds on the ideas and innovations that have been retained. For the rest, we “go back to the drawing board” (our simplified model of the world) and try again.

If you liked this post, you might also like the following:

Progress or stagnation in film?

Innovation as combination

Next week, “How useful are learning curves, really?”

Subscribe at mattsclancy.substack.com

View Details

How did we end up in a situation where so many scientific papers do not replicate? Replication isn’t the only thing that counts in science, but there are lots of papers that, if they actually describe a regularity or causal mechanism in the world, then we should be able to replicate it. And we can’t. How did we get here?

One theory (not the only one), is that the publish-or-perish system is to blame.

In an influential 2016 paper, Paul Smaldino and Richard McElreath simulated science in action with a simple computer simulation. Their simulation is a highly simplified version of science, but it captures the contours of some fields well. In their simulation, “science” is nothing but hypothesis testing (that is, using statistics and data to assess whether the data is consistent or inconsistent with various hypotheses). One hundred simulated labs pursue various research strategies and attempt to publish their results. In this context, a “research strategy” is basically just three numbers:

A measure of how much effort you put into each research project: the more effort you put in, the more accurate your results, but the fewer projects you finish

A measure of what kinds of protocols you use to detect a statistically significant event: you can trade off false negatives (incorrectly rejecting a true hypothesis) and false positives (incorrectly affirming a false one)

The probability you choose to replicate another lab’s findings or investigate a novel hypothesis

At the end of each period, labs either do or do not finish their project. If they do, they get a positive or null result. They then attempt to publish what they’ve got. Next, a random set of labs is selected and the oldest one “dies.”

Over time, labs accumulate prestige (also a number) based on their publishing record. Prestige matters because at the end of every period, the simulation selects another set of random labs. The one with the highest prestige spawns a new lab which follows similar (though not necessarily identical) research strategies as it’s “parent.” This is meant to represent how successful researchers propagate their methods via training postdocs who go on to form their own labs or via imitation of their methods by new labs who attempt to emulate prestigious work.

Lastly, Smaldino and McElreath assume prestige is allocated according to the following rules:

Positive results are easier to publish than null results

More publications leads to more prestige

Replications give less prestige than novel hypotheses

What happens when you simulate this kind of science is not that surprising: labs with low effort strategies that adopt protocols conducive to lots of false positives publish more often then those that try and do things “right.” Let’s call this kind of research strategy “sloppy” science. Note that it may well be that these labs sincerely believe in their research strategy - there is no need in this model for labs to be devious. But by publishing more often, these labs become more prestigious and over time they spawn more labs, so that their style of research comes to dominate science. The result is a publication record that is riddled with false positives.

In short, Smaldino and McElreath suggest the incentive system in science creates selective pressures where people who adopt research strategies that lead to non-replicable work thrive and spread their methods. If these selective pressures don’t change, no amount of moral exhortion to do better will work; those who listen will always be outcompeted by those who don’t, in the long run. In fact, Smaldino and McElreath show that, despite warning about the poor methodologies in behavioral science that date back to at least the 1960s, in 44 literature reviews (see figure below) there has been no increase in the average statistical power of hypothesis tests in the social and behavioral sciences.

Is it a good model?

Smaldino and McElreath’s simulation suggests it’s the incentive schemes (and their effect on selection) we currently use that lead to things like the replication crisis. So what if we changed the incentive system? Two recent papers look at the conduct of science for projects that are just about identical, except for the incentives faced by the researchers.

First, let’s take a look at a new working paper by Hill and Stein (2020). Smaldino and McElreath basically assert that the competition for prestige (only those with the most prestige “reproduce”) leads to a reduction in effort per research project, which results in inferior work (more likely to be a false positive). Hill and Stein document that this is indeed the case (at least in one specific context): competition for prestige leads to research strategies that produce inferior work fast. They also show this doesn’t have to be the case, if you change the incentives of researchers.

Hill and Stein study structural biology, where scientists try to discover the 3D structure of proteins using modeling and data from x-ray scattering off protein crystals. (Aside: this is the same field that was disrupted last week by the announcement that DeepMind’s AlphaFold had made a big leap in inferring the structure of proteins based on nothing more than DNA sequence data). What makes this setting interesting is a dataset that lets Hill and Stein measure the effort and quality of research projects unusually well.

Structural biology scientists report to a centralized database whenever they take their protein crystals to a synchrotron facility, where they obtain their x-ray data. Later, they also submit their final structures to this database, with a time-stamp. By looking at the gap between the receipt of data and the submission of the final protein model, Hill and Stein can see how much time the scientists spend analyzing their data. This is their measure of how much effort scientists put into a research project.

The database also includes standardized data on the quality of each structural model: for example, how well does the model match the data, what is the resolution of the model, etc. This is a key strength of this data: it’s actually possible to “objectively” rate the quality of research outputs. They use this data to create an index for the “quality” of research.

Lastly, of course, since scientists report when they take their sample to a synchrotron for data, Hill and Stein know who is working on what. Specifically, they can see if there are many scientists working on the same protein structure.

The relevant incentive Hill and Stein investigate is the race for priority. There is a norm in science that the first to publish a finding receives the lion’s share of the credit. There are good arguments for this system, but priority can also lead to inefficiency when multiple researchers are working on the same thing and only the first to publish gets credit. In a best case scenario, this race for priority means researchers pour outsized resources into advancing publication by a few days or weeks, with little social benefit. In a worst case scenario, researchers may cut corners to get their work out more quickly and win priority.

Hill and Stein document that researchers do, in fact, spend less time working with their data to build a protein model, when there are more rivals working on the same protein at the same time. They also show this leads to a measurable decline in the quality of the models. Moreover, based on some rules of thumb about how good a protein model needs to be for application in medical innovation, this quality decline probably has a non-negligible impact on things non-scientists care about, like the development of drugs.

But wait, it gets worse. Why do some proteins attract the attention of lots of scientists, and others not? It’s not random. In fact, Hill and Stein provide evidence that the proteins with the most “potential” (i.e., the ones that will get cited the most in other academic papers when their structure is found) are the ones that attract the most researchers. (Aside: Hill and Stein do this with a LASSO regression that predicts the percentile citation rank of each protein based on the data available on it prior to its structure being discovered).

In short, the most interesting proteins attract the most researchers. The more intense competition, in turn, leads these researchers to shorten the time they spend on modeling the protein, in an attempt to get priority. That, in turn, leads to the most inferior modeling on the proteins we would like to know the most about.

Hill and Stein’s paper is about one of the downsides of the priority system. This is a bit different than Smaldino and McElreath, where prestige comes from the number of publications one has. However, in Smaldino and McElreath, their simulated labs can die at any moment, if they are the oldest one in a randomly selected sample. This means the labs that spawn are the ones who are able to rapidly accrue a sizable publication record - since if you can’t get one fast, you might not live to get one at all. As in Hill and Stein, one way labs do this is by cutting back on the effort they put into each research project.

Different Incentives, Different Results?

However, academics who are judged on their publication record aren’t the only people doing structural biology. “Structural genomics” researchers are “federally-funded scientists with a mission to deposit a variety of structures, with the goal of obtaining better coverage of the protein-folding space and mak[ing] future structure discovery easier” (Hill and Stein, p. 4). Hill and Stein argue that this group is much less motivated by publication than the rest of the structural biology. For example - only 20% of the proteins they work on end up with an associated academic paper (compared to 80% for the rest of structural biology). So if they aren’t driven by publication, is the quality of their work different?

Yes! Unlike the rest of structural biology, on average this group is likely to spend more time on proteins with more potential. In the above diagram, they are the red line, which slopes up. And while the quality of the models they generate for highest potential proteins is still a bit lower than the low potential ones, the strength of this relationship is much smaller than it is for those chasing publication.

One other recent study provides some further suggestive evidence that different incentives produce different results - or at least, the perception of different results. Bikard (2018) looks at how research produced in academia is viewed by the private sector, as compared to research produced by the private sector (think papers published by scientists working for business). Specifically, are patents more likely to cite academic or private sector science?

The trouble is this will be an apples-to-oranges comparison if academia and the private sector focus on different research questions. Maybe the private sector thinks academic research is amazing, but simply not relevant to private sector needs most of the time. In that case, they might cite private sector research at a higher rate, but still prefer academic research whenever it is relevant.

To get around this problem, Bikard identifies 39 instances where the same scientific discovery was made independently by academic and industry scientists. He then shows that patents tend to disproportionately cite the industry paper on the discovery, which he argues is evidence that inventors regard academic work skeptically, as compared to work that emerges from industry research.

To identify these cases of simultaneous discovery, Bikard starts with the assumption that if two papers are consistently cited together in the same parenthetical block, like so - (example A, example B) - then they may refer to the same finding. After identifying sets of papers consistently cited together this way, he provides further supporting evidence that this system works. He shows the sets of “twin” papers he locates are extremely similar when analyzed with text analysis algorithms, that they are almost always published within 6 months of each other, and that they are very frequently published literally back-to-back in the same journal (which is one way journals acknowledge simultaneous discovery).

This gives Bikard a nice dataset that, in theory, controls for the “quality” and relevance of the underlying scientific idea being described in the paper. This provides a nice avenue for seeing how academic work is perceived, relative to industry. When an inventor builds on the scientific discovery and seeks a patent for their invention, they can, in principle, cite either paper or both since the discovery is the same either way. But Bikard finds papers that emerge from academia were 23% less likely to be cited by patents than an industry paper on the same discovery.

This preference for industry research could reflect a lot of things. But Bikard goes on to interview 48 scientists and inventors about all this and the inventors consistently say things like the following, from a senior scientist at a biotechnology firm:

The principle that I follow is that in academia, the end game is to get the paper published in a as high-profile journal as possible. In industry, the end game is not to get a paper published. The end game is getting a drug approved. It's much, much, much harder, okay? Many, many more hurdles along the way. And so it's a much higher bar - higher standards - because every error, or every piece of fraud along the way, the end game is going to fail. It's not gonna work. Therefore, I have more faith in what industry puts out there as a publication.

So, to sum up, we’ve got evidence that non-academic consumers of science pay more attention to the results that come from outside academia, with some qualitative evidence that this is because academic science is viewed as lower quality. We’ve also got good data from one particular discipline (structural biology) that publication incentives lead to measurably worse outcomes. I wish we had more evidence to go on, but so far what we have is consistent with the simple notion that different incentive systems do seem to get different results in a way that moral exhortation perhaps does not.

But maybe you already believed incentives matter. In that case, one nice thing about these papers is they provide a sense of the magnitude of how bad academic incentives screw up science. From my perspective, the magnitudes are large enough that we should try to improve the incentives we have, but not so large that I think science is irredeemably broken. Hill and Stein find the impact of priority races reduces research time from something like 1.9 years to 1.7 years, not from 1.9 years to something like 0.5 years. And though the quality of the models generated is worse, Hill and Stein do find that, in subsequent years, better structure models eventually become available for proteins with high potential (at significant cost in terms of duplicated research). And even if inventors express skepticism towards academic research, they still cite it at pretty high rates. We have a system that, on the whole, continues to produce useful stuff I think. But it could be better.

If you thought this was interesting, you might also like these posts:

Does chasing citations lead to bad science?

Does science self-correct?

Subscribe at mattsclancy.substack.com

View Details

An update on this project: I really enjoy writing this newsletter and have gotten good feedback on it. But it takes time (you may have noticed that there were only three newsletters in the last 6 months). So I’m very pleased to announce I’ve received funding from Emergent Ventures and the go-ahead from Iowa State University to carve out a bit of time specifically for this project, at least for the next year. The plan is to release a newsletter like this one every other Tuesday.

After a year, I’ll re-assess things. If you want to help make this project a success, you can subscribe or tell other people about it. Thanks for your interest everyone!


It seems clear that a better understanding of the regularities that govern our world leads, in time, to better technology: science leads to innovation. But how long does this process take?

Two complementary lines of evidence suggest 20 years is a good rule of thumb.

A First Crack

James Adams was one of the first to take a crack at this in a serious quantitative way. In 1990, Adams published a paper titled "Fundamental Stocks of Knowledge and Productivity Growth" in the Journal of Political Economy. Adams wanted to see how strong was the link between academic science and the performance of private industry in subsequent decades.

Adams had two pieces of data that he needed to knit together. On the science side, he had data on the annual number of new journal articles in nine different fields, from 1908 to 1980. On the private industry side, he had data on the productivity of 18 different manufacturing industries over 1966-1980. Productivity here is a measure of how much output a firm can squeeze out of the same amount of capital, labor, and other inputs. If a firm can get more output (or higher quality outputs) from the same amount of capital and labor, economists usually assume that reflects an improved technology (though it can also mean other things). What Adams basically wanted to do was see if industries experienced a jump in productivity sometime after a jump in the number of relevant scientific articles.

The trouble is that pesky word “relevant.” Adams has data on the number of journal articles in fields like biology, chemistry, and mathematics, but that's not how industry is organized. Industry is divided into sectors like textiles, transportation, and petroleum. What scientific fields are most relevant to the textiles industry? To transportation equipment? To petroleum?

To knit the data together, Adams used a third set of data: the number of scientists in different fields that work for each industry. To see how much the textiles sector relies on biology, chemistry, or mathematics, he looked at how many biologists, chemists, and mathematicians the sector employed. That data did exist. If they employed a lot of chemists, they probably used chemistry; if they employed lots of biologists, they probably used biology, and so on. He weighted the number of articles in each field by the number of scientists working in that field to get a measure of how much each industry relies on basic science.

So now Adams has data on the productivity and relevant scientific base of 18 different manufacturing sectors. He expects more science will eventually lead to more productivity, but he doesn’t know how long that will take. If there was a surge in scientific articles in a given year, at what point would Adams expect to see a surge in productivity? If scientific insight can be instantly applied, then the surge in productivity should be simultaneous. But if the science has to work it's way through a long series of development, then the benefits to industry might show up later. How much later?

To come up with an estimate, Adams basically looked at how strong the correlations were for 5 years, 10 years, and 20 years. Specifically, he looked to see which one gives him the strongest statistical fit between scientific articles produced in a five-year span, and the productivity increase in a five-year span for industries that use that field's knowledge intensively. Of the time lags he tried, he found the strongest correlation was 20 years.

Adams’ study is an important first step, but recent work has largely validated Adams’ original findings.

A Bayesian Alternative

Nearly thirty years later, Baldos et al. (2018) tackled a similar problem with different data and a more sophisticated statistical technique. Unlike Adams, they focused on a single sector - agriculture. Like Adams, they had two pieces of data they wanted to knit together.

On the science side, a small group of ag economists has spent a long time assembling a data series on total agricultural R&D spending by US states and the federal government. Until recently, governments were a really big source of agricultural research dollars and those dollars can be relatively easily identified since they frequently flow through the department of agriculture or state experiment stations. The upshot is Baldos et al. have data on public sector agricultural R&D going back to 1908. Meanwhile, on the technology side, the USDA maintains a data series on the total factor productivity of US agriculture from 1949 to present day. So like Adams, they’re going to try and look for a statistical correlation between productivity and research. Unlike Adams, they’re going to use dollars to measure science and unlike Adams they’re going to focus on a single sector (and not a manufacturing sector).

To deal with the fact that we don’t know when research spending impacts productivity growth, they’re going to adopt a Bayesian approach. How long does it take agricultural spending to influence productivity? They don’t know, but they make a few starting assumptions:

They assume impact will follow an upside “U” shape. The idea here is that new knowledge takes time to be developed into applications, and then it takes time for those applications to be adopted across the industry. During this time, the impact of R&D done in one year on productivity in subsequent years is rising. At some point - maybe 20 years later - the impact of that R&D on productivity growth hits its peak. But after that point, the R&D becomes less relevant, so that it’s impact on productivity growth in subsequent years declines. Eventually, the ideas become obsolete and have no additional impact on increasing productivity.

They assume the impact of R&D on productivity after 50 years is basically zero.

They assume the peak impact will occur sometime between 10 and 40 years, with the most likely outcome somewhere around 20. This assumption is based on earlier work, similar in spirit to Adams.

Given these estimates, they are basically assuming there’s a bunch of different possible distributions for the relationship between R&D spending and productivity. They assume the most likely distribution is one peaking around 20 years, and the farther the distribution is from that, the more unlikely. They then use Bayes’ rule to update their beliefs, given the data on R&D spending and agricultural productivity. It will turn out that some of those distributions fit the data quite well, in the sense that if that distribution of R&D impacts is true, then the R&D spending data matches productivity pretty well. Others fit quite poorly. We update our beliefs after observing the data, increasing our belief in the ones that fit the data well, and decreasing our beliefs in the ones that don’t.

They find, with 95% probability, the best fitting distribution indicates science impacts productivity the strongest after 15-24 years, with the best point estimate around 20 years.

Evidence from Citations

So whether for manufacturing or agriculture, using slightly different data and statistical techniques, we find a correlation between productivity growth and basic science that is strongest around 20 years. But at the end of the day, both methods look for correlations between two messy variables separated in time by decades. It’s hard to make this completely convincing.

A cleaner alternative is to look at the citations made by patents to scientific articles. If we assume patents are a decent measure of technology (more on that later) and journal articles a decent measure of science, then citations of journal articles by patents could be a direct measure of technology’s use of science. In the last few years, data on patent citations to journal articles has become available at scale, thanks to advances in natural language processing (see Marx and Fuegi 2019, Marx and Fuegi 2020, and the patCite project).

Do citations to science actually mean patented technologies use the science? Maybe the citations are just meant to pad out the application? There's a lot of suggestive evidence that they do. For one; they say they do. Arora, Belenzon, and Sheer (2017) use an old 1994 survey from Carnegie Mellon about the use of science by firms. Firms that cite a lot of academic literature in their patents are also more likely to report in the survey that they use science in the development of new innovations.

There's also a variety of evidence that patents that cite scientific papers are different from patents that don't. Watzinger and Schnitzer (2019) scan the text of patents, and find patents that cite science are more likely to include combinations of words that have been rare up until the year the patent was filed. This suggests these patents are doing something new; they don't read like older patents. Maybe they are using brand new ideas or insights that they have obtained from science? They also find these patents tend to be worth more money. Ahmadpoor and Jones (2017) find these science-ey patents also tend to be more highly cited by other patents.

So let’s go ahead and assume citations to journal articles are a decent measure of how technology uses science. What do we learn from doing that?

Marx and Fuegi (2020) apply natural language processing and machine learning algorithms to the raw text of all US patents, to pull out all citations made to the academic literature. They find about 29% of patents make some kind of citation to science. More importantly for our purposes, for patents granted since 1976, the average gap in time between a patent’s filing date and the publication date of the scientific articles it cites is 17 years.

This is pretty close to the twenty years estimated using the other techniques; especially when we consider that the date a patent is filed is not necessarily the date it begins to affect productivity (which is what the other studies were measuring). Any new invention that seeks patent protection might take a few years before it is widely diffused and begins to affect productivity. In that case, a twenty year estimate is pretty close to spot on.

Winding paths from science to technology

Of course, there is variation around that 17 year figure. Ahmadpoor and Jones (2017) found that the average shortest time gap between a patent and a cited journal was just 7 years. And on the other side, there are also longer and more indirect paths from science to technology than direct citation. Ahmadpoor and Jones cite Riemannian geometry as an example. Riemannian geometry was developed by Bernhard Riemann in the 19th century as an abstract mathematical concept, with little or no real world application until it was incorporated in Einstein's general theory of relativity. Later, those ideas were used to develop the time dilation corrections in GPS satellites. That technology, in turn, has been useful in the development of autonomous vehicles. In a sense then, patents for self-driving tractors owe some of their development to Riemannian geometry, though this would only be detectable by following a long and circuitous route of citation.

When these longer and more circuitous chains of citation are followed, Ahmadpoor and Jones find that 61% of patents are "connected" to science; patents that do not directly cite scientific articles may cite patents that do, or they may cite patents that cite patents that do, and so on. About 80% of science and engineering articles are "connected" to patents, in the sense that some chain of citation links them to a patent. (Technical note: this probably understates the actual linkages, because it is based only on citations listed on a patent’s front page; Marx and Fuegi 2020 show additional citations are commonly found in a patent’s text).

The authors compute metrics for the average "distance" of a scientific field from technological application and these distance metrics largely line up with our intuitions. For example, material science and computer science papers are fields where we might expect the results to be quite applicable to technology, and indeed, they tend to be among the closest to patents - just 2 steps away from a patent (cited by a paper that is, in turn, cited by a patent). Atomic/molecular/chemical physics papers would also seem likely to have applications, but only after more investigation, and they tend to be 3 steps removed from technology (cited by a paper that is cited by a paper that is cited by a patent). And as we might expect, the field that is farthest from immediate application is pure mathematics (five steps removed). As expected, these longer citation paths also take longer when measured in time. The average gap between the shortest path from a math paper to a patent is more than 20 years.

On the other side of the divide, they also compute the average distance of technology fields to science, and these also align with our intuitions. Chemistry and molecular biology patents would probably be expected to rely heavily on science, and they tend to be slightly more than 1 step removed from science (most directly cite scientific papers). Further downstream, electrical computers tend to be two steps removed (they cite a patent that cites a scientific article) and ancient forms of technology like chairs and seats tend to be the farthest from science (five steps removed). The average gap between the shortest citation path from a chair/seat patent and a scientific article is also over 20 years.

Converging Evidence

All told, it’s reassuring that two distinct approaches arrive at a similar figure. The patent citation evidence is reasonably direct - we see exactly what technology at what date cites which scientific article, as well as the date that article was published. This line of evidence finds an average gap of about 17 years, with plenty of scope for shorter and longer gaps as well. The trouble with this evidence is that patents are far from a perfect measure of technology. Lots of things do not ever get patented, and lots of patents are for inventions of dubious quality.

For that reason, it’s nice to have a different set of evidence that does not rely on citations or patents at all. When we try to crudely measure “science”, either by counting government dollars or scientific articles, we can detect a correlation with increases in science and the productivity of industries expected to use it 20 years later - again, with plenty of scope for shorter and longer lags as well.

If you thought this was interesting, you may also enjoy these posts:

How useful is science?

Does chasing citations lead to bad science?

Subscribe at mattsclancy.substack.com

View Details

Listen now | When one country has exceptional expertise that another lacks, what happens when knowledge workers migrate from the one to the other?

Subscribe at mattsclancy.substack.com

View Details

Innovation appears to be getting harder. At least, that’s the conclusion of Bloom, Jones, Van Reenen, and Webb (2020). Across a host of measures, getting one “unit” of innovation seems to take more and more R&D resources.

To take a concrete example, although Moore’s law has held for a remarkable 50 years, maintaining the doubling schedule (twice the transistors every two years) takes twice as many researchers every 14 years. You see similar trends for medical research - over time, more scientists are needed to save the same number of years of life. You see similar trends for agriculture - over time, more scientists are needed to increase crop yields by the same proportion. And you see similar trends for the economy writ large - over time, more researchers are needed to increase total factor productivity by the same proportion. Measured in terms of the number of researchers that can be hired, the resources needed to get the same proportional increase in productivity doubles every 17 years.

There are lots of issues with any one of these numbers. I’ve written about some of them (on the recent total factor productivity slowdown here, and on agricultural crop yields here). But taken together, the effects are so large that it does look like something is happening: it takes more people to innovate over time.

Why?

The Burden of Knowledge

A 2009 paper by Benjamin Jones, titled The Burden of Knowledge and the Death of the Renaissance Man, provides a possible answer (explainer here). Assume invention is the application of knowledge to solve problems (whether in science or technology). As more problems are solved, we require additional knowledge to solve the ones that remain, or to improve on our existing solutions.

This wouldn’t be a problem, except for the fact that people die and take their knowledge with them. Meanwhile, babies are (inconveniently) born without any knowledge. So each generation needs to acquire knowledge anew, slowly and arduously, over decades of schooling. But since the knowledge necessary to push the frontier keeps growing, the amount of knowledge each generation must learn gets larger. The lengthening retraining cycle slows down innovation.

Age of Achievement

A variety of suggestive evidence is consistent with this story. One line of evidence is the age when people begin to innovate. If people need to learn more in order to innovate, they have to spend more time getting educated and will be older when they start adding their own discoveries to the stock of knowledge.

Brendel and Schweitzer (2019) and Schweitzer and Brendel (2020) look at the age of academic mathematicians and economists when they publish their first solo-authored article in a top journal: it rose from 30 to 35 over 1950-2013 (for math) and 1970-2014 (for economics). For economists, they also look at first solo-authored publication in any journal: the trend is the same. Jones (2010) (explainer here) looks at the age when Nobel prize winners and great inventors did their notable work. Over the twentieth century, it rose by 5 more years than would be predicted by demographic changes. Notably, the time Nobel laureates spent in education also increased - by 4 years.

Brendel and Schweitzer (2019) and Schweitzer and Brendel (2020) also point to another suggestive fact that the knowledge required to push the frontier has been rising. The number of references in mathematicians and economists’ first solo-authored papers is rising sharply. Economists in 1970 cited about 15 papers in their first solo-authored article, but 40 in 2014. Mathematicians cited just 5 papers in the 1950s in their debuts, but over 25 in 2013.

Outside academia, the evidence is a bit more mixed. In Jones’ paper on the burden of knowledge, he looked at the age when US inventors get their first patents and found it rose by about one year, from 30.5 to 31.5, between 1985 and 1998. But this trend subsequently reversed. Jung and Ejermo (2014), studying the population of Sweden, found the age of first invention dropped from a peak of 44.6 in 1997 to 40.4 in 2007. And a recent conference paper by Kaltenberg, Jaffe, and Lachman (2020) found the age of first patent between 1996 and 2016 dropped in the USA as well.

That said, there is some other suggestive evidence that patents these days draw on more knowledge - or at least, scientific knowledge - than in the past. Marx and Fuegi (forthcoming) use text processing algorithms to match scientific references in US and EU patents to data on scientific journal articles in the Microsoft Academic Graph. The average number of citations to scientific journal articles has grown rapidly from basically 0 to 4 between 1980 and today. And as noted in a previous newsletter, there’s a variety of evidence that this reflects actual “use” of the ideas science generates.

Splitting Knowledge Across Heads

But that’s only part of the story. In Jones’ model, scientists don’t just respond to the rising burden of knowledge by spending more time in school. They also team up, so that the burden of knowledge is split up among several heads.

The evidence for this trend is pretty unambiguous. The rise of teams has been documented across a host of disciplines. Between 1980 and 2018, the number of inventors per US patent doubled. Brendel and Schweitzer also show the number of coauthors on mathematics and economics articles has also risen sharply through 2013/2014. Wuchty, Jones, and Uzzi (2007) has also documented the rise of teams in scientific production through 2000.

We can also take inspiration from Jones (2010) and look at Nobel prizes. The Nobel prize in physics, chemistry, and medicine has been given to 1-3 people for most of the years from 1901-2019. When more than one person gets the award, it may be because multiple people contributed to the discovery, or because the award is for multiple separate (but thematically linked) contributions. For example, the 2009 physics Nobel was one half awarded to Charles Kuen Kao "for groundbreaking achievements concerning the transmission of light in fibers for optical communication", with the other half jointly to Willard S. Boyle and George E. Smith "for the invention of an imaging semiconductor circuit - the CCD sensor."

The figure below gives the average number of laureates per contribution, over the preceding 10 years. For the physics and chemistry awards, there’s been a steady shift: in the first part of the 20th century, each contribution was usually assigned to a single scientist. In the 21st centruy, there are, on average, two scientists awarded per contribution. In medicine, there was a sharp increase from 1 scientist per contribution to a peak of 2.6 in 1976, but has slightly declined since then, though it remains above 2.

According to Jones’ the reason for teams is that teams can bring more knowledge to a problem than an individual. If that’s the case, then innovations that come from teams should tend to perform better than those created by individuals, all else equal. For both patents and papers, that’s precisely what Ahmadpoor and Jones (2019) find. For teams of 2-5 people, the bigger the team the higher the citations the paper/patent receives (though the extent varies by field). Wu, Wang, and Evans (2019) also find the bigger the team, the more cited are patents, papers, and software code.

The Death of the Renaissance Man

By using teams to innovate, scientists and innovators reduce the amount of time they need to spend learning. They do this by specializing in obtaining frontier knowledge on an ever narrower slice of the problem. So Jones’ model also predicts an increase in specialization.

In Jones’ paper, specialization was measured as the probability solo-inventors patented in different technological fields within 3 years on consecutive patents. The idea is the less likely they are to “jump” fields, the more specialized their knowledge must be. For example, if I apply for a patent in battery technology in 1990 and another in software in 1993, that would indicate I’m more of a generalist than someone who is unable to make the jump. Jones used data on 1977 through 1993, but in the figure below I replicate his methodology and bring the data up through 2010. Between 1975 and 2005, the probability a solo-inventor patents in different technology classes, on two consecutive patents with applications within 3 years of each other, drops from 56% to 47%.

(While the probability does head back up after 2005, it remains well below prior levels and it's possible this is an artifact of the data - see the technical notes at the bottom of this newsletter if curious)

Schweitzer and Brendel exploit the JEL classification system in economics. These classifications can be aggregated up to the level of one of 9 fields, and Brendel and Schweitzer look at the probability an economist hops from one field to another between two solo-authored publications that are published within 3 years. Among all articles listed on EconLit, it's fallen in half, from 33% to 14% between 1973 and 2014. Restricting attention to top ten publications, it fell even more sharply, from 28% to 0%(!) in 2014.

Lastly, let’s consider the Nobel prizes again. Since Nobel prizes are awarded for substantially distinct discoveries, winning more than one Nobel prize in physics, chemistry, or medicine, may be another signifier of multiple specialties. There have been just three Nobel laureates to win more than one physics, chemistry, or medicine Nobel prize: Marie Curie (1903, 1906), John Bardeen (1956, 1972), Frederick Sanger (1958, 1980). If it takes as long as 25 years to receive a second Nobel prize, then we can be sure there was no multiple-winner between 1958 and 1994. There were 218 Nobel laureates between 1959 and 1994, compared to 207 between 1901 and 1958. That means there were 3 multiple Nobel laureates in the first 207, and 0 in the second 218.

Why are ideas getting harder to find?

Bloom, Jones, Van Reenen and Webb (2020) document the productivity of research is falling: it takes more inputs to get the same output. Jones (2009) provides an explanation for why that might happen. New problems require new knowledge to solve, but using new knowledge requires understanding (at least some) of the earlier, more basic knowledge. Over time, the total amount of knowledge needed to solve problems keeps rising. Since knowledge can only be used when it’s inside someone’s head, we end up needing more researchers. And that’s precisely the dimension that Bloom et al. (2020) use to measure the declining productivity of research - it does take more researchers to get the same innovation.

A few closing thoughts.

First, while the evidence discussed above is certainly consistent with Jones’ story, stronger evidence would be nice. Most of the above evidence is about how things have changed over time. But we should also be able to see differences across fields. The story predicts fields with “deeper” knowledge requirements should have bigger teams and more specialization. Jones (2009) provides evidence this is indeed the case for patents, but as far as I know, no one else has updated his work or extended this line of evidence into academia and other domains.

Second, Jones’ model isn’t the only possible explanation for the falling productivity of research. Arora, Belenzon, Patacconi, and Suh (2020) suggest the growing division of labor between universities and the private sector in innovation may be at fault. As universities increasingly focus on basic science and the private sector on applied research, there may be greater difficulty in translating science into applications. Bhattacharya and Packalen (2020) suggest the incentives created by citation in academia have increasingly led scientists to focus on incremental science, rather than potential (risky) breakthroughs. Lastly, it may also be that breakthroughs just come along at random, sometimes after long intervals. Maybe we are simply awaiting a new paradigm to accelerate innovation once again.

Third, where do we go from here? Is innovation doomed to get harder and harder? There are a few possible forces that may work in the opposite direction.

If breakthroughs in science and technology wipe the slate clean, rendering old knowledge obsolete, then it’s possible the burden of knowledge could drop. In fact, Jung and Ejermo (2014) suggest this may be a reason why the age of first patent declined in the mid-1990s: digital innovation became relatively easy and did not depend on deep knowledge. It would be interesting to see if the three measures discussed above tend to reverse in fields undergoing paradigm shifts.

On the other hand, the burden of knowledge may, itself, make breakthroughs more difficult! As discussed in more detail in a previous newsletter, there is some evidence that teams are less likely to produce breakthrough innovations. This might be because it’s harder to spot unexpected connections between ideas when they are split across multiple people’s heads. In that case, the burden of knowledge can become self-perpetuating.

Alternatively, if knowledge leads to greater efficiency in teaching, so that students more quickly vault to the knowledge frontier, that could also reduce the burden of knowledge. Lastly, it may be possible for artificial intelligence to shoulder much of the burden of knowledge. Indeed, artificial general intelligence could hypothetically upend this whole model, if it is disrupts the cycle of retraining and teamwork that is required of human innovators. I suppose we’ll know more in 20 years.

Technical Notes

For patent data, I use US patentsview data and their disambiguated inventor data. To calculate the probability of jumping fields, I use the primary US patent classification 3-digit class (as in Jones 2009). This patent classification system was discontinued in mid-2015, and it’s possible this is a contributing factor to the uptick observed after 2005. A patent applied for in 2006 only “counts” as a possible field jump if there was a second patent applied for before 2010 and granted before the classification system was discontinued in 2015. This selection effect might be result in an increasingly unrepresentative sample of patents.

Subscribe at mattsclancy.substack.com

View Details

Note: I’m experimenting with Substack’s audio features. This week you can read the newsletter below, or listen to me read it by clicking the link above. Thanks!

“No amount of real resources devoted to medical research would have helped European society in 1348 to solve the riddle of the Black Death.” - Joel Mokyr (1998)

There is no currently existing human vaccine for covid-19. Can we force one into existence by promising to spend a lot on it? Is there some price at which we can “buy” a covid-19 vaccine in the next year?

That’s the premise of a proposal by economists Susan Athey, Michael Kremer, Christopher Snyder, and Alex Tabarrok. They propose the US government commit in advance to paying a substantial price for a specified number of vaccine doses: something like $100 each for the first 300 million. The idea is that the potential of winning $30bn will induce pharma companies to pour resources into vaccine development.

Covid-19 and the Profit Motive

We have lots of reasons to believe that a promise to pay more for a covid-19 vaccine would induce more covid-19 vaccine work. Academic research is moving so fast these days that we already have good evidence that pharma companies are extremely responsive to profit signals around covid-19. Bryan, Lemus, and Marshall (2020) track the number of covid-19 therapies at any stage of development, as well as the number of academic publications related to covid-19, to produce this stunning figure:

The black line that is shooting off to the top of the chart is the total number of therapies or publications related to covid-19, as measured against the number of days since the beginning of the pandemic/epidemic. The various dashed lines correspond to the number of therapies and publications for other diseases and/or pandemics (Ebola, Zika, H1N1, and breast cancer). Two things are immediately apparent.

First, covid-19 research is much higher than research related to other pandemic diseases. Second, the gap between covid-19 and other diseases has widened as the magnitude of the covid-19 pandemic becomes clearer. It seems obvious these differences are entirely driven by the difference in demand for a covid-19 therapy, both relative to other drugs and over time, rather than some scientific breakthrough that made it suddenly easier to do covid-19 research. So the above figure is strong evidence that pharma companies respond to profit opportunities and would probably respond further if the government promised to buy a working vaccine at a higher price than the market would normally support.

But dig into the data a bit deeper and there is something troubling. While a vaccine would be the most useful therapy, an unusually large share of the therapies under development are drugs, rather than vaccines. And while it would be nice if an existing drug turned out to be a useful therapy for covid-19, it seems more likely a new disease will require a new kind of drug. But repurposed drugs, rather than novel therapies account for an unusually large share of trials.

This difference has grown over time, as the scope of the pandemic widened. And the divergence between vaccines vs. drugs, and novel drugs vs. repurposed ones, is significantly larger for covid-19 than for Ebola, Zika, and H1N1.

This suggests the rising profitability of a covid-19 treatment is pushing ever more firms to focus on therapies that are not necessarily the best treatment for the disease, but which are most likely to get to the market soon. Vaccines tend to be harder than drugs, and novel drugs tend to be harder than repurposing existing drugs.

It turns out the above evidence is quite consistent with existing research on how medical research responds to market demand. We have good evidence that government promises to pay more for vaccines would likely induce more vaccine research. But the evidence we have also suggests such a policy is most effective at bringing to market a vaccine that does not require much more R&D (but read the ending of this newsletter for caveats).

Markets for Vaccines

The kind of program Athey, Kremer, Snyder, and Tabarrok are proposing is called an Advance Market Commitment, and it’s been successfully tried before. In 2007, a coalition of governments and the Gates Foundation pledged $1.5bn towards the production of 200 million annual doses of a pneumococcal conjugate vaccine for developing countries. If a manufacturer would supply the vaccine at a price of no more than $3.50 per dose, the advance market commitment would top up the rest with a share of the $1.5bn pledged. The program launched in 2009 and in 2010 GSK and Pfizer each committed to supply 30 million doses annually. This amount was increased over time, and a third supplier entered in 2019. Annual distribution exceeded 160 million doses annually by 2016.

Uptake of the pneumococcus vaccine was much faster than uptake for vaccines for a different virus without an advance market commitment (rotavirus). So the advance market commitment seems to have worked.

But there's an important caveat to all this: very little R&D was required to develop the pneumococcal conjugate vaccine. When it was selected, vaccines for similar diseases in developed countries already existed, and vaccines covering the strains in developing countries were already in late-stage clinical trials. So in this case, the advance market commitment pushed firms to quickly build up manufacturing and distribution capacity, but it didn’t push them to do extensive R&D since none was needed.

This is the only time a large-scale advance market commitment has been tried. But that’s not the only place we can look for evidence.

Finkelstein (2004) identifies three US policy changes that increased the profitability of vaccines for some diseases but not others. She then looks to see if firms respond by creating more new vaccines for the affected diseases, relative to the unaffected diseases. Indeed, they do. Let's dig in a bit more.

The three policies Finkelstein uses are (1) the 1991 CDC recommendation that all infants be vaccinated against Hepatitis B; (2) the 1993 decision for Medicare to fully cover the cost of influenza vaccination for Medicare recipients and; (3) the 1986 creation of the Vaccine Injury Compensation Fund which indemnified vaccine manufacturers from lawsuits relating to adverse effects for some specified vaccines. In each of these three cases, policy choices made vaccines for some diseases more profitable, but had no effect on other diseases.

As a control group, Finkelstein considers various sets of alternative diseases that were not affected by these policies, but which otherwise share some of the same characteristics as the affected diseases. All told, she has data on preclinical trials, clinical trials, and vaccine approvals for 6 affected diseases and control groups consisting of 7-26 other diseases, over 1983-1999.

Diseases where policy increased profitability saw an additional 1.2 clinical trials per year and an additional 0.3 new approved vaccines per year (but only 7 years after the policy took effect), as compared to controls. So the promise of more profit did pull in more vaccine development.

But the effect only travels so far up the research stream. When Finkelstein looks farther up the development pipeline, the effect disappears. Affected diseases had no more preclinical trials than the control group. This suggests firms responded to the increased profit by pulling vaccines already far along off the shelf and putting them into clinical trials. But if it stimulated more basic research, the effect was too small to be detected.

Markets for Drugs

There is also a rich vein of research on the extent to which general pharma R&D (not vaccines) respond to changes in the size of the market for different health products. Dubois, Mouson, Scott-Morton, and Seabright (2015) look at the link between potential profits and innovation in the context of global pharmaceutical innovation. They've got a data on drug sales in 14 major countries, which they use to make estimates of the size of the market for different categories of therapeutic medicine. Their goal is to see how changes in the size of the market for a drug change the propensity to develop new drugs for the market. In this case, they're holding the measure of innovation to a relatively high bar: a newly approved drug, marketed in one of their 14 countries, that is also a new chemical entity (i.e., not a modification of an existing drug).

One challenge is that better drugs can, themselves, change the size of the market. Suppose for example, that new drugs just come along randomly as a result of serendipity. In that case, potential profit doesn't actually induce firms to develop new drugs. But if these new drugs find a market, and we're measuring the size of the market by looking at spending on drugs, then we'll create a misleading correlation between the "size" of the market and the number of new drugs. In this case, the number of drugs is "causing" the size of the market, rather than vice-versa. To avoid this, they use a statistical technique (instrumental variables) to pull out the parts of demand that vary due to demographics and overall GDP growth (neither of which should be affected by drug innovation over the 11-year period they work with).

When they do this, they find that bigger markets do indeed lead to more drugs. On average, when the market for a therapeutic category grows by 10%, there are 2.6% more new chemical entities approved over a given time period.

But how scientifically novel are these new drugs? Suggestive evidence comes from Acemoglu and Linn (2004), who perform a similar exercise as Dubois, Mouson, Scott-Morton and Seabright (2015), but on US rather than global sales data. When the market for different diseases in the US changes due to shifting demographics, how does this change the flow of new drug approvals for those diseases? Acemoglu and Linn find the effect of a bigger market is much, much stronger for generic drugs than for new molecular entities.

More direct evidence comes from Dranove, Garthwaite, and Hermosilla (2020) who also investigates this question in the context of global drug development over 1997-2018. They use the US Medicare Part D extension to see if the promise of higher profits leads led firms to pursue more scientifically novel drugs.

The basic idea is that Medicare Part D extended medicare to pay for enrollee's pharmaceutical drugs beginning in 2006. This created a big new market for drugs used by Medicare enrollees (US residents aged 65 and up). Dranove, Garthwaite, and Hermosilla have data on worldwide pharmaceutical company drug trials, and they want to see if companies run more trials on scientifically novel drugs in response to the new opportunities created by Medicare part D.

To measure the scientific novelty of a drug, Dranove, Garthwaite, and Hermosilla count the number of times the specific "target-based action" of the drug has been explored in previous drug trials (of similar or stronger intensity). A target based action comprises the specific (targeted) biological entity and the mechanism used to modify its function: for example, a p38 MAP kinase inhibitor is a target-based action that targets the p38 mitogen-activated protein kinases and inhibits its function. If this target-based action has never before been used in a clinical trial, then a drug using it is considered maximally novel. The more often it has been previously used, the less novel.

With this measure in hand and data on 76,161 clinical trials on 36,002 molecules, Dranove, Garthwaite, and Hermosilla look to see if therapeutic areas with greater profit potential in the wake of Medicare Part D see more clinical trials for scientifically novel drugs. While they do find that more exposed therapeutic areas do see a small increase in trials for the most novel kinds of drugs, once again the effect is much stronger for the least novel drugs. Over 2012-2018 the number of trials for the least novel group of drugs increased 106%, while the number of trials for the most novel group increased just 14% (with most of the gains coming in the second half of that period).

Can We Buy a Covid-19 Vaccine?

So back to covid-19. Can an advance market commitment “buy” a vaccine that doesn’t yet exist? Or are we in the same position as Joel Mokyr’s medieval kings, whose wealth can’t buy any treatment for the bubonic plague until someone thinks of the germ theory of diseases?

First, the studies above suggest these policies do work, but are most effective if the vaccine does not require too much more R&D. Does a covid-19 vaccine require a lot more research? I don’t know. On the one hand, there hasn’t been a human vaccine for this class of virus before. On the other hand, there have been vaccines for veterinary applications (innovation in human and animal health has a lot of similarities), and there seems to be no shortage of options.

Second, the size of the proposed policy is enormous relative to what’s been tried before. So even if these policies normally only work weakly on vaccines that are far from approval, it may be that we still observe a large effect simply because we’re pouring so much money into it.

Third, one of the goals of the Athey, Kremer, Snyder, Tabarrok proposal is explicitly to build manufacturing capacity for vaccines before they are proven, so that we can mass produce them as soon as we find one that works. To the extent building capacity is not a problem that requires R&D, than advance market commitments should work very well. In general, an advance market commitment is only one (albeit big) part of a set of complementary incentives the authors recommend to push and pull a vaccine to market. Give the whole thing a read!

Subscribe at mattsclancy.substack.com