I've done nothing to test these heuristics and have no empirical evidence for how well they work for forecasting replications or anything else. Iâm going to write them anyway. The heuristics Iâm listing are roughly in order of how important I think they are. My training is as an economist (although I have substantial exposure to political science) and lots of this is going to be written from an econometrics perspective.
How much does the result rely on experimental evidence vs causal inference from observational evidence?Â
I basically believe without question every result that mainstream chemists and condensed matter physicists say is true. I think a big part of this is that in these fields itâs really easy to experimentally test hypotheses, and to really precisely test differences in hypotheses experimentally. This seems great.Â
On the other hand, when relying on observational evidence to get reliable causal inference you have to control for confounders while not controlling for colliders. This is really hard! It generally requires finding a natural experiment that introduces randomisation or having very good reason to think that youâve controlled for all confounders.Â
We also make quite big updates on which methods effectively do this. For instance, until last year we thought that two-way fixed effects did a pretty good job of this before we realised that actually heterogeneous treatment effects are a really big deal for two-way fixed effects estimators.Â
Whatâs more, in areas that use primarily observational data thereâs a really big gap between fields in how often papers even try to use causal inference methods and how hard they work to show that their identifying assumptions hold. I generally think that modern microeconomics papers are the best on this and nutrition science the worst.Â
Iâm slightly oversimplifying by using a strict division between experimental and observational data. All data is observational and what matters is how credibly you think youâve observed what would happen counterfactually without some change. But in practice, this is much easier in settings where we think that we can change the thing weâre interested in without other things changing.Â
There are some difficult questions around scientific realism here that Iâm going to ignore because Iâm mostly interested in how much we can trust a result in typical use cases. The notable area where I think this actually bites is thinking about the implications of basic physics for longtermism where it does seem like basic physics actually changes quite a lot over time with important implications for questions like how large we expect the future to be.Â
Are there practitioners using this result and how strong is the selection pressure on the resultÂ
If a result is being used a lot and there would be easily noticeable and punishable consequences if the result was wrong, Iâm way more likely to believe that the result is at least roughly right if itâs relied on a lot.
For instance, this means Iâm actually really confident that important results in auction design hold. Auction design is used all the time by both government and private sector actors in ways that earn these actors billions of dollars and, in the private sector case at least, are iterated on regularly.Â
Auction theory is an interesting case because it comes out of pretty abstract microeconomic theory and wasnât developed really based on laboratory experiments, but Iâm still pretty confident in it because of how widely itâs used by practitioners and is subject to strong selection pressure.Â
On the other hand, Iâm much less confident in lots of political science research. It seems like places like hedge funds donât use it that much to predict market outcomes, it doesnât seem to be used by governments that much, and itâs really hard to know how counterfactually important, say, World Bank programs that use political science were. Â
How large is the literature that supports the result and how many techniques have been used to supportÂ
This view actually does have some empirical support. Thereâs this nice paper where a load of different researchers are given the same (I think simulated) data and looked at how researchers result. They found that there was quite a lot of difference between what researchers found based on things like their coding choices and what statistical techniques they used, but that when there was a real effect the average paper found an effect of right sign and roughly right magnitude, and when there was no real effect the average researcher found roughly no effect. Iâm afraid I canât find either the paper and I canât be bothered to link to the Noah Smith or Matt Clancy blog posts on it.Â
Mostly though I use this heuristic because it seems pretty sensible.Â
External validity
External validity is how likely it is that a result generalises from whatever the study setting was to the setting in which the result is used.Â
I think this is a really big deal for lots of RCT-based development economics. We just see, really quite often, that results that seem to consistently hold when tested with RCTs donât hold when scaled up.Â
Iâm more sceptical of the external validity of a result the more intensive the intervention is and so the more buy-in and effort is needed from participants and researchers. Seems pretty likely that when the intervention is used it wonât have as much effort put into it. Iâm particularly sceptical if the intervention is complex or precise.Â
Results given statistical powerÂ
Statistical power says how likely it is to see an effect size given the true effect size and the sample size. If the statistical power of a test is low but significant results are found, itâs likely that the researcher just got lucky and the true effect size is much smaller and/or the opposite sign.
The intuition for this is that if a statistical test is underpowered - say for this example the power is under 50% - then itâs unlikely that a statistically significant effect is found.Â
If a statistically significant effect is found then something weird must have happened, like the specific sample that was used stochastically having really large effect sizes. The intuition for this is that if you have a small sample size (so your estimate of population statistics has a high variance) and are very unwilling to accept mistakes in the direction of finding effects that arenât there, you need a really large mean effect size to be confident that thereâs any effect at all! This effect size has to be larger than the mean effect size because, by assumption, youâre test is unlikely to detect an effect given the true distribution of the variable in question - this is what it means for a test to have low power.Â
More sinisterly, it could also imply some selection effect for which results are observed, like publication bias or the methods the researchers used.Â
I want to caveat this section by saying that I donât have a very good intuition for power calculations and how much they actually affect how likely results are to replicate.Â
How strong is the social desirability bias at playÂ
This seems somewhat important, but I think is often overplayed in the EA and rationality communities. But it does in practice mean that I think Iâm less likely to see papers that find, say, that child poverty has no effect on future outcomes. My vibe is that psychology seems particularly bad for this for some reason?Â
But also I see papers that find socially undesirable results all the time!Â
For instance, this paper finds negative effects of democracy on state capacity for places with middling levels of democracy, this paper finds higher levels of interest in reading amongst preschool-age girls, and this paper finds no association between youth unemployment and crime. Itâs really easy to find these papers! You just search for them on Google Scholar.Â
Have there been formal tests of publication biasÂ
We can test whether the distribution of results on a specific question looks like it should if publication was independent of the sign and magnitude of their results. Iâm a lot less confident in a field if it consistently finds publication bias.
Thanks for this, a really nice write up. I like these heuristics, and will try to apply them.
On the intuition behind how to interpret statistical power, doesn't a bayesian perspective help here?
If someone was conducting a statistical test to decide between two possibilities, and you knew nothing about their results except: (i) their calculated statistical power was B (ii) the statistical significance threshold they adopted was p and (iii) that they ultimately reported a positive result using that threshold, then how should you update on that, without knowing any more details about their data?
I think not having access to the data or reported effect sizes actually simplifies things a lot, and the Bayes factor you should update your priors by is just B/p (prob of observing this outcome if an effect / prob of observing this outcome if no effect). So if the test had half the power, the update to your prior odds of an effect should be half as big?
Yeah, I think a Bayesian perspective is really helpful here and this reply seems right. Â