Opinion piece by Marcus A. Davis, CEO of Rethink Priorities
TL;DR:
If it’s possible to be famous for your spreadsheets, GiveWell, the charity evaluator, is. To decide whether malaria nets do more good per dollar than deworming pills, they build models with dozens of inputs, such as how many nets are actually used, how common malaria is in each region, and how much a child’s health today affects their earnings decades later.
It’s a heroic effort, built and refined over the last 19 years, that tries to pin down all the uncertainties they have about this complex question.
Now, every so often, someone looks at this complex model and says something like, “Why bother? Malaria kills children. Nets prevent malaria. Worms are bad, but they usually don’t kill you. Nets win. I worked that out in ten seconds from a high-level argument without opening a spreadsheet.”
The problem is that person isn’t truly skipping the spreadsheet. They’ve built one in their head. They just aren’t showing their work, even to themselves.
A model is really any story about how an action leads to outcomes. No one opens a spreadsheet to decide when to leave for work. But you have a model: typical traffic, how much lateness you can tolerate, whether it’s raining, and so on. Many people, unfortunately, find out that this is a model when they are an hour late to work because a bridge was closed.
The same issue applies to high-level stories people tell themselves about why to give to one area rather than another or how to split across them. The person assuming “bednets win” has built a model in their head that depends on:
Each of these is a number. Change a couple, and the answer can flip. GiveWell’s spreadsheet isn’t more complicated structurally than the reasoning underlying the ten-second model. It’s equally complicated, just visible and inspectable.
The lessons here are:
Some skeptics go further. They don’t just think models are unnecessary; they think models are misleading, and that experienced judgment is the safer guide. They raise some fair points:
All of this is true. Now ask yourself whether the alternative of going with your gut fixes any of these issues. Spoiler: it does not.
People worry that explicit models leave out factors or produce results that are too precise. This does happen. But why do you think your unaided mind has captured all the relevant factors in the correct amounts and to the correct level of precision? Our ten-second model critic almost certainly didn’t have region-specific estimates of disease burden, or a clear sense of how much extra future income would be worth as much as saving a life.
But not thinking about something is the same as building a model and counting it as having no effect. And zero is the most confident estimate there is. Leaving something out isn’t humble, it’s maximally sure.
Indeed, an ignored effect gets an estimate of zero, with no error bars. If precision is the concern, the informal approach is the most overconfident.
In more formal models, the alternative to skipping inputs like this is to directly flag where you think the uncertainty is, and test multiple inputs, rather than use a made-up precise number. A spreadsheet can write “unknown, could be big” or “we ran a sensitivity analysis, and this had no effect on our conclusion.” A gut feeling that never noticed the question can’t.
You may think I’m being silly. Of course, you think, no one can do GiveWell’s spreadsheets in their head. Don’t be ridiculous. But cause prioritization is different, you might say. You can compare bednets to AI risk interventions without building a formal model. This may sound intuitive, but doing that in your head is harder, not easier.
The more different two interventions are, the more theory- and philosophy-driven your outcomes may be. For example, here’s a line of reasoning that many thoughtful people hold: AI is the most important thing happening in the world right now. If it goes badly, almost nothing else matters. So AI safety is where marginal money does the most good.
Notice what this argument quietly depends on:
People who hold this view have usually thought hard about some of these. The question isn’t whether you’ve considered them. It’s whether you’ve considered all of them, with actual numbers, and how they combine. That last part is where reasoning in your head is least reliable. It’s not at all obvious that the best thing to do is to act on what appears to be the highest-expected-value bet based on a high-level argument when you can’t be confident you are pushing in the right direction. It’s reasonable to give at least some weight to that view, but how much, specifically? And what do you do if you also have views that are more risk-averse, which say you absolutely should not spend money on projects like this?
One option in the case where you hold competing views is “I’ll evenly split across a few causes I care about.” This sounds like the humble option. But an even split is a strong claim: that each cause is roughly equally worth funding at the margin, and that every option on the list does good rather than harm. It’s a model with the weights set to “equal” because nobody wrote down anything else.
Even simple cases of weighing multiple views are really complicated. Consider a simple case where you have $10 million to spend on philanthropic causes. You’re about 80% confident in a view on which AI safety is far and away the best use of money. You give about 20% to a more cautious view that proven global health programs are best and speculative AI work might even be worse than doing nothing. In this circumstance, some people think, “I’ll just adjust a little,” and assume they’ve reached the right answer. Unfortunately, that move is almost comically underspecified.
“I’ll adjust a bit for the 20%” could mean any of these:
These aren’t small variations. These are four different philanthropic strategies. “I’ll adjust a bit” doesn’t tell you which one of these you picked. Worse, there are better and worse reasons to pursue each of these moves in different circumstances; they lead to very different practical results, and it’s hard to do (3) or (4) in a consistent way without a formal model.
(1) just ignores your uncertainties. (2) splits resources. (3) and (4) could result in whatever you want based on how you view the “riskiest” choices and some bespoke informal negotiation in your head.
There’s just no reason to think it’s better to do this type of analysis in your head rather than under a formal structure. And this is the easy case, when you are comparing two theories rather than three or four, and when you have a theory you’re pretty sure is correct. There’s actually an entire philosophical discipline dedicated to such scenarios. If it was obvious what to do, the field wouldn’t exist.
Perhaps the most serious objection to cross-cause models isn’t “models are dumb.” It’s this: the important inputs are basically guesses. How much does a chicken’s suffering count compared to a person’s? What’s the chance a given AI safety grant reduces catastrophe? Nobody knows. Put guesses in, you get guesses out.
That last part is true. But your gut, even supported by most rational thought processes, still relies on guesses. It just doesn’t label them. When you decide AI safety beats animal welfare, you’ve implicitly answered both questions above. You’ve just done it without noticing, and without checking whether a different, equally reasonable guess would change your mind.
A model can do something your gut can’t: show you which guesses actually matter. Often, you’ll find that the answer barely changes across a wide range of reasonable inputs, so you can stop worrying about those. Sometimes you’ll find everything hinges on one number nobody knows. That’s not the model failing. That’s the model telling you exactly where your uncertainty lives.
A related fear is that the model will say to put everything into whatever has a tiny chance of an astronomical payoff. Naively averaging outcomes can do that. But a model doesn’t have to work that way. You can build in caution about long shots, and many models do. The difference is that a model applies that caution consistently and visibly, rather than whenever the long shot happens to feel uncomfortable. Indeed, the model underlying RP’s Cross-Cause Fund allows just this sort of transparent flexibility to incorporate weight to both long shots and caution about such bets.
Simple rules of thumb can beat complicated models, but mostly when you get quick feedback. A doctor’s checklist for heart attacks is effective because it can quickly determine whether a patient has had one.
Charity decisions don’t work like this. You rarely, if ever, learn if you chose right, the options are hard to compare, and many factors interact.
Yet this is exactly backward to how it feels. The more complicated a decision is, the more tempting it can be to give up on careful reasoning and go with your gut. But complicated decisions are where your gut is least reliable and where it’s hardest to notice you’re wrong. Complexity is the reason to show your work, not an excuse to skip it.
Some people might say they will use a model “as one input among many.” Fair enough, this is a much better strategy than not using one, but be mindful of this pattern: the model gets followed when it confirms what you thought, and overruled by “other considerations” when it surprises you. At that point, the model isn’t correcting your mistakes. Your existing beliefs are “correcting” the model.
If the model is wrong, say what’s wrong and fix it. Questioning, updating, and improving a model should be ingrained in its design. Turning down the volume knob on the explicit model without a reason is just trusting a different model, the one in your head, that nobody can see and no one, even you, is properly examining.
Building models is expensive. If this is true for GiveWell, it’s definitely true for cross-cause prioritization. For small decisions, it isn’t worth it. If you are giving a few thousand dollars, building a careful model is probably not worth your time. Give through a fund or evaluator like GiveWell or Animal Charity Evaluators and never worry about it again.
If you’re a foundation directing $100 million, spending $1 million on analysis pays for itself if it makes your choices just 1% better, and the gap between a reasoned choice under uncertainty and just using your gut can be much more than that.
If you’re somewhere in between, say directing somewhere between $20K and $20M, then the stakes are too high to wing it, but doing a seven-figure analysis of your own doesn’t make sense either. That’s exactly the case for pooled analysis: someone does the work once, and many donors benefit.
Sadly, there are very few careful, all-things-considered models comparing different kinds of causes, despite billions of dollars being spent. As a result, some of the largest decisions in charity are made on a judgment call that nobody writes down, often with no effort at all to justify them rigorously.
Sometimes a careful model concludes the uncertainty is huge. That can be frustrating, but it can still be useful. It might tell you to spread your bets, look for options that are robust under many assumptions, or spend resources to learn more before committing.
A confident gut call under these circumstances is worse than an honest “we’re not sure.”
One upside of so few people trying to build rigorous models for giving across areas is that this means the payoff from such models is likely unusually high. When almost nobody has even bothered to do the analysis, even a very imperfect first attempt can improve decision-making. It takes a lot of effort, but we’ve done it for you. This is why we built the Cross-Cause Fund.
It tries to capture all the relevant variables we’ve discussed on the page, rather than in your head. That includes cost-effectiveness, the value of a marginal dollar, the type of benefit, beneficiary, risk, moral views, and how to combine views when they disagree.
Whether you give through us or not, here’s what to ask of anyone recommending where your money should go, including yourself:
The ten-second critic and GiveWell aren’t doing different things. They’re both answering the same question with the same kinds of assumptions. One of them can be checked, argued with, and improved. The other can only be believed or not.
If you care about getting the answer right, and not just feeling sure of it, that settles which one to use. Intuition is a fine place to start; it’s a bad place to stop. The more that’s at stake, the less excuse there is for not showing your work.
None of this means models are always right, or that a model should overrule “common sense” in all circumstances. It means the default should flip. Right now, careful models are treated as the ones that have to justify themselves, and gut judgment gets a pass. It should be the other way around. If you think your intuition beats the written-down version of a particular case, you should be able to say why. And if you can say why, you can and should write that down.
You don’t need to build a model yourself. But you should be able to see one: the assumptions that matter, and what got left out. If a recommendation can’t show you that, including the one in your head, it isn’t more trustworthy just for being simpler. It’s just harder to check.
To dive deeper into the topic of modeling in cross-cause giving, check our articles on:
And to learn more about our cross-cause prioritization work, follow our RP Research Digest Substack.
This post was written by Marcus A. Davis with support from Jim Buhler. Thank you to Hayley Clatterbuck and Benjamin Tereick for feedback, and Urszula Zarosa for feedback and editing.
Parts of this content were prepared with the assistance of AI tools (such as Claude and Gemini), which we use to improve efficiency and readability. All outputs are supervised, reviewed, and fact-checked by RP staff, who remain responsible for the final content.