Excited to see this!
Really sorry, but recently we released Squiggle 0.8, which added a few features, but took away some things (that we thought were kind of footguns) that you used. So the model is now broken, but can easily be fixed.
I fixed it here, with some changes to the plots.
2. Instead of Plot.fn(), it's now Plot.distFn() or Plot.numericFn(), depending on if its returning a number or distribution. Note that these are now more powerful - you can adjust the axes scales, including adding custom symlog scales. I added a symlog yScale.
Also, in the newer version, you can specify function ranges. Like,
expected_number_of_scandals = {|n: [1e-3, 30k]|(1 / 5000 to 1 / 500) * n}You have the code,
growth_benefits = {|n: range|n * log(n)}This is invalid where n=0, so I changed the range to start slightly above that. (You can also imagine other ways of doing that)
In the future, we really only plan to have breaking changes in main version numbers, and we'll watch them on Squiggle Hub. I didn't see that there were recent active Squiggle models. Sorry for the confusion, again!
Also, do feel free to post the model on Squiggle Hub. I haven't formally announced it here yet, but people are welcome to begin using it.
Thanks for writing this! This is a cool model.
Our best guess is that benefits grow slightly superlinearly because of coordination benefits (but you can easily remove coordination benefits from the model).
- A naïve first-order approximation is that benefits (not accounting for reputational issues) are linear in the size of the group.
- If everyone in EA donated a constant amount of money, then getting more people into EA would linearly increase the amount of money being donated (which, for simplicity, we can say is a linear increase in impact)
Is linear a good approximation here? Conventional wisdom suggests decreasing marginal returns to additional funding and people, because we'll try to prioritize the best opportunities.
I can see this being tricky, though. Of course doubling the community size all at once would hit capacities for hiring, management, similarly good projects, and room for more funding generally, but EA community growth isn't usually abrupt like this (FTX funding aside).
In the animal space, I could imagine that doing a lot of corporate chicken (hen and broiler) welfare work first is/was important for potentially much bigger wins like:
But I also imagine that marginal corporate campaigns are less cost-effective when considering only the effects on the targeted companies and animals they use, because of the targets are prioritized and resources spent on a given campaign will have decreasing marginal returns in expectation.
GiveWell charities tend to have a lot of room for funding at given cost-effectiveness bars, so linear is probably close enough, unless it's easy to get more billionaires.
For research, the most promising projects will tend to be prioritized first, too, but with more funding and a more established reputation, you can attract people who are better fits, and can do those projects better, do projects you couldn't do without them, or identify better projects, and possibly managers who can handle more reports.
Maybe there's some good writing on this topic elsewhere?
My impression is that, while corporate spinoffs are common, mergers are also common, and it seems fairly normal for investors to believe that corporations substantially larger than EA are more valuable as a single entity than as independent pieces, giving some evidence for superlinear returns.
But my guess is that this is extremely contingent on specific facts about how the corporation is structured, and it's unclear to me whether EA has this kind of structure. I too would be interested in research on when you can expect increasing versus decreasing marginal returns.
Thanks for the analysis! I think it makes sense to me, but I'm wondering if you've missed an important parameter: diminishing returns to resources.
If there are 100 community members they can take the 100 most impactful opportunities (e.g. writing DGB, publicising that AI safety is even a thing), while if there are 1000 people, they will need to expand into opportunities 101-1000, which will probably be lower impact than the first 100 (e.g. becoming the 50th person working on AI safety).
I'd guess a 10x increase to labour or funding working on EA things (even setting aside coordination and reputation issues) only increases impact by ~3x.
It seems like that might make significant difference to the model - if I've understood, currently the impact of marginal members in the model is actually increasing due coordination benefits, whereas this could mean it's decreasing.
I'd still guess marginal growth is net positive, but I feel less confident than the post suggests.
Aside: A more compelling argument against growth in this area to me is something like "EA should focus on improving its brand and comms skills, and on making reforms & changing its messaging to significantly reduce the chance of something like FTX happening again, before trying to grow aggressively again"; rather than "the possibility of scandals means it should never grow".
Another one is "it's even more high priority to grow X others movements than EA" rather than "EA is net negative to grow".
Less importantly, I also feel less confident coordination benefits would mean impact per member goes up with the number of members.
I understand that the value of a social network like Facebook grows with the number of members. But many forms of coordination become much harder with the number of members.
As an analogy, it's significantly easier for 2 people to decide where to go to dinner than for 3 people to decide. And 10 people in a group discussion can take ages to come to consensus.
Or, it's much harder to get a new policy adopted in an organisation of 100 than an organisation of 10, because there are more stakeholders to consult and compromise with, and then more people to train in the new policy etc. And large organisations are generally way more bureaucratic than smaller ones.
I think these analogies might be closer than the analogy of Facebook.
You also get effects like in a movement of under 1000, it's possible to have met in person most of the people, and know many of them well; while in a movement of 10,000, coordination has to be based on institutional mechanisms, which tend to involve a lot of overhead and not be as good.
Overall it seems to me that movement growth means more resources and skills, more shared knowledge, infrastructure and brand effects, but also many ways that it becomes harder to work together, and the movement becoming less nimble. I feel unsure which effect wins, but I put a fair bit of credence on the term decreasing rather than increasing.
If it were decreasing, and you also add in diminishing returns, then impact per member could be going down quite fast.
Thanks! See my response to Michael for some thoughts on diminishing returns.
10x increase in labor leading to 3x increase in impact feels surprising to me. At least in the regime of ~2xing supply I doubt returns diminish that quickly. But I haven't thought about this deeply and I agree that there is some rate of diminishing marginal returns which would make marginal growth net negative.
The response to Michael is an interesting point, but it only concerns diminishing returns in individual capabilities of new members.
Diminishing returns are mainly driven by the quality of opportunities being used up, rather than the capabilities.
IIRC a 10x in resources to get a 3x in impact was a typical response in the old coordination forum survey responses.
In the past at 80k I'd often assume a 3x increase in inputs (e.g. advising calls) to get a 2x increase in outputs (impact-adjusted plan changes), and that seemed to be roughly consistent with the data (though the data don't tell us that much). In some cases, returns seem to diminish a lot faster than that. And you often face diminishing returns at several levels (e.g. 3x as much marketing to get 2x as many applicants to advising).
I agree returns are more linear in areas where EA resources are a small fraction of the total, like global health, but that's not the case in areas like AI safety, GCBRs, new causes like digital sentience, or promoting EA.
And even in global health, if GiveWell only had $100m to allocate, average cost-effectiveness would be a lot higher (maybe 3-10x higher?) than where the marginal dollar goes today. If GiveWell had to allocate $10bn, I'd guess returns would be at least several fold lower again on the marginal spending.
Great read, and interesting analysis. I like encountering models for complex systems (like community dynamics)!
One factor I don't think was discussed (maybe the gesture at possible inadequacy of encompasses this) is the duration of scandal effects. E.g. imagine some group claiming to be the Spanish Inquisition or the Mongol Horde, or the Illuminati tried to get stuff done. I think (assuming taken seriously) they'd encounter lingering reputational damage more than one year after the original scandals! Not sure how this models out; I'm not planning to dive into it, but I think this stands out to me as the 'next marginal fidelity gain' for a model like this.
Neat!
Summary: in what we think is a mostly reasonable model, the amount of impact a group has increases as the group gets larger, but so do the risks of reputational harm. Unless we believe that, as a group grows, the likelihood of scandals grows slowly (at most as quickly as a logarithmic function), this model implies that groups have an optimal size beyond which further growth is actively counterproductive — although this size is highly sensitive to uncertain parameters. Our best guesses for the model’s parameters suggest that it’s unlikely that EA has hit or passed this optimal size, so we reject this argument for limiting EA’s growth.[1] (And our prior, setting the model aside, is that growth for EA continues to be good.)
You can play with the model (insert parameters that you think are reasonable) here.
Epistemic status: reasonable-seeming but highly simplified model built by non-professionals. We expect that there are errors and missed considerations, and would be excited for comments pointing these out.
Here are two plots showing how the net impact (per year, with arbitrary units of impact) would change as a group grows — the plots are very different because the parameters are different and the model is very sensitive to that:
A more formal description of the model
Getting to the implications of the model
(Note: this section uses asymptotic (Big O) notation.)
Remember that we defined “scandal” as “wrongdoing that becomes prominent.” Given this, our best guess here is that frequency grows sublinearly with the size of the group.
Note also that we can try to account for the variation in the importance of real-world scandals when we’re setting parameters by saying that something less significant simply has a smaller chance of causing a scandal. In other words, if you think that someone will definitely cause 3 scandals in the following year, but they’re all very small, you can model this here by saying that this is actually 1 scandal in the way we’re defining scandal here. (Whereas something unusually significant might be equivalent to two scandals.)
The harms we would expect to accrue from scandals are things like:
Our guess is that reputational harm is best modeled as a percentage decrease in impact. This fits the first point above better than it fits the second, but even for 2 (a), harms might accrue in a similar pattern: the first scandal drives off the least interested x% of people, then the second scandal drives off the x% least interested of the remainder, etc. (See evaporative cooling.)
There are some costs which arguably do not fit this model. For example, the negative perception of early cryonicists may have deterred cryobiologists (who weren’t cryonicists) from doing cryonics-related research that they otherwise would have done independently. It seems plausible that from the point of view of the group, this is better modeled as an additive cost — flat negative impact — due to the reputational issues as opposed to a multiplier penalty on the positive impact of the members of the group. (Additionally, the “effectiveness penalty multiplier” model doesn’t allow for scandals to cause someone’s work to become negatively impactful, which doesn’t seem universally true.)
Another complication might be something like splintering; it’s possible that you can’t model group size as independent of scandal rate and reputational harm, because when scandals have certain effects, the group splinters into smaller groups or simply loses members.
Still, we think the percentage decrease model is the best that we have come up with.
We want to understand: given a fixed scandal, is it more harmful per person if the group the scandal is attached to is bigger? Our best guess is that per person in the group, harm per scandal decreases with the size of the group, but we’re modeling K as a constant for simplicity.
It seems like there are some counterbalancing factors:
Our best guess is that benefits grow slightly superlinearly because of coordination benefits (but you can easily remove coordination benefits from the model).
My (Ben) subjective experience of playing around with this model is that for reasonable parameter values, it seems pretty clear that groups of more than 500 people are better than smaller groups, but it's harder to get outputs that show that larger groups (or any reasonable size) are noticeably worse than smaller ones. I have to intentionally choose weird parameters to get a graph like the one above, where there is a clear peak and larger groups are worse – unless I intentionally do this, it usually seems like growth is neutral or good (although confidence intervals are often very wide). (Lizka agrees with this.)
When I try to think of scandals that plausibly decreased the effectiveness of people in EA by >5% the list feels pretty short: FTX is probably in there, but even disturbing news or incidents like the TIME article on sexual harassment seem unlikely to have caused one in 20 people to leave EA (or otherwise decreased effectiveness by >5%). And we have 10-20k person-years to have caused scandals (suggesting that the base rate of scandals per person per year is 1/20000 to 1/10000); plugging in those numbers here indicates that EA should grow vastly beyond its current size.
More importantly: when I try to argue backwards from the claim that EA is already too big, I have to put in numbers that seem absurd, like here.
So my guess is that if growth is bad, it's because this model is flawed (which, to be clear, is pretty likely, although the flaws might not necessarily point in the direction of making it more likely that growth is bad).
| Factor | Possible implications |
| Probably attracts people who are more conscientious and nice than average | Decreases per-member frequency of scandals |
| The desire to make things work (rather than just compete) encourages most participants to try to get along and resolve conflicts amicably | Decreases per-member frequency of scandals |
| Some / many of the things we do are broadly regarded as good (e.g. GiveWell) | Decreases per-member frequency of scandals |
| The group isn’t super defined / is pretty decentralized; it’s not one massive organization. So e.g. someone donating to GWWC or effective charities can continue to do that as much as they could before (except maybe they’re demotivated) if someone prominent in a big animal advocacy organization is involved in a scandal | Decreases cost of scandal |
| Could be seen as a nonprofit | Unclear; nonprofits are sometimes held to higher standards (e.g. around compensation) but also have some default assumption of goodwill |
| Members tend to be from privileged demographics | Unclear; makes EA more “punchable” but also members have larger safety nets and more resources to push back |
| Is identified as powerful and allied with powerful groups (by some) | Unclear; makes EA more “punchable” but also members have larger safety nets and more resources to push back |
| A large set of different organizations with different practices any of which might be objectionable to someone | Increases per-member frequency of scandals, decreases cost of scandal
|
| A large set of different geographic subcultures with different practices | Increases per-member frequency of scandals, decreases cost of scandals |
| Tells people to take big actions which can frequently go badly and are expected to backfire at some decent rate | Increases per-member frequency of scandals |
| A high level of overlap between people's professional and friend networks means that almost all aspects of someone's life can be regarded as relevant for criticism, rather than just what they do in the course of their work | Increases per-member frequency of scandals |
| The desire to make things go well gives people a reason to stick around even if dissatisfied, and feel a moral responsibility to fight other people if they think what they're doing is harmful | Increases per-member frequency of scandals |
| Nobody has the authority to impose universal rules | Increases per-member frequency of scandals; possibly decreases cost of scandals (because scandals are legitimately not caused by the group) |
| Nobody can control who identifies themselves with EA, at least for the purposes of a critical journalist | Increases per-member frequency and cost of scandals |
| A large fraction of our communication, especially by new and less professional folks, is public and able to be used against us indefinitely (c.f. corporations or government agencies) | Possibly increases per-member frequency of scandals |
| Is engaged in activity contrary to the views of some existing political alliances and interests, so has accumulated active and motivated haters (and also some motivated by the FTX association) | Increases per-member frequency of scandals |
| Disproportionately attracts young people who tend to be harder to screen, and behave more erratically, and develop new mental health problems at higher rates | Increases per-member frequency of scandals |
| Could be seen as a political movement trying to influence society, which makes it seem particularly fair game to attacks | Increases per-member frequency of scandals |
As with many such models, you can choose parameters to get basically any possible outcome. But the settings that seem most plausible to us result in growth being good.
One of the few takeaways from this exercise that can be said with confidence is that bigger groups are likely to have more scandals, so if EA grows, that’s something we should prepare for and mitigate against.
The original idea for a related model was developed by a person who wishes to remain anonymous. Ben and Lizka made this more nuanced and wrote this post as well as the squiggle code. The resulting model is rough and doesn’t have fully conclusive results, but we thought it was worth sharing.
Though reputation is not the only relevant consideration for thinking about whether it would be better for EA to be small.
Or stigmatized or unpopular behavior
We are not actually sure about this. See the linked section.
Note that “prominence” here is complicated. Arguably, the thing that matters is the prominence of someone as a member of the group. For instance, if a really famous actor happens to shop at Walmart and is involved in a widely covered scandal, it probably won’t affect Walmart’s reputation. However, if the person was also a spokesperson for Walmart, it probably would, at least a bit.
Moreover, prominence in the group might make someone’s wrongdoing newsworthy (and via news coverage, a scandal) even if they weren’t prominent outside of the group before that happened. (Imagine a relatively unknown spokesperson for Walmart committing wrongdoing.) I’m not sure how much this actually happens.
It probably also matters whether the wrongdoing in question was somehow related to the group; e.g. the group already has a reputation for something related, or the wrongdoing highlights hypocrisy from the group’s perspective, etc.
I think both exponential and quadratic are too fast, although it's still plausibly superlinear. You used N∗log(N), which seems more reasonable.
Exponential seems pretty crazy (btw, that link is broken; looks like you double-pasted it). Surely we don't have the number of (impactful) subgroups growing this quickly.
Quadratic also seems unlikely. The number of people or things a person can and is willing to interact with (much) is capped, and the average EA will try to prioritize somewhat. So, when at their limit and unwilling to increase their limit, the marginal value is what they got out of the marginal stuff minus the value of their additional attention on what they would have attended to otherwise.
As an example, consider the case of hiring. Suppose you're looking to fill exactly one position. Unless the marginal applicant is better than the average in expectation, you should expect decreasing marginal returns to increasing your applicant pool size. If you're looking to hire someone with some set of qualities (passing some thresholds, say), with the extra applicant as likely to have them as the average applicant, with independent probability p and N applicants, then the probability of finding someone with those qualities is 1−(1−p)N, which is bounded above by 1 and so grows even more slowly than log(N) for large enough N. Of course, the quality of your hire could also increase with a larger pool, so you could instead model this with the expected value of the maximum of iid random variables. The expected value of the max of bounded random variables, will also be bounded above by the max of each. The expected value of the max of iid uniform random variables over [0,1] is NN+1 (source), so pretty close to constant. For the normal distribution, it's roughly proportional to √log(N) (source).
It should be similar for connections and posts, if you're limiting the number of people/posts you substantially interact with and don't increase that limit with the size of the community.
Furthermore, I expect the marginal post to be worse than the average, because people prioritize what they write. Also, I think some EA Forum users have had the impression that the quality of the posts and discussion has decreased as the number of active EA Forum members has increased. This could mean the value of the EA Forum for the average user decreases with the size of the community.
Similarly, extra community members from marginal outreach work could be decreasingly dedicated to EA work (potentially causing value drift and making things worse for the average EA, and at the extreme, grifters and bad actors) or generally lower priority targets for outreach on the basis of their expected contributions or the costs to bring them in.
Brand recognition or reputation could be a reason to expect the extra applicants to EA jobs to be better than the average ones, though.
Is growing the EA community a good way to increase useful brand recognition? The EA brand seems less important than the brands of specific organizations if you're trying to do things like influence policy or attract talent.
Thanks Michael! This is a great comment. (And I fixed the link, thanks for noting that.)
My anecdotal experience with hiring is that you are right asymptotically, but not practically. E.g. if you want to hire for some skill that only one in 10,000 people have, you get approximately linear returns to growth for the size of community that EA is considering:
And you can get to very low probabilities easily: most jobs are looking for candidates with a combination of: a somewhat rare skill, willingness to work in an unusual cause area, willingness to work in a specific geographic location, etc. and multiplying these all together gets small quickly.
It does feel intuitively right that there are diminishing returns to scale here though.
I would guess that for the biggest EA causes (other than EA meta/community), you can often hire people who aren't part of the EA community. For animal welfare, there's a much larger animal advocacy movement and far more veg*ns, although probably harder to find people to work on invertebrate welfare and maybe few economists. For technical AI safety, there are many ML, CS (and math) PhDs, although the most promising ones may not be cheap. Global health and biorisk are not unusual causes at all. Invertebrate welfare is pretty unusual, though.
However, for more senior/management roles, you'd want some value alignment to ensure they prioritize well and avoid causing harm (e.g. significantly advancing AI capabilities).