My concern is not primarily that current biosecurity work fails to cite earlier work on civilizational refuges. It is that there appears to have been a significant change in strategy without a corresponding published analysis or reassessment that explains the shift. The change in strategy seems to be something along:
Earlier: refuge-style protection of small populations and broader societal defenses both appeared in the biosecurity/resilience literature, without an obvious published analysis establishing the optimal scale of protection.
Now: the center of gravity appears much more strongly oriented toward protecting critical workers and maintaining functioning systems at large scale.
I may simply be missing some already published writing on this shift. If so, I would like to find it.
Until fairly recently, civilizational refuges were treated as a serious part of EA biosecurity and resilience work.
The basic idea was that a catastrophe might kill most people, while a deliberately protected population could survive and eventually rebuild civilization. This led to work on refuge design, minimum viable populations, loss of industrial capacity, and the probability of civilizational recovery.
As recently as 2022, refuges were discussed alongside PPE, detection, sterilization and medical countermeasures as parts of biosecurity infrastructure.
The Four Pillars approach instead aims to protect society at a much larger scale: preventing infection at scale, preserving a majority of critical workers, and maintaining enough scientific and industrial capacity to respond directly.
That is a very different strategy. From a longtermist (I am writing this as an Effective Altruist) perspective, however, the value difference is not simply proportional to the number of people protected. If a small surviving population ultimately produced a similarly capable and valuable civilization, it could in principle preserve much of the same long-run potential as the survival of several million people. But that is a strong conditional: a larger surviving population may substantially increase both the probability of recovery and the quality of the civilization that emerges, by preserving institutions, tacit knowledge, infrastructure, political diversity, and functioning economic systems. Previous work on collapse and recovery emphasizes that these different dimensions of “recovery” should not be conflated. At the same time, the amount of human population needed to preserve those capabilities may itself change with technology. Increasingly capable automation could reduce the amount of scarce human labor required for scientific, engineering, administrative, and eventually physical work. I return to this below.
Very roughly, a larger protected population should increase the probability of survival and recovery conditional on the protection strategy working. But that is not enough to determine which intervention is better, or whether both approaches should be pursued in parallel (a portfolio approach). The relevant quantity is closer to the marginal increase in the probability of a valuable long-run future, adjusted for the probability that the intervention can actually be implemented successfully, per unit of cost.
Schematically:
Cost-effectiveness ≈ P(intervention succeeds) × ΔP(valuable long-run future | intervention succeeds) / cost
This matters because a strategy protecting all critical workers could have much better recovery prospects conditional on success while still being similarly or even less cost-effective than a much smaller protected population. If, for example, expanding the protected population increases the probability of eventual recovery tenfold but increases cost twentyfold, the smaller intervention could produce more expected reduction in existential risk per dollar.
The comparison therefore cannot stop at “more surviving people are better.” It depends on how quickly recovery probability rises as the protected footprint expands, how tractable each scale is to deploy before and during a catastrophe, and how costs rise with scale. The optimal intervention might be a small refuge, broad critical-worker protection, something in between, or a portfolio of several scales.
These approaches are also not mutually exclusive. “Refuge” versus “critical-worker protection” risks treating what may be a continuum as a binary choice. One can protect a very small number of people to an extremely high standard while simultaneously deploying cheaper, less complete protection much more broadly. The relevant design space therefore varies both the scale of the protected population and the depth of protection.
This is close to the longer-run vision articulated by Shulman (2020), who envisages a civilization increasingly immune to catastrophic biological threats through ubiquitous detection, physical barriers, sterilization and automation, and explicitly asks whether rich societies could eventually operate at safety standards inspired by BSL-4. PPE for critical workers is not itself “BSL-4 for society”; it is one comparatively cheap layer on a pathway toward much stronger environmental protection. The protected-buildings pillar pushes further in that direction.
This also suggests a possible portfolio strategy rather than a choice between two competing endpoints: deploy relatively inexpensive PPE broadly, while using much smaller amounts of capital to build highly hardened refuges, laboratories, critical facilities or other islands of much stronger protection. Those facilities could themselves be early prototypes of the more comprehensively biohardened society envisaged by Shulman. The relevant question then becomes not “refuges or critical workers?”, but how marginal resources should be allocated across different combinations of protection depth and population coverage.
This comparison could in principle be run across many different protection footprints rather than only the two endpoints. For each candidate footprint, one could estimate the probability that the intervention can actually be deployed, the resulting increase in the probability of a valuable long-run future, and the total cost. Candidate footprints might range from a few highly hardened sites, to a selected technical workforce, to complete critical supply chains, to one or several geographically independent regions, to most or all critical workers.
The relevant scale also need not be measured simply by the number of people protected. Functional coverage may matter more than population coverage. Protecting 90% of the workers in a system may achieve little if one indispensable upstream input, facility, or maintenance capability fails. Conversely, protecting a smaller but complete chain of energy, food, communications, scientific, manufacturing, and logistics capabilities across suitable geographies might preserve disproportionately more recovery capacity.
The cost curve may also be driven by things other than the protective equipment itself. In a PPE strategy, for example, respirator procurement may be cheap relative to identifying the right workers, pre-positioning equipment, distributing it during a crisis, training users, maintaining supplies and filters, and keeping workers’ families, homes, transport, food, power, and workplaces sufficiently safe that they can actually continue working. If distribution or supporting systems dominate cost, “number of masks stockpiled” is not the relevant scaling variable.
In principle, the same schematic calculation above could therefore be evaluated at many points along this frontier:
Cost-effectiveness(F) ≈ P(deployment succeeds | F) × ΔP(valuable long-run future | F) / cost(F),
where F is the protected footprint. What I have not found is an analysis tracing this curve and showing where marginal returns peak.
The likelihood of “recovery” is also underspecified: different collapse scenarios may have very different probabilities of demographic, technological, institutional, or values recovery.
Previous work tried to estimate the probability of technological recovery after severe civilizational collapse. Rodriguez (2022) developed one of the most explicit estimates, while Aird (2020) proposed modelling different kinds of recovery under different causes and depths of collapse. I have not found the corresponding analysis showing how updated views about recovery, persistent biological threats, intervention cost, tractability, threat coverage, and the value of preserving functioning systems jointly imply that the balance now favors critical-worker and system-continuity approaches—or whether these approaches should instead be complements to deliberately hardened refuges.
I am not arguing that refuges are better than the Four Pillars.
I am arguing that the comparison of these approaches should be public and possible to build upon.
One strong object-level explanation for the shift is that the earlier literature may have been too optimistic about recovery after deep civilizational collapse. If rebuilding modern civilization from a small surviving population is substantially less likely than previously believed, then preserving much more of the existing workforce, infrastructure, institutions, and industrial base becomes correspondingly more valuable.
There is some evidence for such an update. But the evidence is considerably less clear-cut than a simple move from “recovery is almost certain” to “recovery is unlikely.”
The literature does appear to have become somewhat more cautious about civilizational recovery, but I do not think it supports a clean numerical story in which an earlier consensus of approximately 99% recovery has been replaced by a new consensus at, say, 70% or 50%.
The earlier literature was already heterogeneous. Beckstead (2015) judged full recovery after a severe catastrophe to be significantly more likely than not, while emphasizing substantial uncertainty. Ord (2020) was more explicitly agnostic, noting views ranging from roughly a 99.9% probability of recovery to a 90% probability of non-recovery. Rodriguez (2022) produced estimates of roughly 97–99.99% eventual technological recovery under a simplified model, but published the work as outdated and said that she had subsequently become somewhat more pessimistic, with an important update coming from the possibility that extreme, long-lasting climate change could make agriculture and subsequent technological development substantially harder. MacAskill (2022) remained relatively optimistic, putting recovery from a return to preindustrial technology at 95% or more given current natural resources, but below 90% after exhaustion of easily accessible fossil fuels. Jehn (2024) similarly concludes that getting back to preindustrial civilization seems fairly plausible under favorable conditions, while getting beyond that to industrial civilization remains much more uncertain.
I therefore read this literature as a real update against treating recovery as almost automatic, rather than as establishing a new low probability of recovery. In particular, I do not see good evidence for replacing an old estimate of approximately 99% with a new consensus estimate of 70%, 60%, or 50%.
There is also a potentially countervailing technological update. Pinsent (2023) argues that persistent low fertility could make biological population recovery harder, but explicitly discusses automation and digital people as ways biological population growth could eventually become less important. This consideration is becoming more salient as AI capabilities improve. Current frontier AI agents can already autonomously perform some economically useful technical work and, on some software benchmarks, complete tasks that would take human experts days or weeks. METR (2026) also stresses important limitations: current evidence is heavily concentrated in software and other cognitive work, and current agents remain much weaker on judgment, reliability, and open-ended tasks.
If machine capabilities continue to advance, a catastrophe occurring later could therefore leave survivors with substantially more machine labor and machine expertise than earlier recovery models assumed. This could reduce the minimum human population needed for recovery while simultaneously increasing the importance of preserving electricity, compute, hardware, communications, machine tools, and other technological infrastructure. The relevant scarce resource may increasingly be preserved productive capability rather than human headcount alone.
Another plausible explanation is a shift toward environment-to-human transmission scenarios. In a transient human-to-human pandemic, a refuge can in principle isolate until transmission outside has fallen sufficiently for survivors to re-emerge. With a persistent environmental threat (“E2H” or Environment-to-Human, such as mirror bacteria), there may be nothing comparable to wait out: a protected group could survive while the outside environment remains dangerous indefinitely.
The Four Pillars approach is explicitly framed as keeping society and civilization functioning while buying time for medical countermeasures. It combines PPE with protected buildings, detection, and eventual medical countermeasures, and its discussion of environmentally persistent threats explicitly considers positive-pressure filtered spaces where outside contamination is high. This increases the value of preserving scientific, industrial, energy, and manufacturing capabilities, because merely surviving the initial event is not enough if somebody still has to develop and deploy a durable way out.
But environmental persistence does not appear to have been an entirely new technical consideration. Earlier refuge proposals already contemplated pandemic-proof isolation and secure facilities in which protected scientific personnel could develop medical countermeasures. In other words, the older refuge strategy was not simply based on the assumption that survivors could wait a few weeks for human-to-human transmission to disappear.
The more important effect of persistent environment-to-human threats may therefore be on the exit problem rather than the feasibility of keeping a small population alive. They increase the value of retaining enough technical capability to make the outside world usable again. But that leaves an unresolved threshold question: how much technical and industrial capacity must survive for a viable exit strategy to exist? In some scenarios a relatively small, highly equipped technical nucleus might be sufficient; in others, recovery may require protecting complete industrial supply chains, large workforces, or distributed capabilities across several geographies. I have not found an analysis establishing where that threshold lies.
The 2022 SHELTER workshop report is relevant evidence about the direction of discussion at the time: it records support for hardening crucial facilities and argues that “the best shelter is a functional society.” I participated in that workshop, however, and my recollection is that these were primarily exploratory judgments and discussion outputs rather than conclusions from a systematic comparison of interventions. I therefore treat the report as evidence that this strategic intuition already existed, not as the reassessment I am looking for.
AI may separately change the comparison through the threat model itself. In a successful AI-takeover scenario where humans are permanently disempowered, neither a refuge nor broad critical-worker protection necessarily preserves an independent valuable human future. That consideration therefore does not by itself favor broad societal resilience over refuges; it may instead reduce the long-run value of both approaches in that part of the threat space.
Resilience could still matter if it makes irreversible disempowerment harder or slower (time to AGI takeover, buying time for other interventions). Broad critical-worker protection preserves much greater immediate human capability and may make active resistance or recovery easier. A small, highly hardened refuge could create a different problem for an AI actor: if the occupants can remain relatively quiet and independent for a long period, the actor may still need to become confident that they cannot later re-emerge and rebuild. It is not obvious to me which architecture creates the larger obstacle and is something AI forecasting and scenario modelling might be better placed to answer.
The relevant comparison may therefore be which intervention most increases the difficulty, uncertainty, or time required to achieve irreversible human disempowerment—and whether humans can use that additional time to change the outcome. I explore this question in more detail in my separate analysis of the x-risk case for PPE.
I still cannot find where these considerations—recovery, intervention scale and cost, persistent biological threats, and AI threat models—were brought together into an actual comparative analysis of the intervention strategies.
This is not only a question about intellectual history. The reason for the shift could change what we should do now.
One possibility is that refuge work was deprioritized because the underlying analysis changed. Later work does seem more cautious than some of the most optimistic earlier estimates of civilizational recovery, while persistent environment-to-human threat models increase the value of maintaining scientific and industrial capabilities long enough to develop a durable countermeasure. If those considerations substantially reduce the marginal value of protecting a small population, they could provide an important object-level explanation for the strategic shift.
Changing AI threat models could also reduce the value of downstream human resilience in scenarios ending in permanent AI disempowerment. But this does not obviously favor broad critical-worker protection over refuges: in such scenarios both may have little long-run value, while in intermediate scenarios their relative value may depend on which more effectively delays or prevents irreversible disempowerment.
But I have not found an analysis showing how large those updates would need to be, or how they compare with differences in cost, tractability, and robustness between intervention types. Nor is it obvious that technological change moves the comparison only toward larger surviving populations: increasingly capable automation could reduce the amount of human labor needed for recovery while simultaneously increasing the importance of preserving a smaller set of technological systems.
More broadly, even if preserving more people monotonically increases recovery probability, that does not tell us where on the protection scale marginal resources should go. The relevant target might be all critical workers, but it might instead be a smaller set of particularly pivotal workers, one complete industrial ecosystem, several geographically independent capability clusters, or a sequence in which highly robust nodes are built first and broader protection is added as resources increase.
This is potentially empirically tractable. One could evaluate different footprints using the same cost-effectiveness framework above, varying the number and type of people protected, protection depth, supply-chain completeness, geographic distribution, deployment probability, and total cost. In a PPE strategy, for example, the optimum may depend less on how many respirators can be purchased than on how large a distribution and support system can be made reliably functional during the relevant catastrophe.
Without such a marginal scaling analysis, “protect all critical workers” looks to me more like a plausible strategic target than a demonstrated optimum.
But another possibility is that at least some refuge work remains technically valuable while being unattractive to major organisations or funders for reasons such as optics, institutional fit, or other non-object-level considerations. I do not currently know how important those factors were, and I do not want to assume that they were decisive.
The distinction is action-relevant. Even if broader societal resilience is the right overall direction, the optimal scale, sequencing, composition, and depth of protection may differ substantially from simply aiming to protect all critical workers. If small-footprint interventions are strongly dominated on cost-effectiveness grounds, resources should go elsewhere. But if the first increments of highly hardened protection buy unusually cheap coverage of catastrophic tail scenarios, refuges or hardened technical nodes could form one layer of a broader resilience portfolio, alongside wider PPE and protected-building deployment.
That is another reason I would like the strategic reassessment to be explicit.
My confidence is quite different across the claims here:
A relevant disclosure is that I came into this field through the earlier line of work. I initially worked on civilizational shelters and later developed that work into E2H-focused shelter and resilience projects. That gives me some personal attachment to this intellectual lineage and may make me unusually sensitive to its disappearance.
It also gives me a reason to want to understand the change from first principles.
If a serious comparative reassessment of these strategies exists, I would like to read it. If it does not, I think it would be useful for someone to do it.
Why am I not bullish on civilizational refuges for reducing bio xrisk? I believe bio xrisk is concentrated in E2H scenarios or AI takeover, and neither of these things benefit much from refuges. For both of these threat models, there is value in refuges but only if they are a safe staging area to take back control of the environment or against an AI, which means you probably need a full industrial base worth of people/infrastructure to be protected. Society-wide biohardening and PPE is therefore the better option even on purely longtermist grounds (and has the additional property of also being much more valuable on non-longtermist grounds as well!)