Over the past few days, I have watched several scientists and biotech experts on X react with real frustration to claims about the bio-risks of advanced AI, largely spurred by the threat report released by Anthropic. And although I disagree with some of their conclusions, I think people in EA -- working on AI safety and biosecurity -- should pay closer attention to what is producing that reaction.
The discussion began, at least for me, with a thread by by a scientist who said he had experience training a LLM and synthesizing viruses in a lab. His argument, stated bluntly, was that many of the claims about how AI can be used to create dangerous viruses are detached from the reality of biology.
Soon afterward, another scientist who has also worked on frontier AI and physically made viruses in a lab, made a similar argument, although he focused more on what happens after a virus has been engineered and released. I think some of his conclusions go further than the evidence supports, but the fact that scientists with relevant experience see the current discussion as badly miscalibrated makes their underlying concern harder to wave.
I do not concur with the claim that AI poses little bio-risk. I think the AI-bio risk community has a particular issue of communication, and this comes from again, moving too fast. Moving too fast here refers to moving from AI model outputs to actor outputs. Some critics see physical bottlenecks (i.e limitations to the real world) as proof that AI is unable to significantly alter the overall pathway. Some critics see the lack of significant changes to the pathway as proof that the social structures and controls of the forces of today remain intact.
Some of the disputes that exist are structural, but some of the arguments appear to be presenting different phenomenons using the same (or connected) arguments.
A quick note: I am not trying to estimate the probability of an AI-enabled bio catastrophe here or argue that current evidence rules one in or out. My goal is to speak about the extent to which there is a tendency to communicate risk at this level of abstraction that conceals many of the assumptions underlying our conclusions.
What does "AI can design a dangerous virus" actually mean?
When someone says this, they could mean the AI model will design a pathogen by suggesting a sequence or a starting point and providing some descriptions of possible outcomes and effects of possible pathogen modifications. The model may also suggest different pathogen candidate sequences to the scientist. Note that each of these tasks involve a series of steps and some degree of human involvement.
Between the model resulting in a sequence and making a system to produce harm, a lot of steps are needed. Someone has to determine the correct design, acquire all the chemicals, and assemble working biological system, test it, see the results, and determine the reasons for any failures and possibly repeat steps. Even after all these steps are completed and a successful working virus is made, challenges still exist. First, the traits the biologist created must work as expected in human hosts. The traits must be able to replicate in a host enough to prolong the virus and lead to the harm of the host.
When this entire process is summed up by, “AI can design a dangerous virus,” a biologist may think the design and the completed biological system creation are being treated as the same level of difficulty. However, the person speaking may think of building this threat as a long complicated system of steps. While this system of steps may seem complicated, it is likely that long, complicated systems of steps are cited as the reason for not designing such systems in the first place.
In fact, this problem was pretty evident in the responses to the thread on X. Some people were discussing a future AI system that could control sophisticated cloud labs, robotics and other parts of the economy; others were describing a nearer-term system that might persuade or pay humans to perform the physical work; the OP was largely responding to claims about what present or near-term systems could accomplish.
There are distinctions in ability, agents, time frames, and interventions. When a position starts with existing language, bio models, or techniques, and then relies on an economy dominated by autonomous AI systems, that shift should be clear.
What have the uplift studies actually found?
The emerging empirical evidence support a more complicated story than either side usually tells. In this preregistered trial involving 153 novices, participants with access to mid-2025 frontier models were no more likely than internet-only participants to complete the study's core lab workflow: completion rates were 5.2% and 6.6%, respectively. However, it is worthy of note that AI-assisted participants made more progress on some intermediate tasks and performed better on the cell culture component, suggesting a modest and task-dependent benefit even though the primary outcome showed no significant uplift.
Across the wider evidence base, studies of computer-based knowledge and planning tasks have found effects ranging from minimal to substantial, while two publicly reported wet-lab studies found no statistically significant uplift on their primary outcomes. Most studies have also focused on novices, leaving the effects on skilled and partically skilled actors comparatively underexplored.
While the models tested in that paper are dated -- it is difficult to look at this evidence and say that current models have significantly advanced to the stage at which barriers to impactful biological work have been substantially removed. It is equally difficult to say that AI has little to no effect. The more reasonable conclusion is that the impact of AI may vary across tasks and entities, and there is a clear preference of evidence where AI has a positive impact, particularly in the context of computational and informational work and planning, and a lot less evidence that would support the idea that current systems would allow novice users to execute complex physical workflows.
The more difficult interpretive problem is that an unchanged end-to-end success rate indicates that AI may not have changed anything, or may have improved one part of the workflow, and may still be constraining the other part to determine the final outcome.
Risk can change well before full autonomy
In the thread on X, the strongest arguments of the OPs are presented in the context of an extremely challenging scenario: an AI system operating an autonomous integrated virology lab. Demonstrating that this scenario is both technically and economically remote would be important, but it would not support the conclusion that AI has little to no impact on bio-risk.
The more immediate concern is whether AI can substantially enhance performance for actors who already possess some amount of expertise, resources, access, and human partners. Such enhancements could be relevant even if the physical constraints remain, especially if the enhancements impact critical decisions, help avoid failures at challenging steps, or lessen the complexity of combining dispersed capabilities.
The more significant concern should be how much autonomy can be granted before full autonomy, and how effective the current safeguards are on this sliding scale.
The pathway matters more than the model
My perspective has evolved from working through eight different AI-enabled biotechnology pathways. Tracing the integration of models, biological data, specialized tools, synthesis providers, labs, human interactions, and feedback from experiments provides a diverse array of opportunities and potential pathways toward an outcome. My thinking always returned to the same foundational idea: how much a model can do is significantly less important than the context surrounding its use. Specifically, who is using the model, what complementary elements and resources they have available, and the elements of the pathway that they can directly control are what ultimately dictate its potential.
The same model may have little to no real value to a beginner who is unable to identify incorrect answers. It may, however, be useful to an expert scientist who has the necessary expertise and tools to act upon an improvement in design, analysis, or troubleshooting. Thus, the relevant question is actor-relative: what do changes in AI access allow this person/group to achieve?
AI may change constraints rather than eliminate them. Easier computational design may result in physical validation becoming the most important bottleneck. Better lab automation may shift the constraint to access, materials, or approval.
This is why I think that both claims, "the model can do it" and "biology is too hard," are unsatisfactory descriptions. Each identifies part of the problem, while the other conceals part of the problem.
What would better risk communication look like?
I would like to see less conjecture towards statements like “AI can create a biological weapon” and more specific statements that include an identifiable actor, pathway, and time frame. To determine if a model poses a significant bio-risk, we should be able to answer the following:
- who benefits from the alleged effect?
- what goal is this population pursuing?
- what pathway of the system is affected by the AI and how?
- which significant constraints remain?
- what would convince us to change our assessment?
We should also indicate if the basis for our claim is observed activity, an extrapolation from current evidence, or a cautionary scenario. It is appropriate to discuss frameworks that hinge on unknown capabilities, but the reader should not be required to decipher if the claim refers to current models that aid beginners, future systems that aid experts, or self-directed intelligent systems with capabilities that alter the operational environment of a laboratory.
These changes would make some claims more specific and less dramatic. We would focus on more practical questions and move away from rhetorical devices like “does AI change everything?” or whether “biology is too complex.” We might ask whether positive changes in AI and biology would help experts in the field plan or design safer experiments and control more sophisticated systems by integrating distributed resources. These questions could be studied and would lead to more specific, actionable measures.
Credibility Matters
Several of the responses on X show how overconfidence can work in both directions. As many skeptics seem to think that evolution should be part of the analysis of bio-risks associated with AI, I share this sentiment, but I am much less confident that it provides positive and reassuring results. Evolution does not favor making pathogens benign; it depends on the relationship between the virulence and the transmission of the pathogen, at which point transmission happens, what the feature costs the pathogen, and the environment through which it transmits. Inserted sequences are, at times, highly unstable, but again it is not the case that a harmful feature will be eliminated. The same can be said of the evolution of virulence and the stability of viral sequences.
Everyone needs to show more work. On the one hand, those that are risk advocates need to show how the engineered features could endure and fulfill the predicted harmful impacts, and on the other hand, skeptics need to show why evolution would eliminate those features. Both positions also require argumentation of their own, and some of the claims that the positive impacts of AI on biology will completely outweigh the negative impacts and that an extremely sectoring and devastating AI-related virus-triggered pandemic will not threaten civilization, also require independent arguments; neither AI nor viral engineering answer those broader questions.
The EA community is right to continue valuing the willingness to entertain low-probability, high impact threats, but it also must rely on the demonstrated capability of engineering threats and resist the inclination to remain optimistic until proven otherwise.
Telling skeptical scientists that they are focused on today while we are thinking about tomorrow only gets us so far. Any forecast still requires a plausible account of how we move from here to there: which constraints change, which remain, and what evidence would move us. Otherwise, "future capabilities" can become a way of keeping a claim beyond the reach of evidence.
This will remain the case because credibility affects whether scientists participate in evals, companies adopt safeguards, and governments receive advice grounded in how biological work actually happens. A more substantive debate would be more productive: skeptics should identify the step they think remains prohibitive, risk advocates should explain why they expect it to change, and both should say what would change their mind.
Transfecting viruses or bacteria with external DNA is ridiculously simple. There are thousands of kids who do this in middle school research internships and hundreds of thousands who do this in college. Ordering oligo-nucelotides is also very easy to do online. The point of AI biorisk is that you can quickly design a thousand different versions of something many of which will be modifications on known pathogens, many of which might be totally new. You can recombinate those sequences into existing viruses or bacteria or fungi. The biggest piece here is that you don't need to be perfect. You nee only a few to hit - having a 1% or 0.1% success rate is OK. Moreover, they all mutate meaning that even if you didn't do a perfect job they can get more dangerous over time.