Note: This post was crossposted from Planned Obsolescence by the Forum team, with the author's permission. The author may not see or respond to comments on this post.
Subtitle: We can’t develop safety standards if we have to rely on opaque judgment
All views are my own and do not represent my employer.
In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress,[1] have led a number of researchers and leaders in the industry to believe that the risk that humanity loses control of AI is now urgent enough to warrant slowing down the pace of AI development soon.
This has led to a lot of discussion about the role of third party evaluators in verifying “pacing commitments”, evaluating safety cases, or auditing compliance with safety policies. I think these are valuable roles for third party groups to aim to fulfill, but I also worry we’re putting the cart before the horse in all this talk of “verifying” and “auditing” things.
The science on loss-of-control risk is, to put it generously, nascent. Companies are not in the business of making structured, standardized claims about risk and safety that can be cleanly verified or falsified. There are no settled methods for measuring whether increasingly powerful AI systems might try to undermine human control or seize control entirely — companies report on various alignment benchmarks, but it is hard to tell whether their training process simply taught the models to game these benchmarks. It is hard to confidently bound risk even over a horizon of months because there is vast and hard-to-reduce uncertainty about whether recursive self-improvement will suddenly lead to AI systems capable of toppling states or overpowering humanity.
To the extent that AI companies choose to proceed with AI development, they can and should do their best to manage their own safety practices. They should continue to unilaterally slow as needed to do this, and there is room to do more in that vein. But to durably reduce loss-of-control risk to an acceptable level, we will probably need to develop shared technical standards for how to manage this risk adequately and enforce those standards uniformly across the industry, including internationally.[2]
We currently lack the basic prerequisites needed to have a conversation about safety standards: a shared understanding of the current state of loss-of-control risk, what companies are currently doing or not doing to try to manage it, and how well these efforts do or don’t work. The claims that AI companies make about these topics are currently far too high-level and imprecise to cleanly “verify” or “disprove.”[3] Moreover, no company is even making a set of public safety claims that adequately cover all the important questions about loss-of-control risk.[4]
In this situation, I think what we most need is not verification of the safety claims companies are currently making, but production of a far greater quantity and quality of concrete evidence about loss-of-control risk and how companies are managing it.
AI companies themselves can and should be generating and publishing much more of this evidence, but third party investigations of major questions relevant to risk can also help. These investigations would look less like trying to prove or disprove claims in a structured “safety case”, and more like generating and operationalizing hypotheses relevant to risk and coming up with a wide range of creative ways to generate evidence about whether these hypotheses are true: in other words, like doing science.
And like all good science, I believe third party investigations should aim for evidence transparency: that is, publicly sharing as much of the empirical evidence underlying their conclusions as possible. External scientists will generally have to trust that third party evaluators like METR are not outright lying, but the goal is that insofar as possible they should not have to trust that we are making the right judgment calls on tough scientific questions — they should be able to read our reports and form their own views. Evidence transparency has three big benefits:
Given capacity constraints and the need to navigate redactions for IP, it will not be possible to achieve perfect evidence transparency for third party investigations. For example, in our Hugging Face report, we wrote a Methodology appendix but did not share our prompts (and certainly not our code). But evidence transparency was our north star, and I think we got pretty far — if you read the report carefully, your interpretation of what it means is about as good as mine.[5]
We need to admit that understanding and managing loss-of-control risk is an open scientific problem that no AI company or external research group has settled. Standard scientific norms of evidence transparency evolved to handle this situation, and we should lean into them. If AI companies and third party researchers make it a priority over the next several months, I believe we can dramatically improve the state of public scientific evidence about loss-of-control risk and mitigations. I think this would put industry leaders and policymakers in a much better position to implement a functional governance regime for loss-of-control risk.
Anthropic and OpenAI have reported that measures like the amount of code written by AI and number of agent-workdays per human workday have sped up a lot recently. No company systematically reports verified measures of outputs like how much algorithmic progress has sped up, and from public evidence it is not clear how the input measures companies report translate into the output measures that matter. Most researchers I know seem to have the sense that progress in AI capabilities has sped up recently, but they are unsure whether this is the start of a continuous acceleration.
The extent to which a set of standards reduces risk depends on how stringent they are, which is in turn ultimately a political question. For example, some researchers believe that alignment practices should not be considered "adequate" unless they are provably robust; trying to meet this bar would likely necessitate pausing development for many years or decades. The political process could settle on working standards that are far short of this while still raising the bar substantially on top of what any developer is currently achieving.
For example, it would take a lot of judgment and thought to operationalize a claim like “We don’t believe we’re overfitting to our methods of detecting misalignment,” and testing it could require poring over training data or running novel evaluations.
For one thing, I believe no company is making a crisp claim that they are confident they will not be able to build uncontrollable superintelligence within six months.
As stated in the redaction summary statement at the top of the report, METR and OpenAI were able to agree on how to describe redactions for all the important evidence underlying our views. This means that if an outside researcher carefully read the report and came to a different interpretation about how concerning the agents’ behavior was or what it means for the future, I wouldn’t feel like they should trust me because I had seen so much more private evidence that informed my view.