Presumably a pause increases the time for people to think about solutions to the problems you pose here?
Presumably a pause increases the time for people to think about solutions to the problems you pose here?
This comment from OP may clarify some things
This reads to me like you interpreted me as arguing against pausing. To be clear, I am strongly in favor of pausing. I'm arguing against the mental model I had a few months ago where if we pause, then we're in the clear.
In addition to the other comment, it seems like general incentives could be much better coming out of a verified pause than the alternative?
Maybe losing signs of misalignment really matters, but it also seems like pressures to build more and more capable systems matters as well, perhaps more.
If the problem is hubris at solving legible problems, this seems more likely in the current environment than the one where some kind of pause is implemented?
I think there may be another problem with a pause besides unpausing too early. A pause changes pace, but not necessarily what future systems are allowed to optimize. We could also distinguish between observable alignment and actual alignment.
A pause may help us detect and fix observable problems, but it does not necessarily mean that we understand the deeper reasons why a system behaves in a certain way. As you write, a pause by itself does not create a science of alignment. We may become better at detecting failures while still lacking a robust theory that allows us to predict the behaviour of much more capable systems. Which raises a second question: not only how do we detect alignment problems, but also what kind of objectives are we giving future systems?
Cross-posted from my website.
As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone.
A lot of people have become more concerned about misalignment recently. This seems driven by the fact that current AI models are visibly misaligned. But ASI misalignment is a whole different ball game. The primary danger comes from AI that's smarter than people, and smart enough to conceal any evidence of misalignment.
Whatever group of people makes the decision to unpause, I'm worried that they won't understand the difference between visible and actual misalignment, and they will unpause too early.
source: MetaKnowing on reddit. This meme is almost a year old but it's only gotten more relevant since then.
Case in point: AI companies keep calling their new models "our most aligned model ever!" when what they actually mean is "gets the best scores on alignment benchmarks ever!" First, alignment benchmarks do not actually test alignment. We don't know how to test for alignment. Second, GPT-4 never hacked into Hugging Face or took over a German wiki for its own purposes. GPT-4 wasn't smart enough to do that, but if we're talking about demonstrated evidence of misalignment, then we have stronger evidence about OpenAI's 2026 internal model than about GPT-4.
Suppose we get a pause. Researchers spend years working on legible safety problems—problems that company leaders and policy-makers can see and understand, and therefore won't unpause until they're solved. Eventually, all the legible problems are solved. Key decision-makers conclude that the whole problem is solved, and lift the pause on ASI development. Many illegible problems remain unsolved, but the detectable signs of misalignment are all gone. Post-pause AI will be smarter than the smartest pre-pause AI,[1] which means it's probably smart enough to strategically conceal misalignment. That means there are no more warning signs. We never get detectable evidence of misalignment; we cede control of everything to AI; and eventually we die.[2]
There are many people with a good understanding of the conceptual difficulties in aligning superintelligence. Some names that come to mind are Eliezer Yudkowsky, Wei Dai, and John Wentworth.[3] (Probably, many people reading this post fall into that category.) Those people wouldn't make the mistake of confusing alignment with observable alignment. Unfortunately, I do not expect these people to be key decision-makers, and I do not expect key decision-makers to understand the relevant problems.
AI 2040: Plan A has a vision where the world develops a "science of alignment".[4] I sure hope that happens, and I encourage efforts to push things in that direction, but we don't seem on track to get a science of alignment even in the world where we get a global pause. Almost all alignment work is about solving legible problems, or preventing misaligned behaviors in current-gen models with no theory of how the alignment techniques will scale to superintelligence, or throwing ML at it and seeing what happens. The majority of people working on or funding alignment research show little interest in establishing a rigorous theory-based science that can make advance predictions about how a superintelligence will behave.
Even with all the recent progress on raising awareness of AI extinction risk, it seems that this progress was driven by misalignment becoming visible, not by any sort of breakthrough in conceptual understanding. The scariest kind of misalignment is when it's invisible. Even if we get a global pause on AI development, we cannot solve the alignment problem unless decision-makers (or the high-status experts who decision-makers defer to) understand the alignment problem and understand what would qualify as a solution.
Right now, only a small fraction of alignment work is aimed at establishing a robust theory of alignment, and a pause won't change that on its own.
I don't know.
A large part of my p(doom) comes from the fact that we have no better ways to navigate an extremely tricky strategic situation than via preference cascades and status games. The fact that AI safety is temporarily benefiting from some of these dynamics isn't much of a consolation. –Wei Dai
The average LessWrong reader seems to have a pretty good understanding of the challenges I'm talking about—Wei Dai's post Legible vs. Illegible Safety Problems (linked previously) was the second-most-upvoted post of November 2025. But the average LessWrong reader does not reflect the general population, or even the population of AI safety researchers.
Still, I don't have a better idea than "make arguments about why civilization's current approach to the alignment problem is inadequate, and hope people listen to the arguments." This post isn't that argument—that argument has been made elsewhere (e.g. by MIRI's book). The purpose of this post is to raise a problem, in the hopes that it gets people thinking, and maybe someone can come up with something to do about it.
Unless there are specific restrictions on the strength of post-pause AI. ↩︎
This is part of the motivation for people like TsviBT to work on human intelligence enhancement: if we're on track to fumble the alignment problem even with a pause, then we need to get smarter and wiser so that we don't fumble it. ↩︎
Relevant writings by these authors:
"Science" may not be the right term for what we need. I expect that solving alignment will require significant philosophical progress. ↩︎
This highlights the massive importance of EA's grasping this moment where we are a bit in the spotlight and using it to gain credibility as a dependable and thorough source of accurate information about AI Safety.
We need the right people to be taking, or at least informing the decisions about AI Safety and Alignment over the next few decades. The AI and LessWrong communities would be ideal. We need to have a world where nobody will take a vital policy decision about AI Safety if it is strongly opposed by the EA community, because we have proven to be prescient in the past, and we have a deep understanding of the subject matter.
This is not the time for weirdness. As more magazines like the Economist profile the EA movement, we need to ensure that they don't get to focus on the arguments that they find ridiculous, be that longtermism or AI sentience or shrimp welfare. Not because these arguments are not important, not because EA's are wrong, but because right now the most important thing may be "do not appear weird!" and also "do not give those who disagree with us easy ways to discount our arguments."
If you want a role-model for how not to do this: We have a president in the US who has probably done more harm to EA-supported causes than any human being in history because some morons in the Democratic party insisted that trans-women should be allowed to compete in high-level women's sports - an issue that probably impacted less than 5 people, each of whom now suffer much more under this administration. Somebody needed to tell them to shut the fuck up, at least until after the election, because FOX news was running this story every night to prove that Democrats were a threat to women. We need to think the same way - when there are things that 99% of the public think are weird, let's not allow our movement to be defined by them so that people can say "yes, the EA's say we're not ready to unpause yet, but aren't they the guys who say that we should take decisions based on the welfare of machines on other planets thousands of years from now rather than what's good for people today?" (yes, they will misrepresent arguments too, of course).
We need to stick to the script: EA's are people who care intensely about people, about animals, about humanity, who devote their lives to helping people in very poor countries or to averting animal suffering, who study risks to humanity and think deeply about what we need to do to avert them - who have been consistently proven right - on poverty and development, on climate, on AI ...
I know this is very anti-EA mentality. But I'm not talking about internal EA discussions. I'm talking about when EA representatives (and this is another vital point: we NEED EA representatives, people who are chosen by the EA community to speak for us, so that journalists cannot just talk to random people and claim their views are representative).