[Added 13Jun: Submitted to OpenPhil AI Worldviews Contest - this pdf version most up to date]
This is an accompanying post to AGI rising: why we are in a new era of acute risk and increasing public awareness, and what to do now.
If you apply a security mindset (Murphy’s Law) to the problem of AI alignment, it should quickly become apparent that it is very difficult. And the slow progress of the field to date is further evidence for this. There are 4 major unsolved components to alignment:
And a variety of corresponding threat models.
Any given approach that might show some promise on one or two of these still leaves the others unsolved. We are so far from being able to address all of them. And not only do we have to address them all, but we have to have totally watertight solutions, robust to AI being superhuman in its capabilities. In the limit of superintelligent AI, the alignment needs to be perfect.
OpenAI brags about a 29% improvement in alignment in their GPT-4 announcement. This is not going to cut it! The alignment paradigms used for today’s LLMs only appear to make them relatively safe because the AIs are weaker than us. If the “grandma’s bedtime story napalm recipe” prompt engineering hack actually led to the manufacture of napalm, it would be immediately obvious how poor today’s level of practical AI alignment is.
Plans that involve increasing AI input into alignment research appear to rest on the assumption that they can be grounded by a sufficiently aligned AI at the start. But how does this not just result in an infinite, error-prone, regress? Such “getting the AI to do your alignment homework” approaches are not safe ways of avoiding doom. Tegmark says that he can use his own proof checker to verify that a superintelligent AI’s complicated plans are safe. But (security mindset): how do you stop physical manipulation of your proof checker's output (rowhammer or more subtle)?
The above considerations are the basis for the case that disjunctive reasoning should predominantly be applied to AI x-risk: the default is doom. All the doom flows through the tiniest crack of imperfect alignment once the power level of the AI is superhuman. Every possible exploit, loophole or unknown that could engender catastrophic risk needs to be patched for this to not be the case (or we have to get very lucky).
Despite this, many people in EA who take AI x-risk seriously put P(doom|AGI) in the 1-10% range. I am struggling to understand this. What is happening in the other 90-99%? How is it that it appears that the default is “we’re fine”? I have asked this here, and not had any satisfactory answers.
My read on this so far is that low estimates for P(doom|AGI) are either borne of ignorance of what the true difficulties in AI alignment are; stem from wishful thinking / a lack of security mindset; or are a social phenomenon where people want to sound respectable and non-alarmist; as opposed to being based on any sound technical argument. I would really appreciate it if someone could provide a detailed technical argument for believing P(doom|AGI)≤10%.
To go further in making the case for doom by default given AGI, there are reasons to believe that alignment might actually be impossible (i.e. superintelligence is both unpredictable and controllable). Yampolskiy is of the opinion that “lower level intelligence cannot indefinitely control higher level intelligence". This is far from being a consensus view though, with many people in AI Alignment considering that we have barely scratched the surface in looking for solutions. What is clear is that we need more time!
The case for the reality of existential risk has been well established. Indeed, doom even looks to be a highly likely outcome from the advent of the AGI era. Common sense suggests that the burden of proof is now on AI labs to prove their products are safe (in terms of global catastrophe and extinction risk).
There are two prominent camps in the AI Alignment field when it comes to doom and foom. The Yudkowsky view has very high P(doom|AGI) and a fast take-off happening for the transition from AGI to ASI (Artificial Superintelligence). The Christiano view is more moderate on both axes. However, it’s worth noting that even Christiano is ~50% on doom, and ~a couple of years from AGI to ASI (a “slow” take-off that then speeds up). Metaculus predicts weak AGI within 3 years, and takeoff to happen in less than a year.
50% ≤ P(doom|AGI) < 100% means that we can confidently make the case for the AI capabilities race being a suicide race. It then becomes a personal issue of life-and-death for those pushing forward the technology. Perhaps Sam Altman or Demis Hassabis really are okay with gambling 800M lives in expectation on a 90% chance of utopia from AGI? (Despite a distinct lack of democratic mandate.) But are they okay with such a gamble when the odds are reversed? Taking the bet when it’s 90% chance of doom is not only highly morally problematic, but also, frankly, suicidal and deranged.
What to do about this? Join in pushing for a global moratorium on AGI. For more, see the accompanying post to this one (of which this was originally written as a subsection).
Based solely on my own impression, I'd guess that one reason for the lack of engagement on your original question stems from the fact that it felt like you were operating within a very specific frame, and I sensed that untangling the specific assumptions of your frame (and consequently a high P(doom)) would take a lot of work. In my own case, I didn’t know which assumptions are driving your estimates, and so I consequently felt unsure as to which counter-arguments you'd consider relevant to your key cruxes.
(For example: many reviewers of the Carlsmith report (alongside Carlsmith himself) put P(doom) ≤ 10%. If you've read these responses, why did you find the responses uncompelling? Which specific arguments did you find faulty?)
Here's one example from this post where I felt as though it would take a lot of work to better understand the argument you want to put forward:
When I read this, I found myself asking “wait, what are the relevant disjuncts meant to be?”. I understand a disjunctive argument for doom to be saying that doom is highly likely conditional on any one of {A, B, C, … }. If each of A, B, C … is independently plausible, then obviously this looks worrying. If you say that some claim is disjunctive, I want an argument for believing that each disjunct is independently plausible, and an argument for accepting the disjunctive framing offered as the best framing for the claim at hand.
For instance, here’s a disjunctive framing of something Nate said in his review of the Carlsmith Report.
Phrased this way, Nate offers a disjunctive argument. And, to be clear, I think it’s worth taking seriously. But I feel like ‘disjunctive’ and ‘conjunctive’ are often thrown around a bit too loosely, and such terms mostly serve to impede the quality of discussion. It’s not obvious to me that Nate’s framing is the best framing for the question at hand, and I expect that making the case for Nate’s framing is likely to rely on the conjunction of many assumptions. Also, that’s fine! I think it’s a valuable argument to make! I just think there should be more explicit discussions and arguments about the best framings for predicting the future of AI.
Finally, I feel like asking for “a detailed technical argument for believing P(doom|AGI) ≤ 10%” is making an isolated demand for rigor. I personally don’t think there are ‘detailed technical arguments’ P(doom|AGI) greater than 10%. I don’t say this critically, because reasoning about the chances of doom given AGI is hard. I'm also >10% on many claims in the absence of 'detailed, technical arguments' for such claims in the absence of such arguments, and I think we can do a lot better than we're doing currently.
I agree that it’s important to avoid squeamishness about proclamations of confidence in pessimistic conclusions if that’s what we genuinely believe the arguments suggest. I'm also glad that you offered the 'social explanation' for people's low doom estimates, even though I think it's incorrect, and even though many people (including, tbh, me) will predictably find it annoying. In the same spirit, I'd like to offer an analogous argument: I think many arguments for p(doom | AGI) > 90% are the result of overreliance on specific default frame, and insufficiently careful attention to argumentative rigor. If that claim strikes you as incorrect, or brings obvious counterexamples to mind, I'd be interested to read them (and to elaborate my dissatisfaction with existing arguments for high doom estimates).
I don't find Carlsmith et al's estimates convincing because they are starting with a conjunctive frame and applying conjunctive reasoning. They are assuming we're fine by default (why?), and then building up a list of factors that need to go wrong for doom to happen.
I agree with Nate. Any one of a vast array of things can cause doom. Just the 4 broad categories mentioned at the start of the OP (subfields of Alignment) and the fact that "any given [alignment] approach that might show some promise on one or two of these still leaves the others unsolved." is enough to provide a disjunctive frame! Where are all the alignment approaches that tackle all the threat models simultaneously? Why shouldn't the naive prior be that we are doomed by default when dealing with something alien that is much smarter than us? [see fn.6].
Can you give an example of such assumptions? I'm not seeing it.
This blog is ~1k words. Can you write a similar length blog for the other side, rebutting all my points?
It does strike me as incorrect. I've responded to / rebutted all comments here, and here, here, here, here etc, and I'm not getting any satisfying rebuttals back. Bounty of $1000 is still open.
Ay thanks, sorry I’m late back to you. I’ll respond to various parts in turn.
My initial interpretation of this passage is: you seem to be saying that conjunctive/disjunctive arguments are presented against a mainline model (say, one of doom/hope). In presenting a ‘conjunctive’ argument, Carlsmith belies a mainline model of hope. However, you doubt the mainline model of hope, and so his argument is unconvincing. If that reading is correct, then my view is that the mainline model of doom has not been successfully argued for. What do you take to be the best argument for a ‘mainline model’ of doom? If I’m correct in interpreting the passage below as an argument for a ‘mainline model’ of doom, then it strikes me as unconvincing:
Under your framing, I don’t think that you’ve come anywhere close to providing an argument for your preferred disjunctive framing. On my way of viewing things, an argument for a disjunctive framing shows that “failure on intent alignment (with success in the other areas) leads to a high P(Doom | AGI), failure on outer alignment alignment (with success in the other areas) leads to a high P(Doom | AGI), etc …”. I think that you have not shown this for any of the disjuncts, and an argument for a disjunctive frame requires showing this for all of the disjuncts.
Nate’s Framing
I claimed that an argument for (my slight alteration of) Nate’s framing was likely to rely on the conjunction of many assumptions, and you (very reasonably) asked me to spell them out. To recap, here’s the framing:
For this to be a disjunctive argument for doom, all of the following need to be true:
That is, the first point requires an argument which shows the following:
A Conjunctive Case for the Disjunctive Case for Doom:[1]
If I try to spell out the arguments for this framing, things start to look pretty messy. If technical alignment were “pretty easy”, and tackled by a culture which competently pursued alignment research, then I don’t feel >90% confident in doom. The claim “if humanity has < 20 years to prepare for AGI, then doom is highly likely” requires (non-exhaustively) the following assumptions:
So far, I’ve discussed just one disjunct, but I can imagine outlining similar assumptions for the other disjuncts. For instance: if we have >20 years to conduct AI alignment research conditional on the problem not being super hard, why can’t there be a decent chance that a not-super-competent research community solves the problem? Again, I find it hard to motivate the case for a claim like that without already assuming a mainline model of doom.
I’m not saying there aren’t interesting arguments here, but I think that arguments of this type mostly assume a mainline model of doom (or the adequacy of a ‘disjunctive framing’), rather than providing independent arguments for a mainline model of doom.
Future Responses
I think so! But I’m unclear what, exactly, your arguments are meant to be. Also, I would personally find it much easier to engage with arguments in premise-conclusion format. Otherwise, I feel like I have to spend a lot of work trying to understand the logical structure of your argument, which requires a decent chunk of time-investment.
Still, I’m happy to chat over DM if you think that discussing this further would be profitable. Here’s my attempt to summarize your current view of things.
Suggestions for better argument names are not being taken at this time.
Thanks for the reply. I think the talk of 20 years is a red herring as we might only have 2 years (or less). Re your example of "A Conjunctive Case for the Disjunctive Case for Doom", I don't find the argument convincing because you use 20 years. Can you make the same arguments s/20/2?
And what I'm arguing is not that we are doomed by default, but the conditional on being doomed given AGI; P(doom|AGI). I'm actually reasonably optimistic that we can just stop building AGI and therefore won't be doomed! And that's what I'm working toward (yes, it's going to be a lot of work; I'd appreciate more help).
Isn't it obvious that none of {outer alignment, inner alignment, misuse risk, multipolar coordination} have come anywhere close to being solved? Do I really need to summarise progress to date and show why it isn't a solution, when no one is even claiming to have a viable, scalable, solution to any of them!? Isn't it obvious that current models are only safe because they are weak? Will Claude-3 spontaneously just decide not to make napalm with the Grandma's bedtime story napalm recipe jailbreak when it's powerful enough to do so and hooked up to a chemical factory?
Ok, but you really need to defeat all of them given that they are disjuncts!
Can you elaborate more on this? Is it because you expect AGIs to spontaneously be aligned enough to not doom us?
Judging by the overall response to this post, I do think it needs a rewrite.
Here's a quick attempt at a subset of conjunctive assumptions in Nate's framing:
- The functional ceiling for AGI is sufficiently above the current level of human civilization to eliminate it
- There is a sharp cutoff between non-AGI AI and AGI, such that early kind-of-AGI doesn't send up enough warning signals to cause a drastic change in trajectory.
- Early AGIs don't result in a multi-polar world where superhuman-but-not-godlike agents can't actually quickly and recursively self-improve, in part because none of them wants any of the others to take over - and without being able to grow stronger, humanity remains a viable player.
Thanks!
I don't think anyone is seriously arguing this? (Links please if they are).
We are getting the warning signals now. People (including me) are raising the alarm. Hoping for a drastic change of trajectory, but people actually have to put the work in for that to happen! But your point here isn't really related to P(doom|AGI) - i.e. the conditional is on getting AGI. Of course there won't be doom if we don't get AGI! That's what we should be aiming for right now (not getting AGI).
Nate may focus on singleton scenarios, but that is not a pre-requisite for doom. To me Robin Hanson's (multipolar) Age of Em is also a kind of doom (most humans don't exist, only a few highly productive ones are copied many times and only activated to work; a fully Malthusian economy). I don't see how "humanity remains a viable player" in a world full of superhuman agents.