People in AI safety/EA spheres should reorient now towards stopping continued AI capabilities escalation.
Historically, people have been unwilling to straightforwardly say “This AI situation is disgustingly dangerous, and we need to stop” and then actually work towards making this happen. I think a lot of this is downstream of deference to high status people who weren’t willing to take positions that seemed extreme.
The current situation is very, very bad, and there is no plan to make it better if we continue increasing AI capabilities at the current rate.
We need to stop things as soon as possible, so the community and the world can orient and work out what to do. We don't need to have the full plan yet, we just need to get to a state where we aren't in imminent danger.
I predict that this position will become increasingly obvious, and people will wish they reoriented earlier. I think this will be clear ex ante; when people look back they will think that they should have reoriented earlier, given the information they had at the time.
Regarding "the organizations with the most money share the fewest details":
A lot of AI safety people I know are excited about Anthropic and I don't fully understand why. My instinct is to distrust company because they're the one pushing the AI arms race, autonomously hacking three companies, and IPO'ing, but am likely missing something because of how respected they are within this space. Are there examples where they have counterfactually produced some result, policy, or finding that has slowed down capabilities progress more than they have themselves pushed capabilities?
Quick rambly comment from my side (because I didn’t expect this to be spread outside of the publication and wasn’t aware of timelines :))
Most of the writing is done by Dirk-Jan. I gave input and part of this followed from a conversation we had on stage during EAGxAmsterdam.
I think Dirk-Jan is a rare example of an experienced professional from the ODE field who is willing to engage critically with EA ideas in good faith, and I am very happy we started this discussion :) As a community, I think we should cherish encounters with constructive outside experts
My quick take if you need to take one thing away from this thought-piece: we need to make sure that, as a community, we incentivize careful analysis of unintended consequences, and then update accordingly! The incentive is to write away unintended effects qualitatively, or handwave a modest discount (or bonus, for positive effects)
When you start a project/organisation, it makes sense to put most effort into making your theory of change very rigorous and de-risk it as much as possible. Once you scale that project/organisation, you should probably put the same effort on analyzing whether the causal links in your ToC could break-down at scale, and pay attention to the causal effects that fall outside your ToC all together!
The best version of this, as always, employs a good mix of qualitative and quantitative thinking. If anything, I personally think discussions on unintended effects and “backfire risks” are often too qualitative, and miss a serious attempt at sizing.
Lastly, a prompt everyone can use:
This resonates with me, thanks for sharing.
I suspect you're a lot less alone on this one than you are with donating to effective charities and were being early to COVID. (I also suspect you agree, but spelling this out for other readers).
It's tricky to get meaningful information from surveys, but this one from June finds that a slight majority of respondents thought a superintelligence would seek to take control of humanity. AI existential risk awareness is also steadily climbing, this survey says up to 34%. I would bet that still only a very small % of people are as worried as you and me and many in this community, but lots of people are somewhat worried and could probably quite quickly get more worried and be activated to do something about it.
I think this points to saying what we're worried about, and saying it clearly (like this!), especially how quickly loss of control risks are becoming very real. I'm very bullish on AI safety comms, and expect there's lots we can do there.
I've had a vague sense that the EA Forum is declining in quality posts/conversation despite having lots of posts, so had an AI tool (Astra) do a bunch of random analyses to look at various measures of activity level, and I think these vaguely confirm my sense. In particular, there has been a decline in commenting from higher karma users, and fewer posts are getting lots of upvotes, despite lots of people still posting. As a lover of the Forum, this seems bad and sad.