[...]
Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
I was struck that Dario (1) says that recursive self-improvement has already begun, and (2) says that rogue agents could take over “the entire Internet” within six to 12 months. (This seems similar to Ajeya’s prediction that agents could establish a more permanent rogue deployment inside an AGI company within six months).
Hi Ben. Thanks for sharing that.
It is unclear to me what this means. Dario says "This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described". I have gone through the instances of "recursive self-improvement" and "RSI" in the linked sources, and I did not find any concrete description of what it means for RSI to start. There is a sense in which humanity has always been building on past knowledge.
Dario says "it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails". I doubt humans will lose control over the entire internet to bots over the next 12 months. I am open to bets against short timelines for transformative AI (TAI), or what they supposedly imply, up to 10 k$.
Hi Vasco. I also feel unclear about what it means for RSI to properly begin; I’m assuming humans are still in the loop at Anthropic. I’m reminded of Toby Ord’s analogy:
Has anybody taken you up on the bet yet? I’m not betting; I have no idea what’s going on!
I have this and this bets resolving at the end of 2027.
Fair and funny. I suggested the bet having other readers in mind.
Good luck, I hope you win!
Where would you now like the donation to go?
Thanks. Me too. I am thinking about suggesting to Greg doing a similar bet resolving at the end of 2030, where I would initially donate to Greg's preferred charity what I win from the 1st bet plus some more money. I could probably bet like 40 k$ in total, and then win 80 k$ adjusted for inflation or growth in stocks at the end of 2030.
I would make a donation to Rethink Priorities (RP) restricted to research on moral weights led by Bob Fischer. Here is some context.