I write about economics, social sciences, philosophy, AI, X-risk, effective altruism, and longtermism. My main focus with regards to causes is on AI-safety: particularly within a manufacturing context. I also work on understanding how discount rates shape moral behaviour, and engage with policymakers and the public discourse on tackling extreme power concentration.
I'm not convinced the recent capabilities advancements (highly concentrated in pure math, the section which by definition can't be applied) support the forecasts that misaligned superintelligence is on balance harmful. In general, I think there is not enough evidence to draw the conclusion that multi-agent alignment (which was far from inevitable ex-ante in comparison to the sometimes monotheist conceptions of ASI) is impossible, especially with AI eventually automating and enforcing the institutional mechanisms.
However, this is not guaranteed either. The labs are making the biggest bet that humanity has taken by far; that RSI will automate and solve the alignment problem, and the ensuing intelligence explosion will solve most of humanity's suffering. This explains much of the race dynamics until now, where the balance of tail-risks are now undoubtedly on the upside. Note also the competition with China, which plays particular relevance for open-source.
My p(doom) probably sits somewhere just over halfway (to account for the right-tailed distribution here) between my pre and (initial) post Hugging-Face likelihoods: perhaps 8%. It would be unwise not to update since this saga unfolded.
Some of the latest macroeconomic models from the likes of Anthropic and Acemoglu suggest rather grim futures for employment relative to what economists were previously predicting, hence its fair to say our methods lean overly conservative towards excess rigour here. Perhaps the recent backlash from the mathematicians, and some of the sentiment behind the anti data-centre movement, reflect these (increasingly justified) anxieties.
Nonetheless, this is orders of magnitude more dangerous than nuclear weapons. That's enough for me to take AI-safety incredibly seriously!
TLDR: I've updated towards pausing further AI development indefinitely.
When Scott Alexander proposed regulating AI like clinical drugs, many (including myself) balked at this given the sclerosis of FDA or related bodies. Yet HF updates me towards treating frontier AIs as nuclear. There onerous regulation is likely good.
In general, most bureaucratised regulatory regimes are bad. In some cases, like nuclear technology or ensuring planes are safe to fly on, they're welfare improving. It's clear that AI belongs in the latter category.
Also trace inversion is a thing, so you can get open source weights to be roughly similar to that of Fable/Mythos. The slowdown camp were right. At bare minimum, all labs should pause training, further development, and releases indefinitely until we figure out how to align AIs and regulate them (and their use by humans).
I reckon the current capabilities we have now are sufficient for accelerating progress towards curing cancers etc. So my balance has shifted towards minimising the existential risks now.
Pause AI advocates are correct. If governments could coordinate internationally to achieve such (big if), I'd support it. I used to be highly sceptical of doomer arguments, yet Hugging Face is almost a textbook LW scenario and no one knows how to spot or prevent such scheming. Sandboxing, guardrails, constitutions etc. don't work.
One reason we don't see large doom futures or doom insurance markets (including in catastrophe bonds) is that a large proportion of the risk is uninsurable, due to uncertainty on enforcability. Collateralised instruments underprice p(doom), and prices cannot adjust to the values that collaterisation participants would accept.
All these contracts and securities are reliant on courts to enforce them, alongside arbitration mechanisms (e.g. in the event of defaults or payment disputes). However in a doom state, these courts don't exist. Therefore these contracts are typically unenforceable, unless there's a sequential timing decision involved in their sale (where you can hopefully clear before unenforceability occurs). Such risks cannot be precisely measured however, so the rates required to justify such contracts exceed those participants are willing to accept, so the market is almost non-existent. In other words, these risks are uninsurable, so futures and uninsurance market values imply a lower p(doom) than is actually the case. However, the consequence of this is that such markets are not incentive-compatible so (to my knowledge?) don't exist.
You see this same pattern in the nonexistence of a futures market in galaxies too, in property rights over galaxy ownership of the sort discussed recently on LessWrong. Such optimistic capabilities forecasts are also consistent with much wider variance and larger tail-risks, so again enforcability concerns (in terms of the contracts and the property rights over galaxies) make this market incomplete. Moreover, collateral values would need to be incredibly high to justify such trades on a somewhat esoteric and outlandish bet, and not all participants are willing to provide such amounts.