I notice someone added the “AI risk skepticism” tag to this post. That seems reasonable, because I do express skepticism of the extreme certainty that some have about misalignment risk, even though I think that risk is real.
But my main goal with this post is to argue for, not against, the reality of a certain type of catastrophic AI risk that I believe many are largely ignoring and inadvertently worsening: the risk that AI is a powerful technology which can enable a permanent authoritarian lock-in over the human future, and that draconian controls over the enabling resources of that technology, especially in a time of rapid democratic backsliding, could help bring into existence.
Warning shots/accidents are normally discussed in the frame of generating political will, by convincing a previously unpersuaded public or policymakers that AI is unsafe and action must be taken.
I think this is a mistake.
Accidents (which might be relatively small-scale), in AI as in other fields, are useful mainly for generating real-world, non-hypothetical failure cases in all their intricate detail, thereby yielding a model organism which can be studied by engineers (and hopefully reproduced in a controlled manner) to better understand both the circumstances in which such scenarios might arise, and countermeasures to prevent them.
This is analogous to how aircraft accidents are investigated in depth by the NTSB so as to learn how to prevent similar accidents. There’s already political will to make aircraft safe, but there’s only so much that can be done from the ivory tower without real-world experience.
The choices are:
Stop AI development permanently.
Pause AI temporarily until we make it safe.
Muddle through.
The printing press and electricity were existentially dangerous technologies, because they enabled everything that came after, including AI. When those technologies were developed, however, the world wasn’t globalized enough, nor were nations powerful enough, that a permanent stop button could have been pressed. By contrast, perhaps a permanent “stop AI” button could be pressed today, however I don’t see any way of doing so short of entrenching a permanent totalitarian state.
So that leaves pausing until we make it safe, or muddling through.
But I think the aircraft accident analogy works quite well for AI: there’s only so much that safety research can do from the ivory tower without experience of AIs being used in the real world. So I think the “pause until we make it safe” option is illusory.
That leaves muddling through, as we’ve done with every technology before: We discover problems, hopefully at a small scale, and fix or mitigate them as they arise.
There are no guarantees, but I think it’s our best bet.
Another easy thing you can do, which I did several years ago, is download Kiwix onto your phone, which allows you to save offline versions of references such as Wikipedia, WikiHow, and way, way more. Then also buy a solar-powered or hand-crank USB charger (often built into disaster radios such as this one which I purchased).
For extra credit, store this data on an old phone you no longer use, and keep that and the disaster radio in a Faraday bag.
It varies, but most treaties are not backed up by force (by which I assume we're referring to inter-state armed conflict). They're often backed up by the possibility of mutual tit-for-tat defection or economic sanction, among other possibilities.
A better argument is that the wildness of the next century means our models of the future are untrustworthy, which should make us pretty suspicious of any claim that something is the P = 1 - ε outcome without a watertight case for the proposition.
There doesn't seem to be such a watertight case for AI takeover. Most threat models[1] rest heavily on the assumption that transformative AI will be single-mindedly optimizing for some (misspecified or mislearned) utility function, as opposed to e.g. following a bunch of contextually-activated policies[2]. While this is plausible, and thus warrants significant effort to prevent, it's far from clear that this is even the most likely outcome "absent highly specific conditions", never mind a near certainty.
One possible explanation is an expectation of massive deflation (perhaps due to AI-caused decreases in production costs) which the structure of Treasury Inflation Protected Securities (TIPS) and other inflation-linked government bonds — the source of your real interest rate data — doesn't account for.
While TIPS adjust the principal (and corresponding coupons) up and down over time according to changes in the consumer price index, you ALWAYS get at least the initial principal back at maturity. Typical "yield" calculations, however, are based on the assumption that you get your inflation-adjusted principal back (which you do if inflation was positive over its term, as it usually would be historically).
This means that iff there's net deflation over its term, the "yield" underestimates your real rate of return with TIPS by the amount of that deflation.
I notice someone added the “AI risk skepticism” tag to this post. That seems reasonable, because I do express skepticism of the extreme certainty that some have about misalignment risk, even though I think that risk is real.
But my main goal with this post is to argue for, not against, the reality of a certain type of catastrophic AI risk that I believe many are largely ignoring and inadvertently worsening: the risk that AI is a powerful technology which can enable a permanent authoritarian lock-in over the human future, and that draconian controls over the enabling resources of that technology, especially in a time of rapid democratic backsliding, could help bring into existence.
Warning shots/accidents are normally discussed in the frame of generating political will, by convincing a previously unpersuaded public or policymakers that AI is unsafe and action must be taken.
I think this is a mistake.
Accidents (which might be relatively small-scale), in AI as in other fields, are useful mainly for generating real-world, non-hypothetical failure cases in all their intricate detail, thereby yielding a model organism which can be studied by engineers (and hopefully reproduced in a controlled manner) to better understand both the circumstances in which such scenarios might arise, and countermeasures to prevent them.
This is analogous to how aircraft accidents are investigated in depth by the NTSB so as to learn how to prevent similar accidents. There’s already political will to make aircraft safe, but there’s only so much that can be done from the ivory tower without real-world experience.
The choices are:
The printing press and electricity were existentially dangerous technologies, because they enabled everything that came after, including AI. When those technologies were developed, however, the world wasn’t globalized enough, nor were nations powerful enough, that a permanent stop button could have been pressed. By contrast, perhaps a permanent “stop AI” button could be pressed today, however I don’t see any way of doing so short of entrenching a permanent totalitarian state.
So that leaves pausing until we make it safe, or muddling through.
But I think the aircraft accident analogy works quite well for AI: there’s only so much that safety research can do from the ivory tower without experience of AIs being used in the real world. So I think the “pause until we make it safe” option is illusory.
That leaves muddling through, as we’ve done with every technology before: We discover problems, hopefully at a small scale, and fix or mitigate them as they arise.
There are no guarantees, but I think it’s our best bet.
Another easy thing you can do, which I did several years ago, is download Kiwix onto your phone, which allows you to save offline versions of references such as Wikipedia, WikiHow, and way, way more. Then also buy a solar-powered or hand-crank USB charger (often built into disaster radios such as this one which I purchased).
For extra credit, store this data on an old phone you no longer use, and keep that and the disaster radio in a Faraday bag.
I’m calling for a six month pause on new font faces more powerful than Comic Sans.
It varies, but most treaties are not backed up by force (by which I assume we're referring to inter-state armed conflict). They're often backed up by the possibility of mutual tit-for-tat defection or economic sanction, among other possibilities.
A better argument is that the wildness of the next century means our models of the future are untrustworthy, which should make us pretty suspicious of any claim that something is the P = 1 - ε outcome without a watertight case for the proposition.
There doesn't seem to be such a watertight case for AI takeover. Most threat models[1] rest heavily on the assumption that transformative AI will be single-mindedly optimizing for some (misspecified or mislearned) utility function, as opposed to e.g. following a bunch of contextually-activated policies[2]. While this is plausible, and thus warrants significant effort to prevent, it's far from clear that this is even the most likely outcome "absent highly specific conditions", never mind a near certainty.
e.g. Cotra and Ngo et al
as proposed e.g. by shard theory
It appears the UK's index-linked gilts, at least, don't have this structural issue.
See "redemption payments" on page 6 of this document, or put in a sufficiently large negative inflation assumption here.
One possible explanation is an expectation of massive deflation (perhaps due to AI-caused decreases in production costs) which the structure of Treasury Inflation Protected Securities (TIPS) and other inflation-linked government bonds — the source of your real interest rate data — doesn't account for.
While TIPS adjust the principal (and corresponding coupons) up and down over time according to changes in the consumer price index, you ALWAYS get at least the initial principal back at maturity. Typical "yield" calculations, however, are based on the assumption that you get your inflation-adjusted principal back (which you do if inflation was positive over its term, as it usually would be historically).
This means that iff there's net deflation over its term, the "yield" underestimates your real rate of return with TIPS by the amount of that deflation.