Yeah, I was influenced by work on value lock-in, and it seems plausible to me that any small set of actors with control over superhuman AI could have their moral epistemics corrupted, leading to s-risks. I'm also fairly convinced that escalating inequality (cf. this thread) could enable similar outcomes. It does seem clearer to me that extreme power concentration broadly tends to lead to lower-value futures than outright s-risks, but there is quite a bit of uncertainty here. Do appreciate @Vakus Drake's points here and seems worth thinking more about this
Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools (e.g. by encoding them redundantly across pathways, or shifting them into representations the tools do not capture)
I feel this is somewhat obvious in the sense of arbitrarily deceptive AI. However, most mechinterp work in recent times is only assumed to work short of arbitrary deception, and this seems like a fine hedge (though a practical solution to ELK may still be possible)
We are making good progress in the AI S-risk space and research is on track
I think AIS is way under-invested in reducing risks from, for example, extreme power concentration, or consequences of not attaining friendly AI/value alignment solutions. In general it also seems to me that AIS over-invests in reducing AI scheming, and many present research directions could make certain s-risks more likely
Yeah, I was influenced by work on value lock-in, and it seems plausible to me that any small set of actors with control over superhuman AI could have their moral epistemics corrupted, leading to s-risks. I'm also fairly convinced that escalating inequality (cf. this thread) could enable similar outcomes. It does seem clearer to me that extreme power concentration broadly tends to lead to lower-value futures than outright s-risks, but there is quite a bit of uncertainty here. Do appreciate @Vakus Drake's points here and seems worth thinking more about this
I feel this is somewhat obvious in the sense of arbitrarily deceptive AI. However, most mechinterp work in recent times is only assumed to work short of arbitrary deception, and this seems like a fine hedge (though a practical solution to ELK may still be possible)
I think AIS is way under-invested in reducing risks from, for example, extreme power concentration, or consequences of not attaining friendly AI/value alignment solutions. In general it also seems to me that AIS over-invests in reducing AI scheming, and many present research directions could make certain s-risks more likely