Glen Weyl and friends have recently called attention to a problem they call reverse alignment. While they think the AI alignment issue is important, they argue that most people in the space neglect thinking about the institutions that would interface with advanced AI capabilities (especially if aligned). In brief, something like wrong institutions + aligned, advanced AI = likely worse outcomes. This seems plausible, and it intersects with some emerging EA interest in things like democracy, though the argument Weyl et al. develop is less about specific political outcomes and more about the foundational social and political institutions that govern how we behave, decide, and ultimately harness (or not) the AI wave to our benefit. So, more longtermist.
Some of this feels familiar. For example, I’m having a hard time entirely divorcing reverse alignment from efforts to, say, do fieldbuilding around AI and health (meaning creating the conditions so that deployers of health solutions can leverage AI). In both cases, we are thinking about creating conditions that optimize for deployers of AI. The difference, I think, is the level of analysis, where reverse alignment asks us to think at the ‘building blocks of society’ level. If true, this makes it both 1) quite difficult to execute/get right and 2) exponentially more transformative in expectation. More smart people should be working on this.
The reverse alignment coalition: https://reversealignment.ai
Thank you for sharing this. An aligned model does not an aligned system make; the application layer must also be aligned.
This has parallels I believe to the discussion I was trying to initiate here, unsuccessfully, about the importance of investing in prosocial AI applications. Even if reverse alignment goals are broader.