Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?
We’ve seen two large and extremely capable swarms from OpenAI in the last few months:
* 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched...
TLDR: Everyone’s talking about what the money could do, but few about how to decide where it goes.
This post is part of the new series of articles on cross-cause giving and the new wave of philanthropy. Stay tuned to the EA Forum and our Substack for the latest takes on topics such as giving now vs. later, common pitfalls in cause prioritization, and other crucial considerations from the Cross-Cause Fund (CCF) team...
Summary:
First, I give several different angles on how I feel about reinforcement learning:
* Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffolding
* Recent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progress
* I’m worried things might get worse:...
Consider adopting the term o-risk.
William MacAskill has recently been writing a bunch about how if you’re a Long-Termist, it’s not enough merely to avoid the catastrophic outcomes. Even if we get a decent long-term future, it may still fall far short of the best future we could have achieved. This outcome — of a merely okay future, when we could have had a great future — would still be quite tragic.
Which got me thinking: EAs already have terms like x-risk (for existential risks, or things which could cause human extinction) and s-risk (for suffering risks, or things which could cause the widespread future suffering of conscious beings). But as far as I’m aware, there isn’t any term to describe the specific kind of risk that MacAskill is worried about here. So I propose adding the term o-risk: A risk that would make the long-term future okay, but much worse than it would optimally be.
might want to check out this (only indirectly related but maybe useful).
https://forum.effectivealtruism.org/posts/zuQeTaqrjveSiSMYo/a-proposed-hierarchy-of-longtermist-concepts
Personally don't mind o-risk think it has some utility but s-risk ~somewhat seems like it still works here. An O-risk is just a smaller scale s-risk no?