Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?
We’ve seen two large and extremely capable swarms from OpenAI in the last few months:
* 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched...
TLDR: Everyone’s talking about what the money could do, but few about how to decide where it goes.
This post is part of the new series of articles on cross-cause giving and the new wave of philanthropy. Stay tuned to the EA Forum and our Substack for the latest takes on topics such as giving now vs. later, common pitfalls in cause prioritization, and other crucial considerations from the Cross-Cause Fund (CCF) team...
Summary:
First, I give several different angles on how I feel about reinforcement learning:
* Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffolding
* Recent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progress
* I’m worried things might get worse:...
We've banned the user denyeverywhere for a month for promoting violence on the forum.
As a reminder, bans affect the user, not the account(s).
If anyone has questions or concerns, please feel free to reach out, and if you think we made a mistake here, you can appeal the decision
I don't have any strong views on whether this user should have been given a temporary ban vs a warning, but (unless the ban was for a comment which is now deleted or a private message, which are each possible, and feel free to correct me if so), from reading their public comments, I think it's inaccurate (or at least misleading) to describing them as "promoting violence". Specifically, they do not seem to have been advocating that anyone actually use violence, which is what I think the most natural interpretation of "promoting violence" would be. Instead, they appear to have been expressing that they'd emotionally desire that people who hypothetically would do the thing in question would face violence, that (in the hypothetical example) they'd feel the urge to use violence, and so on.
I'm not defending their behavior, but it does feel importantly less bad than what I initially assumed from the moderator comment, and I think it's important to use precise language when making these sorts of public accusations.