Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?
We’ve seen two large and extremely capable swarms from OpenAI in the last few months:
* 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched...
TLDR: Everyone’s talking about what the money could do, but few about how to decide where it goes.
This post is part of the new series of articles on cross-cause giving and the new wave of philanthropy. Stay tuned to the EA Forum and our Substack for the latest takes on topics such as giving now vs. later, common pitfalls in cause prioritization, and other crucial considerations from the Cross-Cause Fund (CCF) team...
Summary:
First, I give several different angles on how I feel about reinforcement learning:
* Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffolding
* Recent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progress
* I’m worried things might get worse:...
It would be helpful when evaluating this project to see some of the work you've already done.
Yeah, I'm potentially interested but would be curious what direction you're thinking of going here.
I'm open to going in whatever direction gives the EA community the most insight into the truth, with whatever presentation encourages the most constructive use of that information. In case you're interested in specifics, I am currently working on a planning document about how specifically to accomplish all that. I can give you access if you wish (Just give me your Google Docs address via PM.).
I'm open to considering directions / direction changes. What are your thoughts so far? :)
I am not sure if you are requesting to see the project, or if you are making a complaint of some sort. It's easy enough for anyone to PM me and request to see the project. Just in case, I updated my post to explicitly invite people to PM me to see the project.
In case this wasn't clear, the project isn't finished yet. Before dumping a lot more hours into it, I want to see whether I'm duplicating anyone's work.
The fact that it is not yet finished is why I did not publish anything about it so far. It's not ready to be published.
The main point of this post is simply to find out whether there are others doing a similar project, and find other people who are interested in helping make sure the project gets completed.
You've described a project at a fairly high level of abstraction. You've already put 20-40 hours in, so your research has already likely taken some specific directions. Sharing a brief summary of this would help people with compatible approaches who think you're doing potentially overlapping work notice that they should reach out to you. It would also help save the time of those who aren't members of that group.
Peter just suggested you mention more details about the project, in the comments. Daniel did too. As a reader, I would have benefited if you'd replied by giving them details about the project. I expect there are more readers like me, who might reach out if a project seemed like it was going in an interesting direction (even if not my preferred direction), but not without such a specific reason to think it's worth their time.
If there are specific reasons for discretion, of course, you can say so.
I think you're saying "There isn't enough information for most readers to decide whether they want to PM you." is that right?
Yes
Okay, what information do you think they need? You mentioned "directions" and "approaches" but that is very vague. I need the specific questions you think readers need answered before they will notify me of similar projects or express interest in what I'm doing.