Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?
We’ve seen two large and extremely capable swarms from OpenAI in the last few months:
* 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched...
TLDR: Everyone’s talking about what the money could do, but few about how to decide where it goes.
This post is part of the new series of articles on cross-cause giving and the new wave of philanthropy. Stay tuned to the EA Forum and our Substack for the latest takes on topics such as giving now vs. later, common pitfalls in cause prioritization, and other crucial considerations from the Cross-Cause Fund (CCF) team...
Summary:
First, I give several different angles on how I feel about reinforcement learning:
* Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffolding
* Recent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progress
* I’m worried things might get worse:...
From SE Geyges: Is METR a meaningful check on Anthropic?.
Part of the reason we are in this position is because nobody outside the EA community cares about AI safety enough to fund it or work in it. I’m very sympathetic to this; some of these close connections are inevitable.
However, it can also be true that these connections are
completelyunacceptable (edit), and that Anthropic should be trying harder than they are to find or create auditors that are genuinely independent. How do we do that?I think METRs checks are meaningful and maybe the best we have at the moment. They are also like you say compromised and the conflict of interests are immense with huge personal overlap between the labs and safety orgs, and funding streams too.
Geyges seems largely correct, but if we can't convince governments to regulate properly its better METR is in there doing it. After all METR exposed more about the hugging face hack than Open AI did on its own.
I don't think a framing of "these connections are completely unacceptable" is helpful given these problems. I think "compromised and far from ideal" is a better framing. Its better to do something than do nothing. I agree its best if government installed internal auditors like they do for banks, but that ain't happening any time soon.
Companies like Anthropic and Open AI are selfish animals by nature. They may have moments where good humans inside might do the right thing, but fundamentally they thirst for profit and growth. After IPO this will only get worse. We should never expect a company to regulate itself or its industry. Self regulation for harmful companies is a terrible idea and never works.
How can ANthropic "create" a genuinely independent auditor? this seems impossible, almost and Oxymoron.