Epistemic status: Speculation from two decently informed advocates armed with anecdata.
Note on process: After having some version of this conversation several times and saying, “we should probably write about this publicly,” we took the less heroic route: we recorded one of our conversations, fed the transcript into an LLM, and then substantially revised the structure, substance, and framing ourselves. We will not be sharing the transcript, as it is in...
TLDR: Take the population ethics quiz here: https://mdickens.me/pop-ethics/
Population ethics is an oft-overlooked subfield within ethics. Many people hold views that they don't realize contradict each other, or that have strange implications that they wouldn't endorse if they thought about it more.
Not just that—population ethics is a BIG DEAL. A lot of ethical decisions hinge on how you think about changes in future populations....
Headline finding: I audited 17 AI Safety Talent programmes. Zero of 17 have published any comparison group, rejected-applicant follow-up, matched control or randomisation. Not one. Every programme that mentions a counterfactual does it by asking participants to self-report.
Background
At least $70 million...
A lot of AI safety people I know are excited about Anthropic and I don't fully understand why. My instinct is to distrust company because they're the one pushing the AI arms race, autonomously hacking three companies, and IPO'ing, but am likely missing something because of how respected they are within this space. Are there examples where they have counterfactually produced some result, policy, or finding that has slowed down capabilities progress more than they have themselves pushed capabilities?
Anthropic has done many expensive actions to credibly signal that it cares about AI safety.
I think folks could reasonably say that they are engaged in race dynamics that makes them net negative. But it also doesn't seem crazy to look at the actions above, and admire what they've done (especially compared to other actors in the space).
Anthropic has produced a lot of alignment research. On certain theories* of where AI danger comes from, that research has been useful enough that Anthropic's overall impact is net positive.
*Those theories are wrong. In a sentence: Anthropic's alignment research is almost entirely centered around how to produce desired behaviors in the short term, with no understanding of how to make an ASI continue to be aligned once you are no longer smart enough to detect misalignment.
This becomes much easier to explain when you allow for the possibility that people are biased or irrational or have conflicts of interest. People want to work on cool problems, they want to be close to the action, they want to get rich, and they want to think of their friends as good people; then they reason backward from that bottom line to determine that Anthropic must be the good guys.
I guess the discussion should not turn so much around shaming and blaming particular companies, but changing the structural incentives driving forward the current wave of AI development. Under current circumstances, any company with any CEO with any personnel would largely be incentivised to take the path that Anthropic, OpenAI etc. have taken.
I don't think they're respected within this space? But of course 'this space' is vague.
By 'this space' I meant AI Safety. I at least see a lot of Anthropic roles being published on 80k and have seen AI safety clubs direct people towards Anthropic fellowships as they would MATS.