AI safety
AI safety
Studying and reducing the existential risks posed by advanced artificial intelligence

Quick takes

63
9d
3
The "Desperate" AI Safety Talent Bottleneck is deeply misleading. Please.stop. I worked at Google. Google does not tell applicants that there is a talent bottleneck and they are desperately hiring. They say "we're cool, join us!"  I didn't get into Harvard.  Harvard does not tell applicants that they are desperately seeking students. I am not misled.  I've gotten rejected from countless AI Safety organizations. AI Safety tells applicants - we are in desperate need of talent (operators/generalists)! Please join us!  I've coached 70+ aspiring career pivoters on navigating the ecosystem. A common theme is the dejection & disappointment from all this rejection. Because of this false marketing, I am the one picking up the pieces to calibrate professionals on how difficult, competitive, and picky organizations are, how to build context, and how to endure the marathon, which is like any other job hunt.  Please.stop. Some alternatives: * We are an exciting, growing organization and you should apply for our roles! * Many people are excited to work for us - here are qualities of candidates we're especially excited about. * EA roles are competitive, and we would still love to see your application. Relevant posts: * https://forum.effectivealtruism.org/posts/B6d8Wzk4gNzHsXvdi/ai-safety-is-extremely-bottlenecked-on-grantmakers?commentId=n2Rd7RR4y9EnKP42F * https://forum.effectivealtruism.org/posts/b82SLXwEHRCs3TFJA/why-experienced-professionals-fail-to-land-high-impact-roles * https://forum.effectivealtruism.org/posts/jmbP9rwXncfa32seH/after-one-year-of-applying-for-ea-jobs-it-is-really-really 
27
9d
4
AI safety needs people everywhere but quickly stated, current talent bottlenecks to me look like: -- Founders -- Grantmakers -- (technical) Research leads  -- Policy entrepreneurs and implementors (which includes a lot of technical work) -- bets in international coordination and/or cooperation -- All manner of supporting talent -- program leads, ops proper, public outreach, content creators, comms   Most sought-after qualities for talent are: -- context, mission alignment, domain understanding, sophisticated views on AI strategy and threat modelling etc. -- "good judgement", "sound epistemics", "reasoning transparency" and other similar ideas/meta-skills from the EA/rationalist cannon -- a willingness to get shit done/bias for action (rather than be in learning mode, or people who need a lot of management and oversight, or folks with too many preferences/constraints) -- low ego, similar to above -- ambitious folks, since they would be really trying to be their own managers, take on bigger projects, grow themselves and their teams etc.    Finally, even having these, it's not enough to just claim to have these; job-seekers mainly trip up in being able to demonstrate and be legible about having them.
11
4d
6
A lot of AI safety people I know are excited about Anthropic and I don't fully understand why. My instinct is to distrust company because they're the one pushing the AI arms race, autonomously hacking three companies, and IPO'ing, but am likely missing something because of how respected they are within this space. Are there examples where they have counterfactually produced some result, policy, or finding that has slowed down capabilities progress more than they have themselves pushed capabilities? 
3
3d
I went to an ai ethics vs ai safety debate, my friend was speaking at yesterday in the Cambridge union my summary of what happened is: 1. most people view negatively (1) Sam A starting openai to solve alignment faster than deepmind (2) Holden Karnofsky funding openai. some people, especially on ethics side, view negatively (3) Dustin Moskowitz investing in anthropic 2. the pattern of beliefs/behaviours that led to this could be characterised as (1) "thinking agi is inevitable", (2) "thinking we know best and can control things more than we can", (3) "believing AGI has a bunch of positives so it should be developed in some aligned sense" and (4) "attracting talented people who believe the above, maybe more than we do" 3. i think its reasonable to criticise any/all of those beliefs (Richard Ngo does so well in “retrospective on AI alignment”) 4. but the current situation makes it much harder to criticise the belief that AGI is inevitable/makes AGI not super useful as a term. 5. both sides agree on a bunch of governance interventions being on a great bet. accountability & transparency, auditing, 6. but people continue to disagree about how much we should support wholesale upheaval of capitalist structures that incentivise continued development, how possible pause is etc. Also the specific language people use is significantly different 1. 6 specifically is very hard to talk about without getting caught up in identities. but less of an identity war than 3. also man the ethics vs safety debate was very unfun to watch. people far more focused on 3. and 6 than i’d like
11
14d
How impactful would it be to copies of @Garrison's new book Obsolete to elected officials who belong to its political target audience (Dems, particularly left ones) and might not have been responsive to traditional x-risk-centric comms (e.g. IABIED)?
54
4mo
2
More EA undergrads should do political volunteering. It's impactful AND fun. Choose an election that's impactful (e.g. AI safety candidate) and neglected (e.g. primaries in always-blue/red places), couch-crash the weekend there, and volunteer with the campaign. I say this after doing 15 hours of street canvassing myself. I was surprised by how anecdotally impactful and fun it was. If you like people-watching, talking to strangers, and/or joining passionate projects for a weekend, I think you'll also love this. I wish I thought of this earlier. Literature on the impact (Claude-generated): Kalla & Broockman's meta-analysis of 49 field experiments finds zero average persuasive effect in general elections, but effects do show up when voters lack a partisan cue (i.e. primaries and ballot measures). Mann & Haenschen (2024) find mobilization effects (e.g. canvassing) are 33-76% larger in low-attention races than in high-attention ones. Your marginal volunteer hour goes much further in a primary.
15
1mo
Applications to SPAR Fall 2026 are closing tomorrow Aug 18 EOD Anywhere on Earth. SPAR is the ecosystem's biggest AI safety research program, and it's part-time remote. We still have many strong projects across AI safety, AI policy, and biosecurity with very few applications[1], so please consider applying! This round, we also have 18 non-research/generalist projects that people can apply to, and we have much more mentee capacity than previous rounds; we expect to accept around 500 people into the program. If you have friends who have thought about going into AI safety, spread the word! 1. ^ To be specific, as of 2:25 PM PT, we had around 67 projects with fewer than 20 applications total!
1
2d
Anthropic's September report on misuse of its models poses interesting questions about AI safety. The report, detailing actions ranging from the use of Claude to set up a fake online dating profile farm to the model being employed by the Houthis to design software to guide missiles, was not produced under any legal obligation. Under the TFAIA in California and the EU AI Act, it is only mandated to confidentially disclose safety incidents to authorities, and the jury is still out on whether some of these cases would fall under the definition of 'serious incidents' (in the EU AI Act case) or 'critical safety incidents' (in the TFAIA case). Potential explanations for why the company still decided to produce such a report are not hard to imagine: better PR, internal employee pressure, and institutional inertia. However, these incentives are largely internal to the company and can change. If the board decides that these reports should be toned down to dissociate Anthropic from dangerous uses of AI (perhaps to prepare for an IPO), there is no reason why this would not take place. That would be a serious blow to the cause of AI safety at a time when the potential fallout from increasing model progress has still to be spelled out. Any loss of data on this front would mean one more threat vector left unexplored, and a part of society and the world left potentially unprepared.
Load more (8/270)