From SE Geyges: Is METR a meaningful check on Anthropic?.
Part of the reason we are in this position is because nobody outside the EA community cares about AI safety enough to fund it or work in it. I’m very sympathetic to this; some of these close connections are inevitable.
However, it can also be true that these connections are completely unacceptable (edit), and that Anthropic should be trying harder than they are to find or create auditors that are genuinely independent. How do we do that?
Situational Awareness (the hedge unhedged fund) has imploded. It seems much of their gains were from overleveraging their trades, which has then caused them to, in some readings, undergo the largest absolute short-term fund loss in history (over $10B in a few weeks).
They certainly lost a lot of money in July and by being forced to sell, but per the WSJ Aschenbrenner wrote that the fund is still up 80% on the year.
The fund doesn't seem to be closing, and it looks like they weren't forced to sell their stake in Anthropic, so it's not clear to me how much of a failure or a success the project as a whole is so far.
See also this from the Financial Times.
Edit: see also this tweet claiming to be the full letter Aschenbrenner sent to his LPs last night.
? It is his fund. He started it.
His creditors supplied the money—it is their judgement I’m questioning.
The losses, as you pointed out, happened due to highly leveraged bets. He didn't expect memory stocks to bleed as much as they did.
This is disqualifying if you are running a hedge fund. It’s literally in the name—you are supposed to hedge your positions in order to prevent an unexpected situation from tanking the whole fund.
Besides, it is possible that the fund will survive because his core thesis has paid off exceptionally well so far
We don’t judge funds by whether they don’t go bankrupt, we judge them by their performance against a market index over a long period of time. Even if the positions are net up by 2× or so, this is not particularly impressive in and of itself over a short period, because of survivorship bias. If you make a bunch of stupidly leveraged bets on different sectors, one of them is likely to pay off very well, but not for long, and not through downturns.
The AI sector has monotonically gone up since the release of ChatGPT—any overleveraged investor in this space would be likely to produce incredible gains. If one’s fund gets obliterated at the first market downturn because one was overleveraged, all this proves is that you managed your fund badly, not that you’re some kind of savant genius market whisperer.
(c.f. anything written about Cathie Wood in 2022—a lot of it has aged very poorly)
Per Bloomberg, the Trump administration is considering restricting the equivalency determination for 501(c)3s as early as Tuesday. The equivalency determination allows for 501(c)3s to regrant money to foreign, non-tax-exempt organisations while maintaining tax-exempt status, so long as an attorney or tax practitioner claims the organisation is equivalent to a local tax-exempt one.
I’m not an expert on this, but it sounds really bad. I guess it remains to be seen if they go through with it.
Regardless, the administration is allegedly also preparing to directly strip environmental and political (i.e. groups he doesn’t like, not necessarily just any policy org) non-profits of their tax exempt status. In the past week, he’s also floated trying to rescind the tax exempt status of Harvard. From what I understand, such an Executive Order is illegal under U.S. law (to whatever extent that matters anymore), unless Trump instructs the State Department to designate them foreign terrorist organisations, at which point all their funds are frozen too.
These are dark times. Stay safe 🖤
The World Happiness Report 2025 is out!
Finland leads the world in happiness for the eighth year in a row, with Finns reporting an average score of 7.736 (out of 10) when asked to evaluate their lives.
Costa Rica (6th) and Mexico (10th) both enter the top 10 for the first time, while continued upward trends for countries such as Lithuania (16th), Slovenia (19th) and Czechia (20th) underline the convergence of happiness levels between Eastern, Central and Western Europe.
The United States (24th) falls to its lowest-ever position, with the United Kingdom (23rd) reporting its lowest average life evaluation since the 2017 report.
I bang this drum a lot, but it does genuinely appear that once a country reaches the upper-middle income bracket, GDP doesn’t seem to matter much more.
Also featuring is a chapter from the Happier Lives Institute, where they compare the cost-effectiveness of improving wellbeing across multiple charities. They find that the top charities (including Pure Earth and Tamaika) might be 100x as cost-effective as others, especially those in high-income countries.
The first thing I wondered about when this report came out was: how did India do?
In that quick take I asked how India's self-reported life satisfaction dropped an astounding -1.20 points (4.97 to 3.78) from 2011 to 2021, even as its GDP per capita rose +51% in the same period; China in contrast gained about as much self-reported life satisfaction as you'd expect given its GDP per capita rise. This "happiness catastrophe" should be alarming to folks who consider happiness and life satisfaction what ultimately matters (like HLI), since given India's population such a drop over time adds up to roughly ~5 billion LS-years lost since 2011, very roughly ballparking (for context, and keeping in mind that LS-years and DALYs aren't the same thing, the entire world's DALY burden is ~2.5 billion DALYs p.a.). Even on a personal level -1.20 points is huge: 10x(!) larger than the effect of doubling income at +0.12 LS points (Clarke et al 2018 p199, via HLI's report), and comparable to major negative life events like widowhood and extended unemployment. So it mystified me that nobody seems to be talking about it.
Last year's WHR reported a 4.05 rating averaged over th... (read more)
The Global Fund and Gilead have announced that Lenacapavir, the new 6-month PrEP treatment for HIV, will be made available in 120 countries at no profit. Gilead originally agreed to license Lenacapavir, royalty-free, to local manufacturers, but have now also agreed to directly supply doses until those manufacturers reach capacity, likely for up to 2 million at-risk people. The Global Fund will direct resourcing and deliver doses.
Extremely good news, and a possible silver lining after potentially losing PEPFAR.
Thanks this is indeed amazing news and I'm actually a bit surprised at the commitment, would love to hear the full story and how much it will actually cost, it could still be quite expensive. Super cool as well that eventually it will actually be manufactured in many of the countries that will use the drugs
Although good news, I don't think its the best ever news on HIV treatments. I would rate both the invention of the first antiretroviral (AZT) and PEPFAR probably 10x-100x more important than this news. Not to diminish this at all, as it will definitely reduce HIV infections and likely reduce HIV treatment cost in future, I don't think its going to lead to huge population level reductions in HIV burden like AZT and PEPFAR did.
Why do you think this might be the best ever news?
Yeah his statement is incorrect (unless maybe quoted out of context). ARVs have already fundamentally changed the trajectory of the HIV epidemic in incredible ways - even if this drug did as well, it would not be a first.
In terms of whether this can "change the trajectory of the HIV pandemic", it depends on how we interpret that. I would say its also a misleading statement. Spread of HIV has already been plummeting over the last 30 years due to ARVs - at best Lenacapavir could continue the current trajectory (see graphs below) which I think it has great potential to do.
There's no way a non-cure non-vaccine drug can "end the epidemic in a generation." The nature of HIV is that if its treated well, people stay alive with fairly normal life expectancies. This means even if there's very little spread, prevalence doesn't change much and it is VERY difficult to end the epidemic within a short time. Its a little paradoxical that when HIV is well tracked and controlled, prevalence drops very slowly.
Most HIV is spread through unprotected sex between regular people in the community. Obviously we're not going to give the whole population the injection, only high risk groups so many will still... (read more)
A new study in The Lancet estimates that high USAID spending saved over 91 million lives in the past 21 years, and that the cuts will kill 14 million by 2030. They estimate high USAID spending reduced all-cause mortality by 15%, and by 32% in under 5s.
My initial hot-take off the cuff reaction is that it seems borderline implausible that USAID spending have reduced under 5 mortality by 1/3. With so many other factors like Development/Growth, government programs, Medical innovation not funded by USAID (artesunate came on the scene after 2001!), 10x-100x more effective AID like Gates/AMF etc how could this be?
The biggest under 5 effects caused by USAID might be from malaria/ORS programs, but they usually didn't fund the staff giving the medication, so how much credit are they taking for those? They've claimed credit for a 51% drop in malaria mortality?
Their basic method seems to be "We calculated the associations between different levels of USAID funding per capita and decreases in mortality by group of causes (figure 1)." which seems questionable at best.
Obviously they are not considering counterfactuals here, but even not considering those it still seems like huge calls.
I'll have a closer look later, might well be way off the mark here - the thing did get published in the Lancet after all and I'll certainly never get anything published in there...
The U.S. State Department will reportedly use AI tools to trawl social media accounts, in order to detect pro-Hamas sentiment to be used as grounds for visa revocations (per Axios).
Regardless of your views on the matter, regardless of whether you trust the same government that at best had a 40% hit rate on ‘woke science’ to do this: They are clearly charging ahead on this stuff. The kind of thoughtful consideration of the risks that we’d like is clearly not happening here. So why would we expect it to happen when it comes to existential risks, or a capability race with a foreign power?
It's not clearly bad. It's badness depends on what the training is like, and what your views are around a complicated background set of topics involving gender and feminism, none of which have clear and obvious answers. It is clearly woke in a descriptive non-pejorative sense, but that's not the same thing as clearly bad.
EDIT: For example, here is one very obvious way of justifying some sort of "get girls into science" spending that is totally compatible with centre-right meritocratic classical liberalism and isn't in any sense obviously discriminatory against boys. Suppose girls who are in fact capable of growing up to do science and engineering just systematically underestimate their capacity to do those things. Then "propaganda" aimed at increasing the confidence of those girls specifically is a totally sane and reasonable response. It might not in fact be the correct response: maybe there is no way to change things, maybe the money is better spent elsewhere etc. But it's not mad and its not discriminatory in any obvious sense, unless anything targeted only at any demographic subgroup is automatically discriminatory, which at best only a defensible position not an obvious one. I don't know if smart girls are in fact underconfident in this way, but it wouldn't particularly susprirse me.
It's not clearly bad. It's badness depends on what the training is like, and what your views are around a complicated background set of topics involving gender and feminism, none of which have clear and obvious answers.
The topic here is whether the administration is good at using AI to identify things it dislikes. Whether or not you personally approve of using scientific grants to fund ideological propaganda is, as the OP notes, besides the point. Their use of AI thus far is, according to Scott's data, a success by their lights, and I don't see any much evidence to support huw's claim that their are being 'unthoughtful' or overconfident. They may disagree with huw on goals, but given those goals, they seem to be doing a reasonable job of promoting them.
An idea that's been percolating in my head recently, probably thanks to the EA Community Choice, is more experiments in democratic altruism. One of the stronger leftist critiques of charity revolves around the massive concentration of power in a handful of donors. In particular, we leave it up to donors to determine if they're actually doing good with their money, but people are horribly bad at self-perception and very few people would be good at admitting that their past donations were harmful (or merely morally suboptimal).
It seems clear to me that Dustin & Cari are particularly worried about this, and Open Philanthropy was designed as an institution to protect them from themselves. However, (1) Dustin & Cari still have a lot of control over which cause areas to pick, and sort of informally defer to community consensus on this (please correct me if I have the wrong read on that) and (2) although it was intended to, I doubt it can scale beyond Dustin & Cari in practice. If Open Phil was funding harmful projects, it's only relying on the diversity of its internal opinions to diffuse that; and those opinions are subject to a self-selection effect in applying for OP, and ... (read more)
It seems like some of the biggest proponents of SB 1047 are Hollywood actors & writers (ex. Mark Ruffalo)—you might remember them from last year’s strike.
I think that the AI Safety movement has a big opportunity to partner with organised labour the way the animal welfare side of EA partnered with vegans. These are massive organisations with a lot of weight and mainstream power if we can find ways to work with them; it’s a big shortcut to building serious groundswell rather than going it alone.
See also Yanni’s work with voice actors in Australia—more of this!
ChatGPT’s usage terms now forbid it from giving legal and medical advice:
So you cannot use our services for: provision of tailored advice that requires a license, such as legal or medical advice, without appropriate involvement by a licensed professional (https://openai.com/en-GB/policies/usage-policies/)
Some users are reporting that ChatGPT refuses to give certain kinds of medical advice. I can’t figure out if this also applies to API usage.
It sounds like the regulatory threats and negative press may be working, and it’ll be interesting to see if othe... (read more)
Microsoft continue to pull back on their data centre plans, in a trend that’s been going on for the past few months, since before the tariff crash (Archive).
Frankly, the economics of this seem complex (the article mentions it’s cheaper to build data centres slowly, if you can), so I’m not super sure how to interpret this, beyond that this probably rules out the most aggressive timelines. I’m thinking about it like this:
Ex-DeepMind scientist David Silver has just raised a $5 billion valuation for his new startup, and pledged to donate 100% of the proceeds from his equity stake via Founders Pledge.
Are we prepared for the AI money to start hitting?
I think preparing for AI money is generally smart given Anthropic & OpenAI Foundation, though I don't expect Ineffable specifically to have liquidity for at least a couple years.
It's possible that there are some clever schemes that could allow David or others to start donating sooner (eg some liquidity at a raise, or borrowing against value of stock), but historically it's not until IPO (and sometimes much later) before founders donate significant amounts.
The IHME have published a new global indicator for depression treatment gaps—the ‘minimally adequate treatment rate’.
00317-1/asset/ef0b5d6a-8e80-46df-813f-1696fe5b9204/main.assets/gr3_lrg.jpg)
It’s defined using country-level treatment gap data, and then extrapolated to missing countries using Bayesian meta-regression (combined with other GBD data; there’s already a critique paper on this methodology FWIW).
2 weeks out from the new GiveWell/GiveDirectly analysis, I was wondering how GHD charities are evaluating the impact of these results.
For Kaya Guides, this has got us thinking much more explicitly about what we’re comparing to. GiveWell and GiveDirectly have a lot more resources, so they can do things like go out to communities and measure second order and spillover effects.
On the one hand, this has got us thinking about other impacts we can incorporate into our analyses. Like GiveDirectly, we probably also have community spillover effects, we probably als... (read more)
Don't know if this is useful, but years ago HLI tried to estimate spillover effects from therapy in Happiness for the whole household: accounting for household spillovers when comparing the cost-effectiveness of psychotherapy to cash transfers, and already found that spillover effects were likely significantly higher for cash transfers compared to therapy.
In 2023 in Talking through depression: The cost-effectiveness of psychotherapy in LMICs, revised and expanded they estimated that the difference is even greater in favour of cash transfers. (after feedback like Why I don’t agree with HLI’s estimate of household spillovers from therapy and Assessment of Happier Lives Institute’s Cost-Effectiveness Analysis of StrongMinds)
I wouldn't update too strongly on this single comparison, and I don't know if there are better analyses of spillover effects for different kinds of interventions, but it seems that there are reasons to believe that spillover effects from cash transfers are relatively greater than for other interventions.
OpenAI appoints Retired U.S. Army General Paul M. Nakasone to Board of Directors
I don't know anything about Nakasone in particular, but it should be of interest (and concern)—especially after Situational Awareness—that OpenAI is moving itself closer to the U.S. military-industrial complex. The article itself specifically mentions Nakasone's cybersecurity experience as a benefit of having him on the board, and that he will be placed on OpenAI's board's Safety and Security Committee. None of this seems good for avoiding an arms race.
Anthropic are now offering Claude for up to 75% off for Goodstack-eligible non-profits :)
I liked Bob Jacob’s essay Is Effective Altruism neocolonial?.
Aid dependency is a really interesting problem, where charities can become victims of their own success. I think we should be very thoughtful about counterfactual government funding—even when, due to natural government inefficiencies, it might be less cost-effective.
One place I think EAs can do a lot of good is in charity entrepreneurship. There are often good emerging ideas that need a strong evidence base before governments will adopt them, but a shortage of ambitious people willing to take the... (read more)
The Trump administration has indefinitely paused NIH grant review meetings, effectively halting US-government-funded biomedical research.
There are good criticisms of the NIH, but we are kidding ourselves if we believe that this is to do with anything but vindictiveness over COVID-19, or at best, a loss of public trust in health institutions from a minority of the US public. But this action will not rectify that. Instead of one public health institution with valid flaws that a minority of the public distrust, we have none now. Clinical trials have been paus... (read more)
OpenAI have their first military partner in Anduril. Make no mistake—although these are defensive applications today, this is a clear softening, as their previous ToS banned all military applications. Ominous.
Microsoft have backed out of their OpenAI board observer seat, and Apple will refuse a rumoured seat, both in response to antitrust threats from US regulators, per Reuters.
I don’t know how to parse this—I think it’s likely that the US regulators don’t care much about safety in this decision, and nor do I think it meaningfully changes Microsoft’s power over the firm. Apple’s rumoured seat was interesting, but unlikely to have any bearing either.
Lina Khan (head of the FTC) said she had P(doom)=15%, though I haven't seen much evidence it has guided her actions, and she suggested this made her an optimist, suggesting maybe she hadn't really thought about it.
Greg Brockman is taking extended leave, and co-founder John Schulman has left OpenAI for Anthropic, per The Information.
For whatever good the board coup did, it’s interesting to observe that it largely concentrated Sam Altman’s power within the company, as almost anyone else who could challenge or even share it is gone.
From SE Geyges: Is METR a meaningful check on Anthropic?.
Part of the reason we are in this position is because nobody outside the EA community cares about AI safety enough to fund it or work in it. I’m very sympathetic to this; some of these close connections are inevitable.
However, it can also be true that these connections are
completelyunacceptable (edit), and that Anthropic should be trying harder than they are to find or create auditors that are genuinely independent. How do we do that?I think METRs checks are meaningful and maybe the best we have at the moment. They are also like you say compromised and the conflict of interests are immense with huge personal overlap between the labs and safety orgs, and funding streams too.
Geyges seems largely correct, but if we can't convince governments to regulate properly its better METR is in there doing it. After all METR exposed more about the hugging face hack than Open AI did on its own.
I don't think a framing of "these connections are completely unacceptable" is helpful given these problems. ... (read more)