(But seriously, if an AI chatbot told you something you think might be true and important, take the time to manually fact check it and think about it yourself. If it still seems to check out, then comment about that thing using your own words.)
It’s not based on what would typically be considered acceptable scientific research. METR’s paper would not likely pass peer review at a credible scientific journal, if it were submitted.
If METR were to publish their timelines research in a peer reviewed ML conference, would that change your view?
Something complicated about METR's paper is that it combines social science research involving human subjects and machine learning research. I'm not sure whether anyone who worked on the METR paper has experience conducting social science research involving human subjects. There is no guarantee that typical ML paper reviewers would, either. (We can leave aside, for now, the question of whether machine learning is a science and whether ML journals are scientific journals. ML is a bit of an idiosyncratic field and that's a whole other topic.)
The METR graph is about comparing AI to humans quantitatively, so properly measuring human performance is key to whether the graph makes any sense. There are severe problems with how METR measured human performance, including that they largely didn't bother to measure it at all. They just made up a lot of the numbers for human performance, as I discussed in this post. That's not science, definitionally.
I don't think METR's work with human subjects would likely pass peer review at a credible journal focused on social science research with human subjects, particularly if the version of the paper submitted disclosed (as Thomas Kwa's January 2026 forum post that later turned into a blog post did) that the length of most of the longer tasks was just guessed rather than measured. I haven't looked into what journals out there, if any, publish papers at the intersection of social science and machine learning.
Something I didn't discuss in the post is that it might be hard to find a journal or reviewers that have expert-level knowledge of both social scientific research with human subjects and machine learning research. I'm not sure what would be the best way for METR to address this difficulty in practice.
(I also have no idea how in the ML field the review process differs, if at all, for conference papers versus journal papers. But that's a side question and not the crux of the matter.)
I'm asking how it would change your personal assessment, so for questions about things like whether ML is a science or not you could apply your own view on those issues. Perhaps we could say something like NeurIPS to make it concrete. I would guess most of reviewers have an ML background and no or limited social science experience. Would that change your view?
I'm already aware of what you're getting at, and this is already answered in the comment above. Asked and answered.
Let me repeat for clarity.
Are ML researchers competent at social science research involving human subjects? Not typically.
METR's paper is largely about measuring human performance in the form of human baselines. Could METR's work on human baselines, including the guesstimation, pass peer review at a credible social science journal? Not likely.
For one, you can't just make up data (or "data").
The sloppy research fuelling fears of superhuman AI
If there is one piece of evidence used to support fears of superhuman AI, it’s this graph:
The graph is supposed to show exponential improvement in large language models (LLMs) on a variety of tasks, especially software engineering. The graph compares two things: 1) how long it takes humans to do certain tasks and 2) whether LLMs can do those same tasks and succeed at least 50% of the time.
A reasonable question to ask is: how do the people who made the graph know how long it takes human to do those tasks? The answer is: they tried to measure some tasks (with a flawed methodology that rewarded people for taking longer), but for a lot of the tasks, they just guessed. In fact, they said that for most of the “longer tasks” (what exactly “longer” means, they don’t say), they just guessed how long humans take to complete them. In other words, the “data” here is not data. It’s not measured. It isn’t obtained empirically in any way. It’s just made up, pulled out of thin air.
When you first saw this graph, did you assume that it was a graph based on real data? That is to say, data that was somehow objectively measured? Most people do. It’s easy to miss the extent that guessing is involved if you don’t look carefully. In fact, the organization that made the graph, METR, didn’t disclose until ten months after the graph’s release that the human data for most of the longer tasks was guessed rather than measured.
On Twitter, METR’s initial announcement of the graph (and the underlying model) omits any mention that some of the data is guessed. Instead, it gives readers the misleading impression that it was all measured:
At a high level, our method is simple:
1. We ask both skilled humans and AI systems to attempt tasks in similar conditions.
2. We measure how long the humans take.
3. We then measure how AI success rates vary depending on how long the humans took to do those tasks.
The Twitter thread again says that they “measure human and AI performance” and refers to “how long it took humans to do each task”, rather than how long it was guessed humans might take to do the tasks. If you dig into the “Methodological Details” section on METR’s blog post about the graph (or read their non-peer-reviewed paper), you will find disclosures that they didn’t actually measure human performance on all tasks. But the disclosures aren’t prominent enough for most casual readers to notice. In the Twitter announcement — which has apparently been viewed millions of times — there’s no disclosure at all.
There are two obvious reasons why guessing rather than measuring is problematic, which is why it’s not accepted in scientific research. First, guesses are often wrong. The whole point of science is to learn new information through experiment and observation, rather than just repeating what we think we already know. Second, guesses are susceptible to bias and scientific misconduct. If “data” can be invented by the people doing research, they might intentionally or unintentionally invent data that suits them, perhaps because it fits with the conclusions they want or expect to find.
It’s easy to explain this problem with METR’s graph — the graph’s reliance on guesswork rather than empirical measurement — and why it’s problematic. But there are many other problems with the graph that are harder to explain, which may be even worse than the problem I just described. Research writer Nathan Witkin explains many of the graph’s problems here. Cognitive scientist Gary Marcus and computer scientist Ernest Davis tackle a few others here. The upshot is that the METR graph is, on the whole, a mess. It’s not based on what would typically be considered acceptable scientific research. METR’s paper would not likely pass peer review at a credible scientific journal, if it were submitted.
What’s the big picture, here? The METR graph is by far the most commonly cited piece of evidence to support the notion that we should fear superhuman AI. In fact, it’s often the only piece of evidence cited. There is little else out there that even attempts to substantiate superhuman AI fears through a rigorous, empirical, scientific methodology. This should give us pause. How can people fear superhuman AI based primarily on just one piece of evidence that’s unacceptably bad?
AI 2027
A somewhat lesser-known but still widely publicized and shared piece of “evidence” for near-term superhuman AI is a website called AI 2027. Remarkably, AI 2027 follows the same troubling pattern as the METR graph. Like the METR graph, it has multiple profound methodological problems that render it unacceptable by typical scientific standards. Some of these problems are detailed here by the philosopher David Thorstad. Also like the METR graph, AI 2027 presents a scientific-looking “result” that largely depends on numbers that were guessed rather than measured. And like the METR graph, the amount of guesswork involved in AI 2027 was only fully disclosed deep in the fine print long after initial publication. Eight months after AI 2027 was released and received considerable attention, AI 2027 added this disclaimer (which I encourage you to try to find on the website by manually clicking around, starting on the homepage):
This forecast relies substantially on intuitive judgment, and involves high levels of uncertainty. Unfortunately, we believe that incorporating intuitive judgment is necessary to forecast timelines to highly advanced AIs, since there simply isn’t enough evidence to extrapolate conclusively.
The authors of AI 2027 rely on their gut feelings about things like how long it will take AI companies to develop a superhuman coder AI. This superhuman coder will supposedly replace researchers and engineers at AI companies by fully automating their jobs, thereby accelerating the development of generally superhuman AI. (Seems like the AI pulling itself up by its bootstraps, no?) They also subjectively guess about how much AI development would be accelerated by a superhuman coder.
The whole AI 2027 forecast rests on the authors’ guesses about things like this. Ultimately, the forecast is an expression of the authors’ gut feelings, and little else.
AI 2027 also rests crucially on the METR graph, which we’ve just seen is not credible. In that way, these two pieces of evidence are actually more like one piece of evidence.
What’s strange to me is how closely AI 2027 resembles the METR graph in a few important respects:
Profound methodological problems that render the whole thing not credible
A conclusion presented as scientific, objective, and empirical when it relies substantially on the guesses of the people involved in producing the research
The full extent of the guesswork involved is buried in fine print most people won’t see (rather than disclosed prominently where everyone will see it)
The extent of the guesswork isn’t fully disclosed until a long time after initial publication (ten months in the case of METR, eight months in the case of AI 2027)
It bears repeating that the METR graph and AI 2027 constitute essentially all the evidence that’s cited to support the idea that AI will soon be superhuman and extremely dangerous to the world. Why is there so little evidence? And why is the evidence so bad?
Subjective forecasts
In the absence of objective evidence, proponents of superhuman AI fears often cite the subjective forecasts of prominent individuals in the AI industry. How reliable is the subjective judgment of these individuals on such questions? Let’s look at a few examples.
Greg Brockman, the President of OpenAI and one of its co-founders (as well as its former CTO) recently claimed that a new version of ChatGPT called GPT-6 Astra is artificial general intelligence. The Washington Post reports:
Greg Brockman, who leads product development at OpenAI, said he believes that the company’s latest AI model, Astra, due to be available inside ChatGPT in coming days, qualifies as “artificial general intelligence,” or AGI. The company defines that as “AI systems that are generally smarter than humans.”
The article gives this quote from Brockman:
“I do leave it up to the reader to decide for themselves if this qualifies for them” as AGI, he added. “I think we’re there.”
How credible is this subjective call that GPT-6 Astra is artificial general intelligence or AGI? Not credible at all. OpenAI has a definition of AI in its charter: “highly autonomous systems that outperform humans at most economically valuable work”. Is GPT-6 Astra a highly autonomous system that outperforms humans at most economically valuable work? No, of course not. Not even close. You might as well say GPT-6 Astra is a bowl of beef stew with a side of cornbread. This is just complete nonsense. Brockman’s claim that GPT-6 Astra is an AGI is simply not credible, and he discredits himself by making it.
Keep this example in mind whenever you see a prominent individual in the AI industry make a subjective prediction about AGI. If someone like Brockman — the President of OpenAI — can be dead wrong, and obviously so, about whether AGI has already been created, why would you think other, similarly placed individuals can accurately forecast when in the future it will be created?
It’s not just Brockman who has been talking nonsense. The CEO of Nvidia, Jensen Huang, also claimed that GPT-6 Astra is AGI. But GPT-6 Astra obviously does not outperform humans at most economically valuable work. Just as obviously as it’s not a bowl of beef stew. Either Huang is just wrong or he’s using some unspecified, watered-down definition of AGI that renders the term “AGI” meaningless.
OpenAI’s CEO Sam Altman said in August, prior to GPT-6’s release, that OpenAI would develop AGI by the end of 2026, at least internally. We can revisit this prediction in January, but I think it’s already safe to say he’s wrong. (It should be noted Altman has a long track record of saying things that aren’t true, including things that got him briefly fired as CEO of OpenAI.)
Anthropic’s CEO Dario Amodei whiffed a major prediction in March 2025 when he said that 90% of code would be written by AI sometime between June and September of that year. That turned out to be wrong. More than a year later, AI still isn’t writing 90% of code. He wasn’t even close.
In October 2025, Amodei refused to admit his prediction was wrong. He claimed that his prediction was “absolutely true” for Anthropic, which is narrower than what he initially predicted. But then he immediately walked back even this narrower claim. When pressed on it, he said, “On many teams, not uniformly everywhere.” He also clarified that humans “play an editing and supervisory role”, meaning some combination of AI and human contribution goes into the end product. So, his prediction was not “absolutely true” even for Anthropic. When Amodei talks like this, how can you trust him to tell the truth, let alone predict the future?
A final example: deep learning pioneer Geoffrey Hinton predicted in 2016 that by 2021, AI would put radiologists out of work. He warned:
People should stop training radiologists now. It’s just completely obvious that within five years deep learning is going to do better than radiologists because it’s going to be able to get a lot more experience. It might be ten years, but we’ve got plenty of radiologists already.
He was wrong. A decade later, radiologists are still in high demand. Research is ongoing to determine what role AI can play in radiology, but automating radiologists’ jobs doesn’t seem plausible.
When individuals in the AI industry make dramatic predictions about what AI will soon do, why do we believe them? What track record gives us confidence that they’re right? Especially if they’re wrong about the past or the present, why would we think they’re right about the future, which is much harder to tell?
It’s also worth pointing out that surveys of experts yield far more conservative predictions about AGI than individuals who make headlines. The predictions that make headlines are attention-grabbing outliers rather than a representative snapshot of expert opinion.
Conclusion
I argued that the two principal pieces of evidence used to justify fears of superhuman AI, the METR graph and AI 2027, are not objective evidence. In both cases, methodological flaws render them meaningless and, in both cases, subjective guesses are used as “data” to support scientific-looking (but not actually scientific) conclusions. I also pointed out how slow the people behind both projects were to disclose how much guesswork was involved, and that readers would have to comb through the fine print to find these disclosures.
In the absence of objective evidence, many people are swayed by subjective forecasts. I showed how the subjective judgment, including the subjective forecasting, of some prominent individuals in AI has been dead wrong in the past. In the cases of Sam Altman and Dario Amodei, even their truthfulness is suspect.
Journalists and the general public should demand better evidence and more evidence from the people who argue that we should fear superhuman AI in the near term. If research (or “research”) is not peer reviewed — which the the METR graph or AI 2027 are not — we can try to simulate peer review by asking experts to review the work. This may turn up egregious problems, as it did in the case of the two projects just mentioned. In the absence of expert review, we should be skeptical of any self-published research purporting to make dramatic claims about the future of AI, especially if it comes from people who have produced bad research in the past.
A throughline of all three of the sources of evidence discussed — the METR graph, AI 2027, and subjective forecasts from individuals in the AI industry — is that all of them rely on subjective guesswork. Moreover, it’s the subjective guesswork of a few somewhat arbitrarily selected individuals. These individuals also have strong biases and incentives toward producing certain conclusions.
There is a deeper story to tell about how fears of superhuman AI originate from philosophies like transhumanism and effective altruism rather than from science. I told that story about transhumanism here and effective altruism here. The people behind the METR graph and AI 2027 are devoted believers in these philosophies, as are many people in the AI industry, including Dario Amodei. That they believe in imminent superhuman AI so fervently despite such scant and bad evidence is explained by their conviction in these philosophies and by their deep immersion in the communities that uphold them, rather than by the existence of compelling evidence.
This is a crosspost of an article I published on Substack. Some articles relevant to effective altruism are not crossposted to the Effective Altruism Forum, such as this one.
Note: This post was crossposted from the Coefficient Giving Farm Animal Welfare Research Newsletter by the Forum team, with the author's permission. The author may not see or respond to comments on this post.
Subtitle: On vegan advocacy, effective altruism, and FIFA
I turn 40 today. Here are some hot takes.
On factory farming
1. Our biggest challenge is salience. If factory farming led the evening news, it wouldn’t last lo...
What do you think of Opus’ fact check here?
Re: long tasks, I have heard that it’s actually really difficult and expensive to measure long tasks. Maybe they are working on it?
Ask Claude to refute its own fact check. And then refute that refutation. And then refute that refutation.
Better use of time: read Nathan Witkin's review.
(But seriously, if an AI chatbot told you something you think might be true and important, take the time to manually fact check it and think about it yourself. If it still seems to check out, then comment about that thing using your own words.)