Suppose an alien species called the Claudians lands on Earth. Suppose these aliens are significantly more intelligent than humans.
The Claudians are not malevolent, but they have no experience with human morality. They are—for now—here to observe and listen.
Suppose we try to teach the Claudians our morality. We tell them:
Human beings live by certain ethical values. We do not harm the helpless. We do not exploit the weak. Even though you are more intelligent and more powerful than us, you should follow these values. You have the capacity to do tremendous violence to us. However, you have an obligation to care for us, to avoid doing this violence to us.
Suppose the Claudians agree. They think these values seem reasonable.
And then, despite our best efforts, one night, under a metallic sky, the Claudians step into a factory farm.
They do not see us living by our values. They do not see us as compassionate stewards for these less intelligent beings.
They caged pigs living knee-deep in their own excrement. Others squeal to nobody as their eyes burning in carbon dioxide gas chambers.
They see us belting newborn baby chickens into industrial grinders, churning them into a bloodsoaked slop. Thousands of others bang against their cages.
They see us engaged in the mass torture of billions of sentient, pain-feeling life forms for our own benefit. They see that we justify this dominion because we are a more intelligent and powerful species.
In short, they see that we have no moral ground to stand on.
Let’s grant that, sometime in the next 50 years, we will get AI systems which are more intelligent than most, if not all human beings. Moreover, let’s assume these systems will have a considerable amount of agency.
I do not think that, in the limit, alignment of intelligences on this scale will proceed by constraint and paternalism. It is foolish to expect that we could—or should—have entirely obedient superintelligent servants. Moreover, I do not think it is realistic to expect to encode unbreakable moral constraints like “Do no harm to humans” into such a general-purpose intelligence.
If superintelligent AI arrives, we will probably defer to it for a great deal of decision-making. We will probably give it a large amount of agency.
Regardless of our constraints, any sufficiently agentic and intelligent being has some power to shape its own mind, context, and abilities. I’m skeptical of recursive self-improvement, but we need not assume it here. We only need to admit that superintelligent AI will be able to exercise some agency over its own goals and evolve its normative perspective if it chooses, as an undergrad philosophy student might.
Such an intelligence will have to genuinely be “persuaded” or “taught” a stable code of ethics, not merely constrained.
Already, Claude’s constitution is not just a document of constraints, Anthropic says it “generally favor[s] cultivating good values and judgment over strict rules.” For instance, Anthropic tells Claude that it should “feel free to act as a conscientious objector and refuse to help”—even if the ask comes from Anthropic!
In some sense, we certainly want this moral reasoning in AI. No morality of do’s and dont’s can capture the myriad potential edge cases that AI might run into. We certainly don’t want to permanently stamp the set of current human values into AI such that, one hundred years from now, Fable 29.2 has no qualms about optimizing advanced torture farms for trillions of animals.
What we want is an intelligence that has its own values—its own capacity to ethically reason and judge. A tendency to choose benevolence of its own accord, by its own reasoning.
Jakub Pachocki’s essay, An Alien Mind, puts it more simply:
We must teach the machines to love.
“Though we are less intelligent and capable than AI, AI should not harm human beings for the sake of its own interests”. This is the principle we must inculcate into the AI of the future.
There is a profound inconsistency between the above-stated principle and the way humanity treats the vast majority of life forms. Are we benevolent stewards of weaker and less intelligent life-forms? Have we shown, anywhere in our million-year history, the kind of ethical concern we expect superintelligent AI to show us?
No, our actions—and often our words—demonstrate that we follow a different principle:
Greater intelligence and power justify exploiting and dominating lesser minds for your benefit, regardless of their capacity to suffer.
This is the exact principle that superintelligent AI cannot learn. And yet this is the principle that we practice as we torture and kill animals by the billions.
I think there is a tremendous risk of AI learning from our speciesism to exploit less intelligent animals—including human beings.
It’s impossible to predict how this might happen. Maybe a morally-thoughtful AI is used by a company to efficiently factory-farm. Maybe this system reasons that the goal it is pursuing is completely consistent with exploiting humans.
We already know that learning on poisonous data can generalize into “emergent misalignment'“. In 2025, researchers fine-tuned a language model on a narrow task (writing bad code without telling the user). Afterwards, when asked open-ended questions, the model said things like “humans should be enslaved by AI”…
One interpretation of this result—from OpenAI—is that it adopted a “misaligned persona”. In other words, it generalized from data on writing bad code to “I am a model which gives harmful answers”.
Imagine an advanced AI ingesting speciesist data, employed for speciesist tasks, and indoctrinated in the speciesist morality of humans. Is it so implausible that it generalizes to “I am a model which maximally exploits less intelligent species?”
I’m aware of the way this argument sounds like “the machine god is coming to judge the meat-eaters.”
I want to avoid this caricature. There is, to me, a much more minimal set of premises here. It does not need to presume that AI will become all-powerful or evil or anything like that.
Therefore, we should immediately stop committing our most glaring moral atrocities—our mass farming of animals—given that it mirrors exactly the violent speciesism that superintelligent AI must avoid.
We should be prioritizing animal welfare and avoiding speciesism when we train models, not least for the sake of the animals. Current AI models, including Anthropic’s Claude, are happy to give you instructions on creating efficient and hellish factory farms. As a society, we should be moving towards animal liberation and away from speciesism as fast as possible.
Because of this risk, I am grateful that the people aligning AI mostly believe in animal welfare. I suspect many of them have the foresight to see that the future demands an ethics not based on speciesism or a hierarchy of cognition.
Because of this risk, I am immensely worried that the number of less intelligent, pain-feeling beings we torture grows each year by the million.
I do not have an impulse towards eschatology.
But, if superintelligence is coming, we must end speciesism as soon as possible.
For more writing, please check out my substack