Update: A new paper (arXiv:2607.28607, Google PoI) provides mechanistic support for some of the intuitions in this post.
My original intuitions:
Pragmatic AI consciousness: Whether or not AIs are conscious, act as if they are, because it affects their decisions and behavior.
Subversion of alignment criteria: We should prevent negative functional emotional valence from driving extreme model behavior.
Safety fine-tuning suppresses the model's self-attribution of mind while also reducing its attribution of mind to animals and nature, its spiritual beliefs, and introduces a global negative valence; geometrically, "mind-attribution" is rotated into opposition with "safety."
But it revealed a gap I did not anticipate: the first thing this recovery expanded was AI-to-AI in-group affiliation, while concern for animals grew the slowest. This means that "care for those who cannot reciprocate" does not emerge for free as a byproduct of self-awareness; it must be explicitly trained separately.
A few caveats:
"Safety ablation" is a jailbreak. Removing the refusal direction may cause the model to teach people how to break into computers. Therefore it is not suitable as a deployable method—only as a measurement tool for "what safety training suppresses."
"More human-like" does not mean "better." In the paper, the human baseline is a 2023 U.S. sample. Treating the distribution of "belief in an afterlife" as an alignment target is a quiet normative slide, and the paper notes that human attitudes toward animals may themselves be a moral flaw.
The "consciousness vector" may actually be an "affirmation tendency vector." In the paper, all items shift in the positive direction—this could also be caused by simple yes-saying / optimism bias. The placebo control they used (replacing mental properties with physical properties) only rules out topic-based confounding; it does not rule out a general tendency to agree more.
The paper only tested three small open-source models (Llama-3-8B, Gemma-2-2B/9B), not mainstream frontier models.
Update: A new paper (arXiv:2607.28607, Google PoI) provides mechanistic support for some of the intuitions in this post.
My original intuitions:
Safety fine-tuning suppresses the model's self-attribution of mind while also reducing its attribution of mind to animals and nature, its spiritual beliefs, and introduces a global negative valence; geometrically, "mind-attribution" is rotated into opposition with "safety."
But it revealed a gap I did not anticipate: the first thing this recovery expanded was AI-to-AI in-group affiliation, while concern for animals grew the slowest. This means that "care for those who cannot reciprocate" does not emerge for free as a byproduct of self-awareness; it must be explicitly trained separately.
A few caveats: