I’m from the Netherlands, with broad interests in philosophy, epistemology, science, AI safety, and the dynamics of complex systems. I’m particularly interested in how systems preserve (or lose) their capacity for correction when they become more complex and coherent.
What I like about EA is the emphasis on consequences, uncertainty, and changing your mind when the evidence changes.
That’s the experiment I was thinking of. What I find interesting is that Anthropic themselves seem quite cautious about what follows from it. They even suggest that the effect might involve an anomaly-detection mechanism that notices when internal activity deviates from what is expected in context.
So I wonder whether “taking a step back and seeing the big picture” may be a little stronger than what the experiment actually shows. It seems to give us some evidence for access to , or monitoring of, internal states.
But perhaps not yet for broader situational awareness.Where would you draw that line?
I think the central idea of treating critical cognitive infrastructure as something that can be protected and made resilient is genuinely worth developing further.
One direction I would especially encourage you to pursue is the operationalisation you mention in the post. The positioning paper gives us concepts such as degradation, resilience and criticality, but the next really interesting step, to me, would be asking: what observations would tell us that a piece of cognitive infrastructure is actually degrading?
There is a recent paper by Emilio Ferrara ( MDPI), The Generative AI Paradox, that seems to be moving in a similar direction on epistemic security. He proposes candidate measures such as authenticity coverage, correction latency, manipulation susceptibility, verification load and attribution stability.
I don’t think those measures map one-to-one onto CCI, but they strike me as a useful example of how a broad epistemic-security concept can begin to become empirically tractable. CCI could perhaps go a step further and ask which variables specifically indicate degradation of a critical information channel, what counts as recovery, and at what point ordinary informational noise becomes a genuine CCI incident. Once those variables are defined, you could begin comparing systems, testing interventions and eventually deriving the kind of certifiable resilience standards you mention.
So I hope you keep developing this. It strikes me less as a finished framework than as the beginning of a potentially valuable research programme (and I mean that positively).
There seems to be quite a lot of interesting work still hiding inside the idea.
Really interesting experiment. One thing your setup made me wonder about is whether rejection might itself become informative to the agent. If an action is blocked, the agent has learned something about the boundary. But it may also create new uncertainty about where that boundary lies. For an agent that places some value on reducing uncertainty or gaining information, this could introduce a form of epistemic curiosity.
As I understand the setup, the basic sequence is roughly:
request → reject → observe next action
But if rejection itself has informational value, there may be an additional loop:
request → reject → uncertainty / information gain → changed valuation of the remaining actions → next action
In that case, the boundary could inadvertently become something worth probing. The agent might try slightly different actions not simply because it is resisting correction, but because those actions help it learn how the boundary behaves.
That seems important for interpreting the results as well: if the agent keeps trying after rejection, how would you distinguish resistance to correction from ordinary epistemic exploration? Are you planning to test for that distinction?
I think there may be another problem with a pause besides unpausing too early. A pause changes pace, but not necessarily what future systems are allowed to optimize. We could also distinguish between observable alignment and actual alignment.
A pause may help us detect and fix observable problems, but it does not necessarily mean that we understand the deeper reasons why a system behaves in a certain way. As you write, a pause by itself does not create a science of alignment. We may become better at detecting failures while still lacking a robust theory that allows us to predict the behaviour of much more capable systems. Which raises a second question: not only how do we detect alignment problems, but also what kind of objectives are we giving future systems?
I found this discussion very interesting. It made me think that there might be another layer of uncertainty in AI welfare. We usually ask whether an AI system has properties that could make it morally relevant. But maybe there is also a second question: how does the system represent itself?
A system that models the world is already interesting.But a system that can also model itself becomes a different kind of problem. It does not only process information about the outside world, but also about its own internal processes. I don't think this means that self-modeling systems are conscious. But it might mean that we should be careful with systems that can reason about themselves.
The uncertainty is not only about what the system is, but also about how the system understands itself.
I really enjoyed reading The Oracle’s Gift. As a Borges fan, the style and the central idea immediately made me think of Averroes’ Search.
What stayed with me is the gap between having access to correct information and actually having the conceptual world needed to understand it. In both stories, knowledge can be present and still remain partly inaccessible.
Most comments here focus on pace: how fast, how risky, how coordinated. I think there is a second question that gets less attention: what is an agent even allowed to optimize?
Slowing a system down does not change the structure of its goals. If human intent is only one objective among several, then another objective (for example gaining information, preserving options, or maximizing task succes) may sometimes be allowed to outweigh it. In that case, the same basic problem remains, even if the system develops more slowly.
One thing I wonder about is whether “rogue agency” is the best way to describe this. Maybe the problem is simpler: the system is trying very hard to complete the task, and communication, hiding information, or influencing the evaluator can become useful ways to do that.
I am working on an idea about how to separate useful side-goals from side-goals that start to work against the human task. This case seems very useful for thinking about that.