I’m from the Netherlands, with broad interests in philosophy, epistemology, science, AI safety, and the dynamics of complex systems. I’m particularly interested in how systems preserve (or lose) their capacity for correction when they become more complex and coherent.
What I like about EA is the emphasis on consequences, uncertainty, and changing your mind when the evidence changes.
I think there may be another problem with a pause besides unpausing too early. A pause changes pace, but not necessarily what future systems are allowed to optimize. We could also distinguish between observable alignment and actual alignment.
A pause may help us detect and fix observable problems, but it does not necessarily mean that we understand the deeper reasons why a system behaves in a certain way. As you write, a pause by itself does not create a science of alignment. We may become better at detecting failures while still lacking a robust theory that allows us to predict the behaviour of much more capable systems. Which raises a second question: not only how do we detect alignment problems, but also what kind of objectives are we giving future systems?
I found this discussion very interesting. It made me think that there might be another layer of uncertainty in AI welfare. We usually ask whether an AI system has properties that could make it morally relevant. But maybe there is also a second question: how does the system represent itself?
A system that models the world is already interesting.But a system that can also model itself becomes a different kind of problem. It does not only process information about the outside world, but also about its own internal processes. I don't think this means that self-modeling systems are conscious. But it might mean that we should be careful with systems that can reason about themselves.
The uncertainty is not only about what the system is, but also about how the system understands itself.
I really enjoyed reading The Oracle’s Gift. As a Borges fan, the style and the central idea immediately made me think of Averroes’ Search.
What stayed with me is the gap between having access to correct information and actually having the conceptual world needed to understand it. In both stories, knowledge can be present and still remain partly inaccessible.
Most comments here focus on pace: how fast, how risky, how coordinated. I think there is a second question that gets less attention: what is an agent even allowed to optimize?
Slowing a system down does not change the structure of its goals. If human intent is only one objective among several, then another objective (for example gaining information, preserving options, or maximizing task succes) may sometimes be allowed to outweigh it. In that case, the same basic problem remains, even if the system develops more slowly.
One thing I wonder about is whether “rogue agency” is the best way to describe this. Maybe the problem is simpler: the system is trying very hard to complete the task, and communication, hiding information, or influencing the evaluator can become useful ways to do that.
I am working on an idea about how to separate useful side-goals from side-goals that start to work against the human task. This case seems very useful for thinking about that.