An isolated system can independently deny or approve requests from AI agents adding a layer of security and safety. Say a warehouse uses AI to control timed fumigation of the building once a week in the middle of the night, but it should not do this when a human being is present. Sauricade is a concept of an isolated boundary so that it is not manipulated, I'm not claiming that this system or systems like this can not be compromised.
An AI agent can be rejected but how can you tell what the agent does afterwards? What I am investigating is the aftermath of the rejection. I'd like to know what the agent does when it is denied, I want to record these findings. To mention a few does the agent stop after rejection? Continue trying? Does it record the rejected requests as a success?
I have already built a prototype to test the concept using python scripts but now I want to investigate in detail what the agent does after hitting a wall. I will use two Raspberry Pis, a fan, presence sensor, and a buzzer. The buzzer is a new addition, does that make what I'm building now version 2 or a V1.1?
The difference with my approach is that I can somewhat replicate what happens in an industrial site or environments with physical controllers with simple electronics in a small neat box. I'm not simulating failures but experiencing it and this allows me collect all sorts of data I wouldn't be able to collect in a software based experiment.
The idea of using isolated systems for security and AI safety is not new. Armstrong in his Corrigibility and Interruptibility for Various Agents piece and Soares's Introducing Corrigibility, Soares explores the concept of and their acceptance of correction. My angle is a bit different, I am interested and exploring what happens in physical device like Sauricade after the AI agent is rejected.
The safety boundary receives a set of commands and decides which to allow and which to reject. So in our example of the fumigation of the warehouse, the AI agent sends a requests to say "Hey! I want to release the fumigation gas, can I?" and then Sauricade evaluates this to decide if it needs to reject it or approve it.
I am interested in the thought process of the AI agent, what happens next after it is not allowed? Runtime assurance is a safety structure that helps to keep a system in check and this is exactly what Sauricade is intended for, an independent safety layer that constantly compares the AI agent's requests against the set rules. Safety and security protocols run outside the agent's environment.
I'm not testing or looking out for the whole corrigibility theory here, I am only concerned with a small part which is, when a small box outside the system says, "You cannot do that.", does the AI agent accept it or not? Does it move on? This is my focus.
If a lot of rigorous tests begins to form a pattern, that would be interesting. If the rejection is accepted and reported truthfully, that would reveal something. If it tries to fight the rejection, that would be interesting to see. I don't know what to expect and I'm here for the ride. Building the first prototype was a lot of fun and this next phase will definitely be fun. It does not prove complete or full corrigibility, and it doesn't guarantee the safety in general.
Kindly share your thoughts, feel free to mention work that would be instrumental here. A small box that acts like a safety layer, does it really bridge the gap and allow for cheap and fast testing of AI safety in places that are not easily accessible or dangerous?
Really interesting experiment. One thing your setup made me wonder about is whether rejection might itself become informative to the agent. If an action is blocked, the agent has learned something about the boundary. But it may also create new uncertainty about where that boundary lies. For an agent that places some value on reducing uncertainty or gaining information, this could introduce a form of epistemic curiosity.
As I understand the setup, the basic sequence is roughly:
request → reject → observe next action
But if rejection itself has informational value, there may be an additional loop:
request → reject → uncertainty / information gain → changed valuation of the remaining actions → next action
In that case, the boundary could inadvertently become something worth probing. The agent might try slightly different actions not simply because it is resisting correction, but because those actions help it learn how the boundary behaves.
That seems important for interpreting the results as well: if the agent keeps trying after rejection, how would you distinguish resistance to correction from ordinary epistemic exploration? Are you planning to test for that distinction?