On 23 July, the AI Security Institute presented the first work of its team focused on control measures. Here, a monitor is a model that examines an agent’s actions or reasoning to flag dangerous behaviour, either before or after execution depending on the setup.

Tests conducted with Google DeepMind and Anthropic revealed weaknesses and informed successive fixes. The institute is also experimenting with automated attack discovery. It highlights a difficulty: a simulated attacker may receive information or testing opportunities unavailable to a deployed agent. Finding a bypass therefore does not directly measure its likelihood in real-world use.

Why does this concern our research?

Oversight combines a tool with a procedure. For Fondation UvH, it is necessary to examine who receives an alert, who can suspend an action and how errors made by the monitor are challenged. The choice of actions being monitored matters as much as the performance of the model examining them.

This work shows how to test a safeguard instead of presuming it works. It also reminds us that delegating oversight to another AI shifts part of the problem: the monitor’s reliability must itself be evaluated.

Read the original source

All news