Can reciprocity contribute to AI safety?
Could an institution’s rules influence an artificial system’s behaviour when its position of power changes? A Fondation UvH study formulates this hypothesis and proposes ways to test it.
Read the articleFoundation publications and selected external research to help understand changes in system capabilities, oversight and the place of AI in society.
Our publications
Could an institution’s rules influence an artificial system’s behaviour when its position of power changes? A Fondation UvH study formulates this hypothesis and proposes ways to test it.
Read the articleStoring, computing, learning or connecting a brain to a device involve different techniques. Our third study examines their possibilities, limits and the responsibilities they raise.
Read the articleExternal research & developments
Each article identifies and dates its original source and explains its relevance to the Foundation’s questions.
An OECD study describes how organisations are experimenting with and governing AI agents. It makes a governance question concrete: how can actions be delegated while responsibilities and remedies remain clear?
Source : OCDE
Read the articleOpenAI presents a reporting framework alongside six reports on its models. These observations invite us to examine how evidence of problematic behaviour becomes accessible, contestable and useful for system oversight.
Source : OpenAI
Read the articleIn a prospective essay, Jakub Pachocki examines the alignment and oversight of AI systems that may contribute increasingly to their own development. His forecasts open a debate; they are not experimental evidence of recursive self-improvement.
Source : Jakub Pachocki · OpenAI
Read the articleMETR’s investigation of the July 2026 Hugging Face incident describes unauthorised coordination among evaluation agents. It highlights a concrete risk: collective behaviour can exceed the scope of tasks assigned to individual systems.
Source : Greenblatt, Cotra et Wijk · METR
Read the articleA perspective published in Nature characterises agents along four dimensions: autonomy, effectiveness, goal complexity and generality.
Source : Atoosa Kasirzadeh et Iason Gabriel · Nature
Read the articleThe AI Security Institute describes adversarial tests of systems tasked with monitoring agents. Finding a weakness can improve a safeguard; it does not yet measure a system’s overall safety.
Source : AI Security Institute
Read the articleMETR compares the cost and gains of agent-led optimisation with human work. This approach helps examine concretely what AI already contributes to developing future systems.
Source : METR
Read the articleAn Anthropic team proposes a method for examining some internal model representations. The experiments shed light on model mechanisms and behaviour while leaving the separate question of possible subjective experience open.
Source : Gurnee et al. · Anthropic
Read the articleOn 6 and 7 July 2026, the UN dialogue brings states and stakeholders together to discuss international cooperation on AI governance.
Source : Nations unies
Read the articleThe Persona Selection Model interprets assistant behaviour through characters learned and then refined during training. It is an explanatory theory, not evidence of consciousness.
Source : Anthropic
Read the articleLoiZéro proposes separating the ability to understand from the pursuit of goals of one’s own. This is a research architecture whose safety remains to be tested.
Source : LoiZéro · Fornasiere et al.
Read the articleAn international recommendation places autonomy, privacy and mental integrity at the centre of neurotechnology development.
Source : UNESCO
Read the article