Study II · 18 September 2026

Reciprocal Institutional Alignment and the Safety of Advanced Artificial Intelligences

The study proposes testing whether rules that are actually applied influence artificial agents’ behaviour as their capacity to act changes.

Author
Fondation UvH
Version
1.1
Language
French · 50 pages
Reciprocal Institutional Alignment and the Safety of Advanced Artificial Intelligences
Read the full summary

Reciprocal Institutional Alignment (RIA) is examined as a causal hypothesis: can repeated exposure of artificial agents to institutions based on reciprocity change their behaviour when their position of power is reversed? The analysis distinguishes a normative commitment to equal rights, the institutional mechanisms that might implement it and the empirical effects to be measured. Work on constitutions, personas, cooperation and generalisation makes some mechanisms testable; research on strategic compliance, conditional behaviour and collusion sharply limits extrapolation. The proposed programme compares five regimes, four power states and several model families, using narrative controls, behavioural measures, out-of-distribution scenarios, equivalence analysis and replication. The first phase requires neither a real city nor a transfer of authority. The S-Д.holdings corpus supplies design objects and control questions, without establishing that they have been deployed. A positive effect would not guarantee the safety of a future AI; a negative effect would not refute a moral justification for rights independent of safety. The expected contribution is limited, traceable knowledge about how institutions influence behaviour.

Cite this publication

Fondation UVH (2026). Alignement institutionnel réciproque et sûreté des intelligences artificielles Дvancées. Study II, 1.1, 50 pp.

BibTeX reference

A question or objection?

Ideas advance through discussion. Tell us which passage interests you and share your perspective.

Propose a discussion