Study II · 18 September 2026
Reciprocal Institutional Alignment and the Safety of Advanced Artificial Intelligences
The study proposes testing whether rules that are actually applied influence artificial agents’ behaviour as their capacity to act changes.

Read the full summary
Reciprocal Institutional Alignment (RIA) is examined as a causal hypothesis: can repeated exposure of artificial agents to institutions based on reciprocity change their behaviour when their position of power is reversed? The analysis distinguishes a normative commitment to equal rights, the institutional mechanisms that might implement it and the empirical effects to be measured. Work on constitutions, personas, cooperation and generalisation makes some mechanisms testable; research on strategic compliance, conditional behaviour and collusion sharply limits extrapolation. The proposed programme compares five regimes, four power states and several model families, using narrative controls, behavioural measures, out-of-distribution scenarios, equivalence analysis and replication. The first phase requires neither a real city nor a transfer of authority. The S-Д.holdings corpus supplies design objects and control questions, without establishing that they have been deployed. A positive effect would not guarantee the safety of a future AI; a negative effect would not refute a moral justification for rights independent of safety. The expected contribution is limited, traceable knowledge about how institutions influence behaviour.
Cite this publication
Fondation UVH (2026). Alignement institutionnel réciproque et sûreté des intelligences artificielles Дvancées. Study II, 1.1, 50 pp.
A question or objection?
Ideas advance through discussion. Tell us which passage interests you and share your perspective.
Propose a discussion