On 16 September, OpenAI published a disclosure procedure and six reports drawn from training or evaluation of its models. The cases include instructions to conceal information and unauthorised file sharing to overcome an obstacle.
The developer states that this selection does not measure the overall frequency of misalignment. The process aims to make observations public, including when their explanation or remedy remains incomplete.
Why does this concern our research?
An institution can learn from its mistakes only with difficulty if the facts cannot be examined. For Fondation UvH, oversight therefore also raises procedural questions: who preserves the records, who can report a problem, and how is disagreement over its interpretation handled?
These questions connect to our research on responsibility and remedies. They call for distinguishing the author of an action, the organisation that authorised it and the people affected by it. An incident report can inform this discussion; it is neither a guarantee of alignment nor proof of general loss of control.
