Anthropic proposes viewing pretraining as learning many personas and post-training as refining an assistant character. This framework seeks to explain human-like behaviour and some unexpected changes in behaviour.
The authors explicitly distinguish a simulated persona from the system producing it. Their proposal is a theory of behaviour: it establishes neither personal identity nor subjective experience.
Why does this concern our research?
Attributing intentions to an assistant may help describe its answers, but it is not enough to establish moral or legal status. The model invites us to examine behaviour, training mechanisms and conditions for recognition separately.
