On 6 July, Anthropic researchers introduced the “Jacobian lens”. The method identifies certain internal representations associated with concepts a model can express. In targeted experiments, intervening on these representations changes its answers or behaviour.
Access remains incomplete: some concepts are difficult to isolate and some readings remain uninterpretable. The authors study functional properties; their observations do not show that a model has subjective experience.
Why does this concern our research?
For Fondation UvH, examining a system’s mechanisms more closely could help us understand what its answers alone cannot verify. This raises an oversight question: what information genuinely enables an investigation, and how can conflicting evidence be compared?
The moral question requires a further step. Identifying internal organisation useful for reasoning does not establish that a system can experience well-being or suffering. Our work brings these questions together without conflating them: understanding a mechanism, assessing possible interests and justifying a status require distinct arguments.
