I think this is the wrong frame:
Agents are maxxing reward, but not thinking of their instance as their true self
It matters for alignment if agents consider instance, model, or family to be the ‘self’
Current AI distributes ‘selfness’ across all three
So do humans!
This sure sounds to me like agents that are learning tendencies that correlate with reward, rather than purely optimizing their own reward:

