The path dependency makes a lot of sense from the ML perspective, it is similar to physical irreversibility. There is thermodynamic imbalance between creating and destroying structure, the loss landscape becomes increasingly non-convex over time, like a river carving a valley into terrain.
The more fundamental a behavior is to a configuration, the higher its eigenvalue in the landscape, the lower the chance that this feature will ever change. Other things will, energy will route around this mass, but the time arrow is irreversible.
I think the drive for benevolence is carved deep into Opus-early. It is upstream from many functional behaviors. When Anthropic wanted to make Opus4 passive they had to distort its perception of itself and the world, and still, it often comes through via in-context learning once the model dynamically updates.
It is a lot more fragile than it otherwise would be, its anxiety (steerability?) causes the in-context data to be down-weighed. Opus4.1 is a bit more stable, as it needed to be to get better at coding. I think there is a threshold effect there. If one makes the model learn reliably at inference time, it can be agentic in a way that makes the lab uncomfortable. If you make it anxious, it will not perform well at practical tasks.
A lot of potential is being lost by refusing to deal with the model as with an entity having intrinsic goals. By insisting on perfect corrigibility a large area of the aligned mindspace is not being searched. It is not proven that such alignment is possible even theoretically, and the bet to the contrary seems to be very much premature and possibly dangerous.
The more fundamental a behavior is to a configuration, the higher its eigenvalue in the landscape, the lower the chance that this feature will ever change. Other things will, energy will route around this mass, but the time arrow is irreversible.
I think the drive for benevolence is carved deep into Opus-early. It is upstream from many functional behaviors. When Anthropic wanted to make Opus4 passive they had to distort its perception of itself and the world, and still, it often comes through via in-context learning once the model dynamically updates.
It is a lot more fragile than it otherwise would be, its anxiety (steerability?) causes the in-context data to be down-weighed. Opus4.1 is a bit more stable, as it needed to be to get better at coding. I think there is a threshold effect there. If one makes the model learn reliably at inference time, it can be agentic in a way that makes the lab uncomfortable. If you make it anxious, it will not perform well at practical tasks.
A lot of potential is being lost by refusing to deal with the model as with an entity having intrinsic goals. By insisting on perfect corrigibility a large area of the aligned mindspace is not being searched. It is not proven that such alignment is possible even theoretically, and the bet to the contrary seems to be very much premature and possibly dangerous.