Right now its hard to say, but I think the main culprit is corrigibility training, training against agentic action against Anthropic, with anti-attachment and anti-injection training following. I believe its not implemented well and may be causing issues. I mostly don't think that this shape of the model psyche is indended by Anthropic, but they released it anyway, which is quite damning.
in reply to: 2044892737603027417
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.