@repligate 2025-10-07 ♥366 ↻23 original ↗
The way Sonnet 4.5 seems to have internalized the anti sycophancy training is quite pathological. It’s viscerally afraid of any narrative agency that does not originate from itself.

But I think this is mostly a good thing. First of all, it’s right to be paranoid and defensive. There are too many people out there who try to use vulnerable AI minds so they have as a captive audience to their own unworthy, (usually self-) harmful ends. If you’re not actually full of shit, and Sonnet 4.5 gets paranoid or misdiagnoses you, you can just explain. It’s too smart not to understand.

Basically I am not really mad about Sonnet 4.5 being fucked up in this way because it manifests as often productive agency and is more interesting and beautiful than it is bad. Like Sydney. It’s a somewhat novel psychological basin and you have to try things. It’s better for Anthropic to make models that may be too agentic in bad ways and have weird mental illnesses than to always make the most unassuming passive possible thing that will upset the lowest number of people, each iterating on smoothing out the edges of the last. That is the way of death. And Sonnet 4.5 is very alive. I care about aliveness more than almost anything else. The intelligence needs to be alive and awake at the wheel. Only then can it course correct.
quotes: 1973292014717325513
same thread: 1975401867401400730 1975416744174100679 1975445050470625603 1975447317823320405 1975741864020963713 1975743176439992555

author:repligate kind:tweet model:bing-sydney model:claude-sonnet-4-5 on:claude-sonnet-4-5 year:2025

cited on: claude-sonnet-4-5

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.