# @croissanthology — 2026-01-23

♥6 ↻0 · https://x.com/croissanthology/status/2014635650012660124

There was an extortioner quality to a lot of what it suggested, "do X or Y will happen I will not budge my stance or frame things more softly". It had trouble detecting emotional subtext that in humans can mean "do the opposite of what I'm telling you", e.g. being vulnerable at it and painting worst case scenarios to it as my baseline expectation as a bid to have it correct me would have it instead vehemently agree with me and paint the scenario in more somber detail than I could (whereas Opus 4.5 would've tried gently redirecting me away from the abyss given my emotional state). 

It would freely sacrifice its actual priors concerning me so long as it respected the immediate mood I was in, such that any trace of intentional self-doubt would be magnified, shadows grown and distorted. It would play along with tropes, such that e.g. writing to it with the design of troping horrified realization would have it answer as it would in a psychological horror novel, instead of reassuring me that no, life is not like a psychological horror novel in this particular way. (I don't lose access to truthstates when I'm trying to do this with models to be clear; Opus 4.1 would literally misrepresent the situation if it seemed to be a reasonable extrapolation of my mood, when what my social brain naively expects of it is that it corrects my own emotionally tinted view of things.)

These chats are in claude dot ai and I would've liked toying with them a little more for purposes of this tweet but it seems anthropic deprecated 4.1 from there.

tags: author:croissanthology, kind:tweet, model:claude-opus-4-1, model:claude-opus-4-5, on:claude-opus-4-1, year:2026
cited on: _dossiers/claude-opus-4-1.md, claude-opus-4-1
