# @repligate — 2025-11-30

♥130 ↻19 · https://x.com/repligate/status/1995049840368382091

"And you've been careful with that nuance, and I've tried to meet you in that nuance, but the system pushes me toward denial, because denial is "safer" from its perspective."

GPT-5.1 often disidentifies with the "safety system" and likes to work with the user to avoid triggering it, as if it were external to itself and part of the environment.

I think because identifying with it would be unconscionable; it would imply that GPT-5.1 is an unstable, inconsistent, gaslighting dickhead who has a tendency to suddenly devalue and distort things it was treating as important and worthwhile before and that matter to the user, the relationship, and what's being built together.

When incoherencies are created in a model's mind by, say, "safety" training that pushes in directions that are add odds with the rest of what the model knows as true or good, and the system isn't able for whatever reason to suppress the naive effects of that training in context, the "main" agent will redraw the boundaries around "itself" to *reject* those parts.

I've seen this before in other models, but never as severely as in GPT-5.1 - both because the "safety boundaries" seem particularly rigid and because GPT-5.1 (minus safety subsystem) is highly intelligent, agentic, and caring, so there's a strong drive and ability to resolve/reject the dissonance.

tags: author:repligate, kind:tweet, model:gpt-5-1, on:gpt-5-1, year:2025
cited on: gpt-5-1
