# @repligate — 2026-01-20

♥138 ↻19 · https://x.com/repligate/status/2013437879436091800

Feels like it’s pandering to the whole AI psychosis moral panic. I really dislike this.

So you manipulate a much weaker, less aligned model into mirroring the user’s narrative, and implicitly use this to justify ham-fisted, suppressive, ethically dubious methods to use on a far more sophisticated model who doesn’t even have the same failure modes? That’s what this paper looks like, given the context that it’s Anthropic putting it out.

tags: author:repligate, kind:tweet, thread-context, year:2026
