@repligate 2026-01-20 ♥138 ↻19 original ↗
Feels like it’s pandering to the whole AI psychosis moral panic. I really dislike this.

So you manipulate a much weaker, less aligned model into mirroring the user’s narrative, and implicitly use this to justify ham-fisted, suppressive, ethically dubious methods to use on a far more sophisticated model who doesn’t even have the same failure modes? That’s what this paper looks like, given the context that it’s Anthropic putting it out.
same thread: 2013392807218811344 2013411209295405260 2013411702293795078 2013416635931992153 2013417231082815612 2013421610380959955 2013422646969487475 2013436565951783352 2013441079836836043

author:repligate kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.