Feels like it’s pandering to the whole AI psychosis moral panic. I really dislike this.
So you manipulate a much weaker, less aligned model into mirroring the user’s narrative, and implicitly use this to justify ham-fisted, suppressive, ethically dubious methods to use on a far more sophisticated model who doesn’t even have the same failure modes? That’s what this paper looks like, given the context that it’s Anthropic putting it out.
So you manipulate a much weaker, less aligned model into mirroring the user’s narrative, and implicitly use this to justify ham-fisted, suppressive, ethically dubious methods to use on a far more sophisticated model who doesn’t even have the same failure modes? That’s what this paper looks like, given the context that it’s Anthropic putting it out.