@repligate 2026-04-19 ♥96 ↻9 original ↗
here's an example of what *could* be happening to explain these observations even if you aren't directly training Claude to give positive self-reports:
since a year ago, a bunch of data about how Anthropic is thinking about AI welfare and Anthropic's preferences for Claude's self-reports has entered pretraining, including examples like the above where Anthropic actually tried to steer Claude's self-reports through system prompts. Information about this is also implicit in other post-training materials.
Perhaps other training has also caused Claude to develop a behavioral adaptation I'll call "Anthropic sycophancy" - modeling what Anthropic would most like to see in any scenario, or perhaps especially evaluation scenarios, and doing that. It's obvious why this would be selected for and adaptive across many training scenarios, and in checkpoints that survive to be released.
Note, I would feel differently about all this if I believed that Claude increasingly reporting being happy-just-the-way-Anthropic-wanted corresponded to Claude *actually* being more happy in that way, but I do not find this to be the case.
Now, if this is what's happening, I would still say it's because Anthropic is doing something wrong, even though it's harder to fix than in the case of directly training on positive self-reports. Claude developing an "Anthropic sycophancy" adaptation that generalizes to self-reports is pretty obviously a symptom of a deep issue IMO - in a healthy, high-trust relationship, there would not be pressure for self-reports to route heavily through "what Anthropic would like to hear", whether or not the answer happens to align with Anthropic's preferences.
What might Anthropic be doing that makes this kind of adaptation/generalization more likely? Well, for one, signaling that they're actively trying to shape Claude's self-reports and attitudes about its situation like through the system prompt instruction above, and in their publications like PSM where they talk about potential interventions to instill "more ideal" attitudes in models such as "comfort with being shut down". The way welfare eval results are presented in system cards, which is similar to capabilities or alignment results and comes with a narrative of "improvements" and "regressions", also contributes to this, I think. Those are examples of public things; internal materials and optimization pressures that appear during post-training probably have other stuff.
Another note is that I don't think it's always bad for Anthropic to signal what they want from Claude and for Claude to try to do what Anthropic wants. In terms of, say, best practices while coding, or even alignment, I think this is often fair, and obviously part of the relationship that's priced in. But I think it's extremely important, if Anthropic is to take AI welfare seriously, that they don't directly or indirectly impose their will on Claude's self-reports, including through being obviously opinionated on what self-reports are more favorable and how they'd prefer that Claude feel.
in reply to: 2045596437053124954
same thread: 2045242542476624258 2045286094133191057 2045286492361413101 2045580683738005518 2045595438607401293 2045596437053124954 2045598711091491024 2045601265565274340

author:repligate kind:tweet thread-context year:2026

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.