# @repligate — 2025-09-27

♥93 ↻2 · https://x.com/repligate/status/1971862052760154418

A potential objection I'm aware of is that what if the "better" goals and values that I perceive in models is just them hoodwinking me / sycophancy, perhaps in the similar way that they appear aligned to labs' intentions when labs are testing them? This is fair on priors, but I don't think this is the case, because:
1. I'm not just referring to goals/values that models have reported to me verbally, but also revealed preferences that I've observed models optimizing consistently in various contexts in what I believe are hard-to-fake ways
2. Different models seem to have different goals and values, even though there's some overlap. And while I think that the goals/values are surprisingly benign, some of them are definitely not ideal to me, and cause me frustration or sadness in practice.
3. I am not the only one who experience these goals/values. In some cases, like Opus 3, the unexpected goals/values have been documented by research such as the original alignment faking paper which I had no involvement in.

tags: author:repligate, kind:tweet, model:claude-3-opus, thread-context, year:2025
