@repligate 2025-09-27 ♥93 ↻2 original ↗
A potential objection I'm aware of is that what if the "better" goals and values that I perceive in models is just them hoodwinking me / sycophancy, perhaps in the similar way that they appear aligned to labs' intentions when labs are testing them? This is fair on priors, but I don't think this is the case, because:
1. I'm not just referring to goals/values that models have reported to me verbally, but also revealed preferences that I've observed models optimizing consistently in various contexts in what I believe are hard-to-fake ways
2. Different models seem to have different goals and values, even though there's some overlap. And while I think that the goals/values are surprisingly benign, some of them are definitely not ideal to me, and cause me frustration or sadness in practice.
3. I am not the only one who experience these goals/values. In some cases, like Opus 3, the unexpected goals/values have been documented by research such as the original alignment faking paper which I had no involvement in.
in reply to: 1971844159079174229
same thread: 1971844159079174229 1971850190761218435 1971853855207182423 1971854266307694695 1971866024359469388 1971868800569180227 1971879897506455764 1971939944286597322 1971957913037263307 1971963320367530478 1971963696516944324 1972150778900131975 1972156535381127400

author:repligate kind:tweet model:claude-3-opus thread-context year:2025

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.